Live · Real-time prosody analysis

Hear your voice.
See your emotions.

Emovo is a cross-platform system that streams your microphone audio to a hosted prosody AI model and renders the resulting emotion scores live — then archives each session for replay and long-term analytics. iOS, Android, and the web. One backend.

≤1s
visible latency
16 kHz
mono PCM-16 audio
3
platforms · iOS · Android · Web
48+
emotion dimensions
1
unified Firebase backend
What is Emovo

The gap between how you sound and how you feel — closed.

Speech prosody — pitch, intensity, rhythm, voice quality — carries a rich emotional signal that listeners decode unconsciously. Mainstream wellness apps rely on self-report (mood journals, PHQ-9), which is reflective and low-frequency by design.

One-tap capture

Tap the mic. Audio streams to Hume's prosody model and the top emotions update in real time alongside a rolling chart and a session timer.

Replay in sync

Every session saves the audio plus per-second emotion snapshots. Scrub the timeline and the chart, top emotion, and bars all re-synchronize to the playhead.

Long-term analytics

Streaks, mood-trend lines, duration-weighted top emotions, a 15-week × 7-day activity heatmap, and a peak-moments feed across all your sessions.

System Architecture

Three tiers. One data contract.

Both clients (Flutter mobile, React web) talk to a single Firebase backend for auth, data, files, hosting, and telemetry. A Firebase-Auth-protected Cloud Function provisions a per-session Hume credential, after which each client streams 1-second audio chunks directly to Hume's expression-measurement WebSocket and receives back emotion-score predictions in real time.

Clients
Mobile App
Flutter · iOS / Android
Web App
React + Vite + TypeScript
Firebase Backend
🔑
Authentication
Google sign-in
📄
Firestore
Session metadata
📦
Cloud Storage
WAV files · photos
⚡
Cloud Functions
getHumeCredentials
🌐
Hosting
HTTPS · global CDN
📈
Analytics
Usage events
☁️ Hume AI
Streaming Expression Measurement API
16 kHz mono PCM-16 1 s chunks 5 s sliding window
Clients ⇄ Firebase — sign-in, sessions, files, hosting, telemetry
Cloud Functions ⇢ Clients — Hume credential to authenticated callers only
Clients ⇄ Hume — direct WebSocket; audio up, emotion scores down
Tech Stack

Built with the tools we love.

A deliberate, narrow set of technologies — chosen so the same data contract serves mobile and web with no per-platform plumbing.

Flutter
Dart
React 18
TypeScript
Vite
MUI v6
Firebase
Cloud Functions
Node 22
Hume AI
Zustand
Riverpod
iOS
Android
Features

Designed for daily use.

Every feature lives on every platform. Sessions saved on the phone replay seamlessly on the web — and vice-versa.

01

Live capture

Gradient card for the dominant emotion with a delta arrow. A rolling 3-emotion line chart. Animated top-5 bars. A clean session timer. One-tap to start, one-tap to stop.

  • Web Audio API + AudioWorklet (web)
  • record package (mobile)
  • Identical byte streams cross-platform
Live capture screen
02

Sessions, replayable

Every completed session is stored with its full audio and per-second emotion scores. Tap a peak moment to jump straight to it. The chart, top-emotion card, and right-now bars all re-synchronize to the playhead.

  • WAV stored in Cloud Storage
  • Snapshots tagged with PCM-byte offset (no clock drift)
  • Same Firestore doc on every device
Sessions list screen
03

Analytics that respect your time

Streak, session count, total minutes. Mood-trend line chart over the three dominant emotions. Duration-weighted top-emotions bars. A 15-week × 7-day activity heatmap. A cross-session peak feed.

  • 7-day · 30-day · all-time ranges
  • Hand-rolled charts — identical pixel-for-pixel
  • Firestore onSnapshot · push updates
Analytics screen
Data flow

How a single second becomes a feeling.

  1. 1

    Mic → PCM

    Audio is captured natively, resampled to 16 kHz mono Int16 PCM inside the audio thread.

  2. 2

    Chunk + WebSocket

    Each 1-second chunk is wrapped in a WAV header, base-64 encoded and sent over a WebSocket with a payload_id.

  3. 3

    5 s sliding window

    Hume analyzes each chunk against a 5-second context so prosody continuity isn't lost between chunks.

  4. 4

    Live render

    Predictions arrive asynchronously and feed the rolling chart, top-emotion gradient, and animated bars at 60 fps.

  5. 5

    Persist

    On stop, the full PCM is wrapped into a single WAV, uploaded to Cloud Storage, and a session document is written to Firestore.

Engineering challenges

The hard bits — and how we solved them.

iOS / Safari

AudioContext starts suspended

Defensive audioCtx.resume() after the user-gesture so the worklet never silently no-ops.

AI

Empty predictions on isolated chunks

Use Hume's 5 s sliding window so each 1 s chunk is analyzed against continuous voiced context.

Replay sync

Wall-clock drift desyncs charts

Tag every snapshot with the cumulative PCM byte count → exact audio-clock offset for playback sync.

Cross-platform

Two codebases drifting apart

Single source-of-truth Firestore schema; mirrored module structure; same Material-3 seed colour #6750A4.

Results

What we shipped.

End-to-end realtime emotion streaming with sub-second visible latency on every platform — and full feature parity across iOS, Android, and web.

< 1s
Visible end-to-end latency
From mic capture to rendered emotion score, on both clients.
100%
Cross-platform parity
Every screen, analytics widget, and theme state present on web and mobile.
1
Firebase project
Hosting, Functions, Auth, Firestore, Storage — no separate infra.
⇄
Fully interoperable sessions
A session saved on the phone replays on the web — same Firestore doc, same WAV.
Future Work

Where Emovo goes next.

📊

Longer-horizon analytics

Weekly / monthly aggregates, day-of-week effects, calendar overlays.

💬

Empathic Voice Interface

Integrate Hume's conversational EVI endpoint for two-way voice interaction.

📱

On-device inference

Lightweight prosody models for offline / privacy-preserving capture.

⚙️

PWA & background sessions

Native PWA install on web; background-recording on mobile.

🧪

User study

Does passive voice tracking improve emotional self-awareness vs. mood-journal baselines?

Try it

Scan, speak, see.

Two builds, one experience. Pick the device in your hand and start a session in under five seconds.

Web App
React + Vite · Firebase Hosting
QR — Web app
Drop web.png
into assets/qr/
emovo-e06a8.web.app
Mobile App
Flutter · iOS & Android
QR — Mobile app
Drop mobile.png
into assets/qr/
Install on iOS / Android