← Daming Wu Source on GitHub ↗

Tavus CVI take-home · August 2026

Talk naturally.
Leave with better English.
Coached by you.

A face-to-face English coach built on Tavus CVI — and the coach can be you.

Explore the case study ↓
Live Tavus conversationPhoenix-4 face · Sparrow-1 turn-taking · private Daily room
Evidence-labelled feedbackmeasured, observed, and suggested are never mixed
Your face, your voicethe live coach is the author’s own Face and cloned voice
Sessions survive their roomautomatic resume across provider room-duration caps

The product, in 92 seconds

Narrated tour over real session captures. Read the transcript · Or try the live product — every session runs a real Tavus room, so if the coach is busy, try again in a minute.

One loop, not a lesson

Language learners usually know more English than they can retrieve mid-conversation. Fluent Me keeps one continuous conversation primary and makes coaching a set of callable moves, not locked steps:

01TalkSpeak freely with a responsive video coach. Camera stays off unless visual coaching is useful.
02NoticeOpen a completed turn only when you want help. Every signal is labelled by source and certainty.
03PractiseHear the exact wording via conversation.echo, make two real attempts, change one thing.
04ReturnA grounded recap, a language review, and an opt-in phrase memory on a visible schedule.

Be your own coach

The most immersive reference speaker for a learner is themselves. When the model phrase arrives in your own voice, there is no accent, pitch, or timbre gap to translate across — imitation collapses into repetition. And talking to your own face removes the social cost of being wrong that makes real conversation practice stressful.

Fluent Me treats this as a first-class, consent-gated path: record ~60 seconds for a Phoenix-4 Face, a clean 60–90-second sample for an ElevenLabs Instant Voice Clone, and the server combines both into a personal Tavus PAL. The live deployment ships this way: its default coach is my own trained Face speaking with my consented voice clone, and the demo narration uses the same clone. Recordings stay local until explicitly submitted, provider keys never reach the browser, and only the resulting IDs are stored.

Progressive activation keeps it honest: face-only uses the stock voice, voice-only uses the stock face, and each step reports the provider's real status instead of a fake ready state.

Why Tavus is structural here

The face is not decoration. It gives the learner someone to address and imitate, while timing, interruption, perception, and coaching behavior work together in a way a text tool cannot:

Phoenix-4a person to listen to and imitate — and, personalized, that person is you
Sparrow-1interruptible, naturally timed turns; practice stays conversational
Raven-1 (limited)bounded delivery observations, framed as tentative — never “emotion detection”
PAL + interaction protocolconversation.respond for coaching, conversation.echo for exact models, conversation.utterance as the transcript source of truth

Architecture

Browsercustom learning UI; joins the Tavus-created private Daily room with a short-lived meeting token; microphone publishes straight into the Tavus pipeline
Server boundaryFastAPI (and a Worker twin for hosting) creates and ends rooms; Tavus, ElevenLabs, and Anthropic keys never reach the client
Evidence modeldeterministic transcript metrics, transient browser acoustics, and Raven observations are stored and displayed separately
Session continuityif a provider-side max_call_duration cap ends the room mid-session, the client reconnects and hands the new room a bounded continuation packet — the coach picks the conversation back up and the focus timer keeps accumulating
Privacyno server transcript store; local history and phrase memory are opt-in, compact, and deletable; every exit path ends the remote room

Honest boundaries

Browser waveforms and pitch are descriptive, auto-scaled signals — never pronunciation, accent, fluency, or emotion scores. Raven observations stay tentative. Missing evidence stays missing: the comparison and recap cite only what actually arrived, and there is no invented numeric score anywhere in the product.

Artifacts