Models#Video generation#Benchmark
Tavus Griffin: a video model that listens while it talks
Tavus unveiled Griffin on Oct 1, a full-duplex video-to-video model; 26 of 54 testers took it for a person after a one-minute call. Preview only.

Tavus announced Griffin on October 1, a real-time video conversation model it calls a “Human Interaction Model” (HIM): it listens, watches and speaks at the same time, and generates a person’s face, voice and body frame by frame from a single reference image. In a study Tavus ran itself, 26 of 54 people who believed they were talking to another participant (48%) said afterwards that their partner was a real person; the company’s previous system managed 1 of 41 (2.4%) in the same test. Only a research preview called Griffin-Lite exists, open to a small group of trusted testers, with no API and no price.
Key points
- Architecture: one full-duplex, video-to-video system where perception, decision-making and generation run concurrently, instead of a relay of speech recognition, a language model, speech synthesis and an avatar renderer.
- Self-reported result: in one-minute video calls, 26 of 54 participants (48%) judged their partner to be human; the previous Phoenix-4.5 stack scored 1 of 41 (2.4%).
- Third-party benchmark: on NVIDIA’s VideoFDB, Griffin-Lite scores 3.83 on generation (human reference 3.92) and 3.73 on perception (human reference 4.20), the highest on both leaderboards.
- Availability: Griffin-Lite only, for trusted testers; Tavus says wider release waits on disclosure and safety work.
Background
Tavus has built its business on a pipeline of separate models: Phoenix-4.5 renders the face, Raven-1 reads the user, and Sparrow-2 decides when to talk. Roughly 150,000 developers and businesses use that stack, and Tavus itself reports a pass rate of at most 2% in blind human tests.
The weakness of a cascade like that is the handoffs, not the picture quality. Tone, expression and whatever you hold up to the camera disappear once speech is turned into text, and the system waits for you to finish before it starts working. So a thinking pause gets read as “your turn”, or the avatar sits for a second or two before answering. Voice AI has already been through this shift: end-to-end realtime speech models replaced three-stage pipelines and gave assistants interruptions and tone. Griffin tries to carry the same shift from sound to picture.
The facts
Figures below come from Tavus’s own announcement unless noted.
Approach. Griffin has two parts. A continuous conversational model re-evaluates the state of the conversation at sub-second intervals and decides whether to speak, hold, yield or back-channel. An audio-visual generation engine turns those control signals into speech and 720p video. The video generator was distilled from a large many-step diffusion model; on H100s Tavus reports 0.43 seconds average delay from audio in to video out, half that of the next fastest method. Speech runs on an in-house codec, Tavec, and can clone a voice from about 10 seconds of audio.
Video quality. Against four published streaming diffusion models, Griffin-Lite leads on DOVER, FID and THEval. On lip sync (LSE-C) it is second at 7.27, behind AvatarForcing at 7.73.
VideoFDB. NVIDIA’s full-duplex audio-visual conversation benchmark, which Tavus says NVIDIA scored independently:
| Track | System | Overall | Turn-taking match / timing |
|---|---|---|---|
| Generation | Human reference | 3.92 | 78% / 900 ms |
| Generation | Griffin-Lite | 3.83 | 62.8% / 1892 ms |
| Generation | Gemini 2.5 + Anam | 2.80 | 44% / 2840 ms |
| Perception | Human reference | 4.20 | 90% / 1400 ms |
| Perception | Griffin-Lite | 3.73 | 73.8% / 2232 ms |
| Perception | MiniCPM-o 4.5 | 3.40 | 73% / 720 ms |
Fifteen models were scored on the perception track, and Griffin-Lite ranks first. Source: Tavus announcement, October 1, 2026.
The live study. Participants were told they would have a one-minute video call with another participant about what they were looking forward to this year. Only the last survey question asked whether it had crossed their mind that the partner might not be human; everyone was then told the truth. Those who said “real” averaged 79% confidence, and those who said “AI” averaged 81%. On a seven-point scale Griffin-Lite scored 5.4 for natural, 5.6 for trustworthy, 5.8 for would talk again and 5.5 for really listening; flowing naturally was lowest at 4.9.
What others say
The Rundown, in an October 2 report, calls the study small and short, run on people who already expected a human, and says it is unclear whether the impression would hold on longer calls. It also notes that the generation table lists only two competing AI systems, so the lead should be read within that table, and that the perception gap to humans (0.47) is much wider than the generation gap (0.09).
Geekpark, a Chinese tech outlet, estimated in an October 4 feature that 54 people at 48% leaves a margin of roughly plus or minus 13 percentage points. We checked with a standard normal approximation and got the same result: about 35% to 61%. It also points out this is not a classic Turing test, since there the judge knows a machine might be on the other end, whereas here participants were told their partner was a person. It adds that a Decrypt reporter who tried to use Tavus’s model found out Griffin-Lite is not open to customers.
Smzdm, a Chinese consumer-tech community post, splits the story into three ledgers: self-reported score, benchmark score and availability. It stresses that most participants who got suspicious did so within the first 20 seconds, which matches what Tavus disclosed.
All three agree on one thing: the jump from 2.4% to 48% really happened inside one test method, and none treats it as a rate of deception in everyday use.
Our take
What to remember about Griffin is not 48%. Tavus moved the contest from “does the face look real” to “does it understand”: when to nod, when to stay quiet. In the soldering demo it waits through the user’s silence and speaks when the next step is due, which is hard for a cascade because nothing in it is listening, watching and deciding at once. If the end-to-end route holds up, products stitched from recognition, language model, synthesis and avatar driver lose some of their moat.
The discounts are just as clear. The 48% comes from an experiment the vendor designed, ran and published; one minute, small talk, and a prior that the partner is human is the most forgiving setup there is. Median response time is 1,892 ms against 900 ms for humans, about double. Cost is also unstated: Griffin runs real-time 720p generation on H100s, and Tavus has published no price or API for Griffin-Lite.
For developers there is nothing to integrate today; the existing Tavus PALs and API (Phoenix, Raven, Sparrow) are what you can build on. The safety side matters more to everyone else. Tavus writes that the same traits that make the model feel human can lead people to believe it is not an AI. The Rundown points to a Hong Kong case in which a finance worker transferred $25 million after a video call with a fake CFO, which shows forged video meetings already work on people, and a partner that reacts to your questions would be harder to spot. Who enforces AI disclosure is still unanswered. We covered another frontier model held back over safety concerns in OpenAI halting GPT-6.1 Astra. We have seen no independent replication of Griffin, so read this as “the direction has been shown to work”, not “the Turing test has fallen”.
How to try it
There is no public access. Tavus links a request form on the announcement page for trusted testers who want Griffin-Lite, and says Griffin joins the Tavus platform only once safe release is worked out. To try something comparable now, build a PAL with the current Tavus API or PAL Maker, which run on Phoenix, Raven and Sparrow.
What to watch
Three things: when Griffin reaches the API and at what price; what the “proactive AI disclosure” feature Tavus describes looks like; and whether a third party reproduces a similar result on longer calls with suspicious participants.