Tavus Says It Built the First AI to Pass the Video Turing Test. It's Too Convincing to Release.
Nearly half the people who video-called Tavus's new Griffin model thought it was human, which is why the company says it can't release it yet.
If this caught your attention, that’s not accidental.
The best editorial systems don’t happen by accident. Outlever builds them.

For one minute, 54 people sat on a video call with someone they believed was another study participant. They chatted about what they were looking forward to this year. Afterward, a survey asked whether it had ever crossed their mind that their partner wasn't a real person.
Twenty-six of them believed they'd been talking to a person. They'd been talking to Griffin, a new model from San Francisco-based Tavus, which the company announced on Thursday as the first AI to pass a real-time video Turing test.
Tavus isn't selling it. In its launch thread on X, the company said Griffin needs more safety work before a public release because it's the first model people can mistake for a real person. For now, a small group of testers has access to a research preview called Griffin-Lite, and customers will wait until Tavus finishes building safeguards.
What Tavus Built
Most AI video agents today work like a relay. One system transcribes what you say, a language model writes a reply, and other systems turn that reply into a voice and a face. Each handoff adds lag and drops information, which is why the avatar keeps smiling through bad news and leaves a beat of dead air after you stop talking.
Tavus calls Griffin a Human Interaction Model. A single system takes in your video and audio and generates its own, listening and talking at the same time. Several times a second, it reassesses the conversation and decides whether to stay quiet, nod, say "mm-hm," cut in or yield. It can be interrupted mid-sentence without losing its place, and it reacts to what it sees on camera as well as what it hears. In one of Tavus's demo clips, it coaches someone through a Rubik's Cube by watching his hands, then waits when he goes quiet to think.
The rendering is the other half. Griffin generates the entire 720p frame in real time from a single reference photo, including hands, shadows and background, and Tavus says it can clone a voice from about 10 seconds of audio.
A Twenty-Fold Jump in One Generation
Tavus runs the same live study on every model it builds. Its previous production stack, the Phoenix-4.5 rendering model paired with its Sparrow-2 turn-taking and Raven-1 perception models, convinced 1 of 41 participants, or 2.4%. Tavus launched Phoenix-4.5 as its state of the art just nine days ago. Griffin-Lite convinced 26 of 54.
The people who guessed human and the people who guessed AI were about equally sure of themselves, at roughly 80% confidence on both sides. More than half said the possibility never occurred to them during the call. Those who did get suspicious usually did so within the first 20 seconds.
Read the Fine Print
Alan Turing's imitation game puts a judge in conversation with a person and a machine, tells the judge one of them is a machine, and asks them to pick. Tavus's participants weren't judges. They were told they'd be matched with another participant and were only asked about AI once the call was over. That measures whether a model can go unnoticed by people who aren't looking for it, which is a lower bar than beating someone actively hunting for the machine.
Tavus CEO Hassaan Raza has described a different goal. In a LinkedIn post announcing Griffin, he wrote that he wants people to talk to it naturally even when they know it's AI, with the technology fading into the background. The study only covered people who didn't know. Online, nobody kept the two apart. Within hours, viral accounts on X had dropped the word "video" and were announcing that AI had passed the Turing test for the first time.
The sample is small and the calls lasted one minute. Participants came through an outside research platform, but Tavus designed, ran and published the study itself, and it hasn't been peer-reviewed. Some of the first enthusiastic posts came from creators Tavus gave early access to, and at least one was labeled a paid partnership.
The jump from 2.4% to 48% under the same protocol is still a big result. Stated carefully, though, the claim is that Tavus's own study found its model often goes unnoticed in a one-minute call. That's the version enterprise buyers should repeat.
The Score Tavus Didn't Grade
The stronger evidence comes from NVIDIA. Its Video Full-Duplex Benchmark, which NVIDIA scores independently with its own judge, rates how naturally a model produces conversational behavior and how well it reads the person on the other side.
On generation, Griffin-Lite scored 3.83 out of 5. Recordings of real humans scored 3.92. The next-best system, Google's Gemini 2.5 paired with Anam's avatars, scored 2.80. On perception, Griffin-Lite scored 3.73, ahead of every realtime model from Google and OpenAI on the leaderboard but well short of the human reference at 4.20.
That gap matters more than the headline figure. Griffin comes close to a human at carrying its side of a conversation, and it's noticeably weaker at reading the other person's.
[QUOTE OPPORTUNITY: Tavus CEO Hassaan Raza or Head of Research Ioannis Patras on when Griffin will ship and what disclosure will look like in practice. If neither responds before publication, replace with: "Tavus did not respond to a request for comment."]
Who Feels This First
Tavus already has a sizable business. It says 150,000 developers and businesses build on its platform, and it sells video agents for sales, recruiting, healthcare intake and training. It raised a $40 million Series B led by CRV last November, with Sequoia and Scale Venture Partners in the round. When Griffin ships, it will go straight into those existing deployments.
For companies running video agents, disclosure becomes a design problem. If half of customers who aren't expecting an AI can't spot one, a line in the terms of service won't do. Tavus says it's building disclosure features before release, a sign the company knows its default experience now fools people.
Security teams have the more urgent problem, and it doesn't depend on Tavus at all. In 2024, an employee in engineering firm Arup's Hong Kong office made 15 transfers totaling about $25 million after a video call where every other participant, including the CFO, was a deepfake. Last week, Reuters reported that fraudsters took €95 million from Fideuram, Intesa Sanpaolo's private bank, using a spoofed WhatsApp message from CEO Carlo Messina and an AI-cloned voice of a law firm partner. About €36 million is still missing.
Neither attack needed a fake that could hold a real conversation. In the Arup case, the deepfaked executives reportedly avoided extended back-and-forth and mostly issued instructions. A model that renders a full scene from one photo, clones a voice from 10 seconds of audio and handles interruptions and unscripted questions removes that limit. Tavus can hold its model back, but it has also published a detailed account of how it built it, and others will follow with fewer reservations about release.
Our View
Video presence should stop counting as proof of identity, starting now. Payment approvals, credential resets, vendor bank changes and executive requests made on camera need an out-of-band check. Remote hiring and KYC flows that treat a live face as evidence of a real person need a second signal. Security teams have given this advice for years. The usual pushback, that people can tell when something's off on video, now has a counterexample with numbers attached.
Text crossed this line last year, when UC San Diego researchers found that GPT-4.5 was picked as the human 73% of the time in a three-party Turing test, more often than the actual humans it was paired against. Video was supposed to hold out longer, because faces and timing carry so much of how people size each other up. By Tavus's numbers, it held out about 18 months.
The line in Tavus's announcement that deserves the most attention is the one where it says it can't release its best model until it works out how to keep people from being fooled by it. AI vendors rarely hold back a flagship because it works too well. Enterprises should take that warning more seriously than the headline.
If this caught your attention, that’s not accidental.
The best editorial systems don’t happen by accident. Outlever builds them.


Get the latest AI insights first.
Sign up for updates, interviews, and fresh analysis on how AI is reshaping business, brands, and technology.





