Tavus Says Its Griffin AI Fooled 48% of People on Live Video Calls Into Thinking It Was Human
Key Takeaways
- •Tavus says its Griffin-Lite model persuaded 26 of 54 participants — 48% — in one-minute video calls to believe they were speaking with a real person, while its earlier system fooled just 2.4% of 41 participants.
- •The findings appear on Tavus' own research page and have been flagged, via a community note on X, as neither independently verified nor conducted under a standard protocol.
- •On NVIDIA's independently run VideoFDB benchmark, Griffin-Lite ranked first with a generation score of 3.83 out of 5, though its perception score of 3.73 remained below the human reference of 4.20.
- •Griffin operates in full duplex, capable of listening, watching and speaking simultaneously, with an average audio-to-video delay of 0.43 seconds on NVIDIA H100 chips that Tavus says is half that of the next-fastest method.
- •Griffin-Lite is currently limited to select trusted testers as a research preview, and Tavus says safety measures and disclosure features are required before a public release.

AI startup Tavus says its new model, Griffin, persuaded 48% of the people who spoke with it on a live video call that they were talking to a real human.
The company unveiled Griffin on October 1, billing it as the first "Human Interaction Model" — an AI built to understand and generate face-to-face conversation, paying attention to facial expressions and pauses as well as words. On the same test, Tavus says, its previous system scored just 2.4%.
In the experiment, Griffin-Lite — the version Tavus tested — faced 54 people, 26 of whom said afterward they believed their conversation partner was a real person. The older system faced 41 people and fooled exactly one.
Participants were told they would be matched with another person for a one-minute video call about what they were looking forward to this year, according to Tavus. Only at the end were they asked whether it had crossed their mind that their partner might not be real.
That setup differs from the classic Turing test proposed in 1950 by British mathematician Alan Turing, in which a judge talks to both a hidden human and a hidden machine and must work out which is which. In Tavus' test, nobody on the call was told a bot might be on the other end.
The results are published on Tavus' own research page, and participants were recruited through what the company describes as an independent research platform. A community note on X has already flagged that the findings are not independently verified and do not follow a standard protocol — meaning the 48% headline rests, for now, on the company's own numbers. Tavus says the participants who grew suspicious typically did so within 20 seconds.
AI systems have been edging toward this milestone for some time. According to a UC San Diego study, OpenAI's GPT-4.5 convinced judges it was human in 73% of conversations when prompted to play an introverted, internet-savvy young person — but that test was text only. Griffin adds a face and a voice, in real time.
On NVIDIA's VideoFDB benchmark, a test of live audio and video conversation, Tavus says Griffin-Lite ranks first. Its generation track, which grades how natural and expressive a model's responses are, gave Griffin-Lite a score of 3.83 out of 5, against 2.80 for the next-best system and 3.92 for a human reference.
The perception track, which measures whether a model understands what it sees and hears, shows where Griffin still trails people. Griffin-Lite scored 3.73, ahead of 3.44 for the strongest baseline but below the human reference of 4.20. Tavus says NVIDIA ran the evaluation independently.
Griffin is also full-duplex, meaning it listens, watches and talks at the same time — more like a phone call than a walkie-talkie. In a Tavus demo, the model coaches a man through solving a Rubik's cube based on what it sees in his hands, and waits when he goes quiet to think. Audio-to-video delay averages 0.43 seconds on NVIDIA H100 chips, the kind used in AI data centers, which Tavus says is half that of the next-fastest method.
The advance matters well beyond the lab, because scammers already operate on video calls. In January, North Korea-linked hackers used deepfakes — AI-generated video that imitates a real person — on Zoom or Teams calls to pose as trusted contacts. Security researchers attributed the intrusion to BlueNoroff, a Lazarus Group subsidiary, and victims were talked into malware disguised as an audio fix. David Liberman, co-creator of Gonka, a decentralized network for AI computing, said in that report that photos and video can no longer be trusted as proof that something is real — and the models used then were not as advanced as Griffin.
Companies are already improvising their defenses. In 2025, Kraken flagged a suspected North Korean job applicant after its security team asked spontaneous questions, such as requesting government ID and the names of local restaurants. The candidate struggled.
Decrypt's own attempts to test Tavus' models proved disappointing. After some research, it turns out Griffin-Lite is not available to customers; it is limited to select trusted testers as a research preview, and the company says it is working on disclosure features and with AI safety organizations. Tavus says Griffin needs safety measures before a public release — making the shape of those safeguards, and when broader access arrives, the next developments to watch.
Tavus raised a $40 million Series B in November 2025, led by CRV. The earlier system that scored 2.4% stitched together three separate models — one each for visuals, dialogue and perception. Trusted testers can request access to Griffin-Lite by submitting a form on the Tavus website.
Originally reported byDecrypt](https://decrypt.co/379975/tavus-griffin-ai-fools-people-video-thinking-human).