Tavus Griffin: The First Human Interaction Model Passes the Video Turing Test

Tavus announced Griffin on October 1, 2026: 48% of participants mistook it for a human on live video calls. How it works, how solid the numbers are, what it means for creators — and when you can try it.

Tavus Griffin: The First Human Interaction Model Passes the Video Turing Test
Table of contents

On October 1, 2026, San Francisco startup Tavus announced Griffin, the first of a category it calls the Human Interaction Model (HIM). The central claim is striking by any measure: in a live study, 48% of participants who held a video call with Griffin believed they had been talking to a real human — where the previous best performance on the same test topped out at 2%. And to make clear this is not marketing inflation, we walk through the details exactly as published in the company's official research post — signed by its CEO — alongside the honest caveats it deserves.

The story matters to working creators for two reasons. First, it sketches the next stage beyond today's video tools: from an avatar that speaks after you write a script, to an avatar that answers you live on a call. Second, Tavus itself is holding back the more powerful release "until safety concerns are addressed" — an unmistakable signal about the scale of open ethical questions around indistinguishable synthetic video.

Official Tavus Griffin announcement image
Official image from the Griffin announcement page on Tavus's site (source: Tavus)

What is Griffin, and what is a "Human Interaction Model"?

According to the official announcement post — authored by Hassaan Raza, co-founder and CEO, and Ioannis Patras, head of research — Griffin is a full-duplex, video-to-video model operating in real time. It receives your image and voice as you speak, understands expressions and pauses rather than words alone, and generates its response — voice, face, and movement — in the same instant. In other words, it is not a voice assistant with a painted-on face; it is a complete system that generates every pixel of every frame from a single reference image and behaves like a genuine conversation partner: it laughs, shifts tone, interrupts, and gets interrupted without losing its place.

This new category — Human Interaction Models — is Tavus's attempt to name something distinct from today's chat assistants: a model that does not make you meet it halfway, but comes to you, learning the language humans use face to face: gaze, expression, gesture, and timing.

The "video Turing test" results, exactly as published

The numbers in the official post split into two parts:

  • The live video Turing test: 48% of participants who conversed with Griffin thought they had talked with a real human. By comparison, previous systems — including Tavus's own production stack built on Phoenix 4.5, Sparrow-2, and Raven-1 — maxed out at a 2% pass rate on the same test.
  • NVIDIA's independent test: on an evaluation of how "human" an AI feels in direct audio-video conversation, Griffin scored 3.83 points versus 3.92 for actual humans and 2.80 for the previous best AI. The company also claims Griffin is 37% ahead of the next best AI "at reacting in the moment."

Translating those numbers into practice: roughly half of the audience could not tell in a short call, and the gap with real humans narrowed to 0.09 points on an independent scale. These are the company's figures, of course — but the NVIDIA test cited is external, and the general direction, that live synthetic video has crossed the "genuinely suspicious" threshold, is now documented this clearly for the first time.

How Griffin works, in plain terms

The official post explains the architecture in respectable depth. It boils down to three layers running concurrently:

  • Continuous Conversational Modeling engine: instead of waiting for you to finish speaking and then replying (as cascaded systems do: speech recognition, then a language model, then text-to-speech, then an avatar), Griffin reassesses the conversation state at sub-second intervals. At each "mini-turn" it decides: stay quiet, nod, backchannel an encouraging "mm-hm," or interrupt — based on what it hears and sees, not merely on when audio stops.
  • Audio-Visual Generation engine: streaming speech generation and streaming video generation convert those decisions into voice and video simultaneously, so a decision made mid-sentence — a pause, a glance away, a smile, an interruption — shows up in the voice and the face within the same mini-turn, not after a hand-off between separate systems.
  • The Tavec audio codec: a Tavus-built convolutional autoencoder that maps 48 kHz audio into a compact continuous representation — 40 values per 10-millisecond frame, no codebooks — enabling speech generation faster than real time, streamed in packets as small as 10 milliseconds. One consequence: the model can clone a speaker's voice from roughly 10 seconds of audio.

On the video side, Tavus dropped the heavy 3D priors that constrained its earlier models — wide gestures and dynamic backgrounds were previously out of reach — in favor of a diffusion-based generator that renders the whole scene: face, hands, the chair the person sits in, the shadows they cast, and the background, all from one reference image.

Official Griffin announcement page on the Tavus website
The official Griffin announcement page on Tavus's site, October 1, 2026 (screenshot of the official page)

The five capabilities, as the company presents them

The post documents five capabilities with demo clips from real conversations between Tavus researchers and first-time users:

  • Behavior and emotion modeling: it laughs, changes its tone, and responds to what the other person says and does according to conversational context.
  • Full-scene generation: every pixel of every frame is generated in real time — including the movement of the chair the person sits in, the shadows they cast, and the background behind them.
  • Full-duplex conversation: it can interrupt, adjust, backchannel, or be interrupted without losing its place in the discussion.
  • Perception: it understands what appears in your camera — your environment, an object you show it, a screen you share — and reacts to it, not to audio alone.
  • Temporal understanding: it knows how long a silence has lasted, what a sustained silence means, and when to speak again on its own.

Can you try it today? The full truth

Here the announcement and the availability must be separated. What actually exists is Griffin-Lite, a research preview open only to a select group of early testers — not a product you can subscribe to today. The more capable, wider release is explicitly deferred "until safety concerns are addressed." The realistic reading: the company has proven the technical feasibility and is now moving carefully before putting the capability in everyone's hands.

That caution is not a formality. A model passing the video Turing test at 48% is by definition a first-class potential fraud instrument: fake video calls from "real people" requesting money transfers or credentials. Personalized marketing video — Tavus's core business since its 2020 founding and roughly $64 million in raised funding — is a different universe from a digital twin sitting across from you in a call that is hard to distinguish from your child or your manager.

What this means for content creators

In the near term: nothing directly — the product is not publicly available. But within months, the door Griffin opens reorders three ideas in a creator's mind:

  • Avatars move from recording to broadcasting: today's tools generate avatar video from a written script — like the Gemini Live avatar whose launch we covered previously. The next generation makes an avatar available in a live stream that interacts with the audience — promotional broadcasts, live Q&A sessions, customer support with a face that speaks every customer's language.
  • Education and training are the nearest use cases: the company itself leads with tutoring, "practicing difficult conversations," and camera-based tech support — applications international audiences grasp instantly, from language learning to job-interview rehearsal.
  • Whoever owns your face owns your content: if cloning a voice takes 10 seconds and animating a face takes one photo, the value of your digital assets — your voice and likeness — rises, and so does the need to protect and license them deliberately.

Quick comparison: Griffin versus your current tools

The fair comparison is not Griffin versus today's video generators — those produce recorded clips while this runs live conversations. The emerging division of labor has three camps: text-to-video models (for recorded content), voice tools (for narration and podcasts), and live avatars (for real-time interaction). A creator building a channel around a digital persona will soon face a genuine question: produce recorded content at lower cost, or open live sessions with my digital twin? Griffin makes the second option thinkable for the first time. For those working in Arabic and English voice and text today, mature tools already exist — from Gemini's voice studio to the always-on assistants we covered in our dots reporting — while Griffin remains a project to watch, not to buy.

Where Tavus came from: the company context behind Griffin

To grasp the size of the leap, a little context helps. Tavus was founded in San Francisco in 2020, raised roughly $64 million, and started with a clearly defined business: personalized AI marketing video — sales messages generated for each customer by name and company, a category with proven conversion. It then expanded step by step into "live digital personas" through its previous model stack: Phoenix for face animation, Sparrow for conversation management, and Raven for perception — the very stack the company says never exceeded 2% on the video Turing test.

In other words, Griffin is not a unknown startup chasing headlines; it is the next generation from a company that has known this market for six years and chose to declare a new category — Human Interaction Models — because it judged that its result no longer fits any existing classification. That context matters when weighing the claims: an established company staking its research reputation on a checkable number, and submitting its architecture to an external test like NVIDIA's, is a materially different signal than a vague demo from an unknown label.

What to watch in the coming months: four decisive markers

  • Opening the preview to the public: moving Griffin-Lite from selected testers to an open waitlist is the first seriousness signal; only then will genuine independent testing appear at scale and across languages.
  • The announced safety framework: the guardrails accompanying the stronger release — provenance watermarking, identity verification, cloning limits — will determine whether the industry proceeds responsibly or replays the deepfake story from the top.
  • Responses from Google and OpenAI: Google already ships a live avatar in Gemini Live and OpenAI has advanced voice systems; if either announces a full live video conversation upgrade, the shift from "stunning demo" to "market category" has genuinely begun.
  • Real-world operating cost: generating every frame in real time is computationally extreme; public-release pricing will reveal the true audience — support and education companies first, most likely, before individuals.

Honest limitations that need to be said out loud

  • Closed research preview: you cannot try Griffin-Lite today unless you are among the chosen testers; there is no public price and no announced general-release date.
  • Company numbers first: the 48% figure comes from the company's own study; the NVIDIA test is independent but singular, and the full methodology awaits research-community review.
  • The safety question is deferred, not solved: Tavus itself gates the stronger release on safety work — meaning the model's most dangerous capabilities exist technically and await unfinished guardrails.
  • No announced Arabic support: the entire post covers English conversations; Griffin's Arabic performance is entirely undocumented so far.
  • Unknown operating cost: generating every pixel of every frame in real time is extreme compute; future pricing will decide who actually uses it.

Frequently asked questions

What is Tavus Griffin?

The first "Human Interaction Model" from US startup Tavus, announced October 1, 2026: a real-time, full-duplex, video-to-video system that holds live video conversations, understands expressions and pauses, and generates voice and video together moment by moment.

Did Griffin really pass the Turing test?

In the company's live study, 48% of participants believed they had spoken with a real human — which Tavus describes as the first pass of the live video Turing test, against a 2% previous best. The result is promising and company-documented; broader independent review has not yet been published.

Can I use Griffin now?

No. The available version, Griffin-Lite, is a research preview limited to selected early testers; the more powerful release is deferred until safety concerns are addressed, with no public pricing or date announced.

How does Griffin relate to AI video tools like Sora or Veo?

Those models generate recorded video clips from text prompts, while Griffin holds a live, interactive video conversation in real time — two different categories: the first for recorded production content, the second for live interaction.

Does Griffin support the Arabic language?

There is no announcement or documentation of Griffin's Arabic performance; all published demos and tests were in English. The general release will be the moment of truth.

What are the risks of this class of model?

The most serious is identity fraud via realistic video calls that are hard to distinguish — financial scams, impersonation of public figures, and interactive deepfakes. The company itself delayed the stronger release for safety reasons, and incoming regulation will inevitably target this category.

The bottom line: Griffin is not a product you buy today — it is a signal about where the market goes next. The gap between "convincing digital video" and "indistinguishable digital call" suddenly collapsed from enormous to 0.09 points on an independent scale. For creators, the practical message is to watch this timeline closely: whoever understands early how to deploy live avatars in education, support, and entertainment will own open ground before it crowds. And until that generation arrives, stay ahead with today's tools — ARWriter writes and schedules your content in Arabic and English from one place.