Tavus, on 1 October 2026, introduced Griffin. Hassaan Raza, co-founder and CEO, and Ioannis Patras, head of research, call it the first Human Interaction Model, their name for one system that holds a face-to-face call, rather than a chain of models that wait on each other.

Most real-time AI, they write, works as a relay, often called a cascade. One system turns speech into text. A language model replies to the transcript. Other systems turn that reply into a voice and a face. Each handoff waits, and each one drops what the next system never sees, such as tone, or what is on camera. Griffin is built to keep one loop running. It listens while it talks. Watching, choosing whether to speak or wait, and producing the speech and the picture all happen at the same time.

The preview on the page is Griffin-Lite. The face on the call is a PAL, their name for the person the model plays. In the films, that face is a woman they call Vanessa, or a man in a denim shirt, against a plain wall. The offices, the red chairs, and the workshop belong to the people on the other side of the laptop.

A laptop showing Vanessa, a woman in a pink shirt, while a man holds a Rubik's cube up to the camera.
Vanessa, in the cube film on Tavus's page. A still from that film.

The page's caption for Simon Says says the face copies the gesture only when the person says "Simon says," and calls the bluff when they do not. The audio in the clip matches that. In the cube film, a man holds the cube up and says he has the first two layers and only a bar on the blue face. Vanessa starts a beginner sequence, a few turns at a time, and waits until he says he is ready. The page says she is watching the cube as he turns it. In the soldering film, a man turns the iron on and asks to be told at twelve seconds. The face says it will watch, and later says the time is up. That film marks the wait with a fast-forward icon and a clock at 00:10.

They say the video side paints every pixel from one reference photo, including the chair and the shadows, in real time. The generated person in these clips is a head and shoulders on a white wall. They also say a diagram of the timing is illustrative, and was not measured from a real session. On the generator, they report 720p video in 320-millisecond chunks. On Nvidia H100 chips, a data-center GPU, the time from a piece of audio arriving to the face showing it averaged 0.43 seconds. They say that is half the next fastest published streaming model, because Griffin does not wait for audio that has not arrived yet.

NVIDIA's VideoFDB is the scored test they point to for the live call. The leaderboard lists Tavus Griffin Lite. A language-model judge rates each response from 0 to 5. On generation, whether the model's own speech and movement fit the moment, Griffin Lite scores 3.83. The human recordings score 3.92. The next system on that board, Gemini 2.5 plus the Anam avatar, scores 2.80. On perception, Griffin Lite's overall score is 3.73. The human recordings score 4.20. The visual-grounding part of that score, whether the response uses what the model sees, is 3.92 for Griffin Lite. The highest other perception overall on the board is 3.44, from MiniCPM-o 4.5 run on audio only. The row directly under Griffin Lite is Gemini 2.5 Flash Native, at 3.17. The board prints two timing figures on each row: a takeover-rate alignment, and a median latency under it. Tavus describes the alignment as how closely the model's decisions about when to speak match the timing in the reference conversations. On generation, Griffin Lite is at 62.8 percent and 1,892 milliseconds. The human recordings are at 78 percent and 900 milliseconds. That latency is not the 0.43 seconds above, which is only the time from audio arriving to the face showing it.

They also ran a live study. They recruited 54 people through a research platform and told them they would be matched with another participant for a one-minute video call, about what they were looking forward to this year. The partner was a PAL on Griffin-Lite. After the call, people wrote down the partner's answer and rated the conversation. Only at the end were they asked whether it had crossed their mind that the partner might not be a real person. Then they were told it was a model. Twenty-six of the 54 said the partner was real. On the older relay, Phoenix-4.5 for the face, Sparrow-2 for when to talk, and Raven-1 for what it sees, 1 of 41 said real, which is 2.4 percent. Tavus calls the 48 percent a pass of a real-time video Turing test, and writes that Griffin is, to the best of their knowledge, the first model to have ever passed one. People who said real averaged 79 percent confidence. People who said it was a model averaged 81 percent. More than half said the thought had not crossed their mind during the call. Those who suspected tended to suspect within the first 20 seconds.

On a 7-point scale, the averages were 5.4 for seeming natural, 5.6 for seeming trustworthy, and 5.8 for wanting to talk again. That last score was still 5.4 among people who said it was a model. The conversation scored 5.5 for feeling listened to, and 4.9 for flowing naturally, the lowest of the five.

They are not shipping it. The same properties that make the call feel natural, they write, can deceive a person into believing the face is not a model. Griffin-Lite is a research preview for select trusted testers. It is not available to customers. They say they are working on a way to disclose that the face is a model, and on further safety checks, before a wider release. A more powerful Griffin is meant to follow.

They name three uses, and they are holding the model while they work those checks. A tutor who notices an explanation is not landing, and tries another. Someone at work, rehearsing a hard conversation with a counterpart who reacts. A customer holding a broken part up to the camera and working out the fix without knowing the name of the part. In each case, they write, the person does not have to turn the need into the right command first. Until that release, the platform still runs the older split. Phoenix generates the face, Raven sees you, and Sparrow decides when to talk.