October 3, 2026 ยท 4 min read

Tavus Griffin: real-time AI avatars that listen while they talk

Tavus Griffin leaves me with two reactions at once: this is impressive, and I'm not quite sure how I feel about where it goes. A video call with an AI that can listen, react and keep up with you could be very useful. It also makes the question of who's actually on the other end feel a lot less straightforward.

The little things are what interest me here. The timing of a smile, making room for an interruption, reacting while someone is still speaking. A fairly ordinary answer can feel different when all of that starts working together.

Here's Tavus's launch video. Pay attention to what happens around the words, including while the other person is talking.

Video: Tavus. This is their demonstration; I haven't tested Griffin myself. You can also watch it on YouTube.

It keeps listening while it talks

Tavus introduced Griffin on October 1. Griffin-Lite is the research preview. The company describes continuous audio and video input, with conversational decisions made at sub-second intervals while speech and video are generated.

In plain English: the system can keep paying attention while it answers. A pause, interruption or facial reaction can influence what it does next. Tavus's architecture combines a conversational engine with streaming speech and video generation.

Think about explaining a problem to someone. They might need a moment to answer, but a nod or a glance at what you're showing them tells you they're still following. The same delay with a frozen face feels like the connection has dropped.

That's why what happens during the wait matters. A transcript misses a lot of what makes a conversation comfortable, or weird.

How close is this to passing for a person?

Tavus reports that 26 of 54 people, or 48%, thought their partner was human after a one-minute call. Participants expected another participant; they were asked about possible AI identity only at the end of the survey.

That's a striking result, with a fairly specific meaning. A small, vendor-run study of brief conversations doesn't establish that the system is undetectable. It doesn't tell us what happens in a longer call, or when someone is actively trying to work out whether they're speaking to AI.

NVIDIA and David AI's VideoFDB leaderboard provides another check. On October 3, Griffin-Lite led the evaluated AI systems on both audiovisual generation and perception. The human reference remained ahead on both overall scores and timing alignment. These are rubric scores from a model-based judge, so they have their own limits too.

At launch, Griffin-Lite was restricted to selected trusted testers. There's still a gap between an impressive demonstration and something people can judge through everyday use.

Tech support is an easy application to imagine

An IT support agent that can follow what you're showing it, ask you to move the camera and notice that you're pointing at a different cable could be useful. So could one that stops explaining the wrong screen as soon as you interrupt.

That would save some of the awkward back-and-forth of describing a problem when you don't know what the part is called. The same kind of interaction could help with field training, or with practising a difficult conversation.

These are possibilities, rather than outcomes established by the demo. The advice still has to be correct. A friendly face confidently suggesting the wrong fix would just make bad support more convincing.

And plenty of tasks only need a short answer or a button. I don't particularly need a photorealistic person to read out a password-reset link. The interesting cases are the ones where seeing, listening and reacting actually make the exchange easier.

Then there's the world around it

If conversations like this become common, a natural video exchange will tell us less about whether another person is present. Being able to talk comfortably with software is useful; being unsure who you're talking to is a different experience.

For routine support, a clearly labelled AI might be perfectly fine. In a personal conversation, or a meeting where someone's presence matters, knowing who has actually turned up matters too. I'd want that information before the call, rather than having to play detective during it.

And of course there's the slightly ridiculous version of this future: a company's AI interviews a candidate's AI. They build excellent rapport. The candidate still gets rejected for lacking a human touch.

Two systems discussing the importance of authentic human connection, while both people are somewhere else. It sounds like a joke. I'd also be completely unsurprised if someone pitched it as a product.

That's speculation, but it gets at something I haven't quite settled in my own head. Some conversations would be easier if software could handle them. Others would feel fairly pointless if everyone sent a substitute.

I can see useful applications for this immediately, and I'm curious about what people will do with it. I'm less sure how I'll feel when a convincing person on a video call becomes something ordinary software can produce. For now, both reactions seem reasonable.


Sources: Tavus's Griffin research page, the launch video above, and NVIDIA / David AI's VideoFDB. Research snapshot: October 3, 2026. Tavus's study is vendor-reported. The possible applications and future scenarios are my own interpretation.