Talking Avatar
Open source byInterviewFlowAI

Give your voice agent a face.

Real-time lip sync for 3D avatars from any audio. Hand it an avatar and your agent's voice; a 1 MB model in the browser moves the lips in sync. Free and open source, from InterviewFlowAI.

npm install @interviewflowai/talking-avatar three
Live in your browser · OpenAI TTS · coral

Works with any voice, from any provider

  • OpenAI
  • ElevenLabs
  • Azure
  • Fish Audio
  • Cartesia
  • LiveKit
  • Daily
  • Pipecat
  • WebRTC

Features

Lifelike lip sync without the heavy lifting

Everything a talking avatar needs, in a package small enough for any web app.

How it works

From sound waves to mouth shapes in 10 ms steps

  1. 01

    Listen

    Audio from any source is resampled to 16 kHz and turned into 10 ms spectral frames as it plays.

  2. 02

    Recognise

    A small causal neural network names the mouth shape (viseme) for every frame, looking 20 ms ahead.

  3. 03

    Animate

    Shapes are blended the way real speech flows from one sound into the next, driving the avatar's face in sync with the audio.

Quick start

Two lines from voice to face

Install the package, point it at an avatar and give it audio. It handles playback, the model, lip sync and rendering. Framework-free, with a React component and TypeScript types.

npm install @interviewflowai/talking-avatar three
Read the documentation
import { TalkingAvatar } from "@interviewflowai/talking-avatar";

const avatar = new TalkingAvatar({
  container: document.getElementById("avatar"),
  avatarUrl: "/avatar.glb", // any .glb with Oculus visemes
});

button.onclick = () => avatar.speak("/hello.mp3"); // or avatar.attachStream(stream)

Use cases

Made for conversational AI

  • AI voice agents

    Put a face on agents built with OpenAI Realtime, LiveKit Agents, Pipecat or your own WebRTC stack.

  • AI interviewers

    Built by InterviewFlowAI for its AI interviewers, which screen candidates over phone and video.

  • Tutors and support

    Friendlier tutoring, onboarding and customer-support agents that look at the user and talk.

  • Narration and presenters

    Turn TTS narration, product tours and announcements into a talking presenter.

Measured, not guessed

Accurate, light and fast

Accuracy measured on 39 speakers the model never trained on, with audio passed through the WebRTC Opus codec.

of 10 ms frames get the right mouth shape on unseen speakers
82%
of p, b and m sounds close the lips
99%
of one CPU core while speaking
~2%
of gzipped code, plus a 1 MB model cached once
8 KB

FAQ

Questions, answered

What is Talking Avatar?

An open-source JavaScript SDK by InterviewFlowAI that lip syncs a 3D avatar to any audio in real time. You give it an avatar URL and an audio stream or clip; a small in-browser model drives the mouth.

Does it work with OpenAI, ElevenLabs or LiveKit voice agents?

Yes. It only listens to the audio, so any TTS provider or WebRTC voice agent works: pass the agent's audio stream to attachStream().

Do I need a server or an API key?

No. Everything runs in the browser and the audio never leaves the page. The only downloads are the avatar and the 1 MB model.

Which voices give the best lip sync?

TTS voices. The model was trained on clean, clearly spoken read speech, which is what TTS produces. Real voices from a microphone work too, but noise and echo make the lip sync less accurate.

Can I use my own avatar?

Yes: any .glb with Oculus viseme morph targets, or map your avatar's own morph names with the morphTargets option.

Which languages does it support?

It was trained on English. On other languages its timing holds up well, with sounds English lacks drawn using the nearest English mouth shape. Try Spanish, Hindi, Japanese and more in the live demo.

Is it open source?

Yes. The SDK and the model are on GitHub, built and maintained by InterviewFlowAI.

Built and open-sourced by InterviewFlowAI

The lip sync built for InterviewFlowAI's AI interviewers, now open source.