Give your voice agent a face.
Real-time lip sync for 3D avatars from any audio. Hand it an avatar and your agent's voice; a 1 MB model in the browser moves the lips in sync. Free and open source, from InterviewFlowAI.
npm install @interviewflowai/talking-avatar threeWorks with any voice, from any provider
- OpenAI
- ElevenLabs
- Azure
- Fish Audio
- Cartesia
- LiveKit
- Daily
- Pipecat
- WebRTC
Features
Lifelike lip sync without the heavy lifting
Everything a talking avatar needs, in a package small enough for any web app.
Any voice, any provider
It only listens to the audio. No text, phoneme timings or provider visemes, so it works with every TTS and voice agent.
Runs in the browser
A 1 MB model in plain JavaScript. No server, no API key, no ML runtime, and the audio never leaves the page.
Two lines to integrate
Point it at an avatar, hand it an audio stream or clip. It plays the audio and moves the mouth in sync.
Anticipates every sound
Like a real speaker, the lips move into each sound on time: streams are delayed 150 ms, clips are analysed ahead.
Lifelike faces
Coarticulated mouth shapes, lips that close on p, b and m, natural blinks and brows that react to listening and thinking.
Bring any avatar
Any .glb with Oculus visemes works out of the box, and other rigs can map their own morph names in one option.
How it works
From sound waves to mouth shapes in 10 ms steps
- 01
Listen
Audio from any source is resampled to 16 kHz and turned into 10 ms spectral frames as it plays.
- 02
Recognise
A small causal neural network names the mouth shape (viseme) for every frame, looking 20 ms ahead.
- 03
Animate
Shapes are blended the way real speech flows from one sound into the next, driving the avatar's face in sync with the audio.
Quick start
Two lines from voice to face
Install the package, point it at an avatar and give it audio. It handles playback, the model, lip sync and rendering. Framework-free, with a React component and TypeScript types.
npm install @interviewflowai/talking-avatar threeimport { TalkingAvatar } from "@interviewflowai/talking-avatar";
const avatar = new TalkingAvatar({
container: document.getElementById("avatar"),
avatarUrl: "/avatar.glb", // any .glb with Oculus visemes
});
button.onclick = () => avatar.speak("/hello.mp3"); // or avatar.attachStream(stream)import { TalkingAvatarView } from "@interviewflowai/talking-avatar/react";
export function Agent({ stream }) {
return (
<div style={{ width: 640, height: 360 }}>
<TalkingAvatarView avatarUrl="/avatar.glb" stream={stream} />
</div>
);
}const pc = new RTCPeerConnection();
// Give the agent's voice to the avatar instead of an <audio> element.
pc.ontrack = (event) => avatar.attachStream(event.streams[0]);import { RoomEvent, Track } from "livekit-client";
room.on(RoomEvent.TrackSubscribed, (track) => {
if (track.kind === Track.Kind.Audio) {
avatar.attachStream(new MediaStream([track.mediaStreamTrack]));
}
});Use cases
Made for conversational AI
AI voice agents
Put a face on agents built with OpenAI Realtime, LiveKit Agents, Pipecat or your own WebRTC stack.
AI interviewers
Built by InterviewFlowAI for its AI interviewers, which screen candidates over phone and video.
Tutors and support
Friendlier tutoring, onboarding and customer-support agents that look at the user and talk.
Narration and presenters
Turn TTS narration, product tours and announcements into a talking presenter.
Measured, not guessed
Accurate, light and fast
Accuracy measured on 39 speakers the model never trained on, with audio passed through the WebRTC Opus codec.
- of 10 ms frames get the right mouth shape on unseen speakers
- 82%
- of p, b and m sounds close the lips
- 99%
- of one CPU core while speaking
- ~2%
- of gzipped code, plus a 1 MB model cached once
- 8 KB
FAQ
Questions, answered
What is Talking Avatar?
An open-source JavaScript SDK by InterviewFlowAI that lip syncs a 3D avatar to any audio in real time. You give it an avatar URL and an audio stream or clip; a small in-browser model drives the mouth.
Does it work with OpenAI, ElevenLabs or LiveKit voice agents?
Yes. It only listens to the audio, so any TTS provider or WebRTC voice agent works: pass the agent's audio stream to attachStream().
Do I need a server or an API key?
No. Everything runs in the browser and the audio never leaves the page. The only downloads are the avatar and the 1 MB model.
Which voices give the best lip sync?
TTS voices. The model was trained on clean, clearly spoken read speech, which is what TTS produces. Real voices from a microphone work too, but noise and echo make the lip sync less accurate.
Can I use my own avatar?
Yes: any .glb with Oculus viseme morph targets, or map your avatar's own morph names with the morphTargets option.
Which languages does it support?
It was trained on English. On other languages its timing holds up well, with sounds English lacks drawn using the nearest English mouth shape. Try Spanish, Hindi, Japanese and more in the live demo.
Is it open source?
Yes. The SDK and the model are on GitHub, built and maintained by InterviewFlowAI.
Built and open-sourced by InterviewFlowAI