Rendered live. Spoken in real time. Behaving as one system.
Athira combines live browser rendering, streaming speech, phoneme-level facial animation, and conversation-driven behavior to create an AI presence that feels responsive rather than pre-recorded.
How it works
Every layer runs on every frame. Select one to see what it contributes.
A GLB asset with an ARKit-52 blendshape rig, drawn by WebGL every frame from the avatar’s current state - never a clip being replayed.
Layers are listed in the order a frame is produced.
Live browser rendering.
Frames come from the avatar current state, never from looped or pre-rendered video.
WebGL and Three.js drive a 3D GLB asset using the ARKit-52 blendshape standard.
Physically based skin and eye shading, environmental lighting, and restrained post-processing.
Real-time media.
Two-way audio and avatar video travel over WebRTC through a managed real-time media layer, plugin-free, in the browser the person already has open. Designed for natural, real-time conversation.
Modular speech-to-speech.
Three stages, each replaceable. Provider-agnostic by design, continuously evaluated for quality and latency, with fallbacks at every stage.
Transcription while the person is still talking.
What the reply should be, decided against the task at hand.
Audio streaming out as it is generated.
Behavior beyond the mouth.
Movement is driven by what is happening in the conversation, not played from canned loops.
Carrying the tone of what is being said.
Following the stress in the sentence.
At human intervals, not on a fixed timer.
Present between turns, so stillness never looks frozen.
Weight shifting the way a person’s does.
Responding to what the person just said.
Fast enough that people talk rather than wait.
Additional languages are planned.