Request a demo
Avatar Technology

Rendered live. Spoken in real time. Behaving as one system.

Athira combines live browser rendering, streaming speech, phoneme-level facial animation, and conversation-driven behavior to create an AI presence that feels responsive rather than pre-recorded.

How it works

Every layer runs on every frame. Select one to see what it contributes.

A GLB asset with an ARKit-52 blendshape rig, drawn by WebGL every frame from the avatar’s current state - never a clip being replayed.

Layers are listed in the order a frame is produced.

A person speaking with Athira, with the render layers it produces annotated alongside

Live browser rendering.

Every frame is generated

Frames come from the avatar current state, never from looped or pre-rendered video.

Rendered in the browser

WebGL and Three.js drive a 3D GLB asset using the ARKit-52 blendshape standard.

Video-call quality

Physically based skin and eye shading, environmental lighting, and restrained post-processing.

Explore Core Avatar
Render loopGeometryDelivery01GLB asset02ARKit-52 blendshapes03Three.js scene04PBR shading05Lighting06Post-processing07WebRTC frame

Real-time media.

Two-way audio and avatar video travel over WebRTC through a managed real-time media layer, plugin-free, in the browser the person already has open. Designed for natural, real-time conversation.

The personAthira avatarVoice in, as they speakAudio and video, streaming backWebRTC / no plugin
The pipeline

Modular speech-to-speech.

Three stages, each replaceable. Provider-agnostic by design, continuously evaluated for quality and latency, with fallbacks at every stage.

Can you walk me through the second stage of the process
Listening live
Speech to text

Transcription while the person is still talking.

StandbyReasoning
StandbyReasoning
StandbyReasoning
StandbyReasoning
Reasoning

What the reply should be, decided against the task at hand.

Streaming audio
Text to speech

Audio streaming out as it is generated.

Behavior beyond the mouth.

Movement is driven by what is happening in the conversation, not played from canned loops.

Emotional expression

Carrying the tone of what is being said.

Emphasis

Following the stress in the sentence.

Blinking

At human intervals, not on a fixed timer.

Breathing

Present between turns, so stillness never looks frozen.

Subtle sway

Weight shifting the way a person’s does.

Nods and head tilts

Responding to what the person just said.

Latency
Responses typically begin within a second

Fast enough that people talk rather than wait.

Language
English is supported today

Additional languages are planned.

Bring a lifelike AI presence into your product.