AVTR-1

Open Stack for Real-Time Interactive Avatars

Artem Kravtsov*, Dmitrii Ziganshin*, Vsevolod Poletaev*, Gleb Balitskiy, Anastasia Tikhonova, Egor Burkov, Vadim Lebedev

Avaturn Live · 2026   ·   * Core contributors: co-wrote the paper

Abstract

Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live system must synchronize the model's output with speech from an external voice agent, schedule video frames for playback, and handle interruptions. We introduce AVTR-1, an open stack for real-time interactive avatar conversations, built around a compact 153M-parameter autoregressive flow-matching motion generator conditioned on both participants' audio. We adapt its audio encoder for streaming through self-distillation. The stack turns the model's chunk-based generation into a continuous, synchronized audio-video stream driven by an external voice agent, and we analytically derive its contribution to the user-facing latencies and validate the resulting bounds with two commercial voice agents. Further experiments demonstrate that AVTR-1 leads the compared dyadic systems on all reported visual-quality metrics and most conventional listening-motion metrics while remaining competitive in lip synchronization. Its inference runtime operates in real time on data-center and consumer GPUs. However, conventional listening metrics do not establish whether the paired speaker's speech contributes to generated motion. We therefore introduce the Reference-Based Directed Granger Gain (R-DGG), which measures the additional predictive information carried by speaker speech after accounting for listener history and speaker motion. R-DGG finds statistically supported predictive dependence for recorded listeners and all evaluated dyadic systems, but not for talking-head generators without paired audio or mismatched speaker–listener pairs. We release the model weights, renderer, and serving backend under component-specific licenses.

Demo

Three dyadic Seamless Interaction calls below.

Model Architecture

The AVTR-1 motion model is a 153M-parameter conditional flow-matching Transformer that autoregressively generates five-frame motion conditioned on two separate audio channels called self and other audio. The motion model is a pre-norm Transformer decoder with 18 layers, 8 heads, and width 512. Each block receives past motion, audio features for both channels, and a static reference. We divide the 75-frame motion history between a near path over the last five frames and a far path over the earlier 70. The flow-matching timestep modulates each sub-block through adaptive layer normalization (AdaLN), with separate scale and shift predicted for each. The outputs are concatenated to form the final 42-dimensional velocity prediction.

AVTR-1 motion model architecture

Inference

We call the inference-side component that serves the motion model the renderer. It registers the source portrait once, then generates five frames at a time. For each motion chunk, the two audio channels are batched into one audio encoder call. Each request predicts a motion chunk, renders it with the cached LivePortrait appearance, and returns the updated state required by the next request.

AVTR-1 inference overview

The streamer's worklet and event-bus architecture. A transport, a conversation engine, and a rendering worklet communicate only through a central event bus. The conversation engine worklet adapts a voice agent, the rendering worklet holds one speech scheduler per speech stream, avatar and user, and drives the renderer, returning fused audio-video frames. Worklets read the stream clock directly.

AVTR-1 streamer worklet and event-bus architecture

Per-chunk latency and real-time factor across GPUs, measured offline on one five-frame (200 ms) chunk.

GPU Latency per chunk Real-time factor
NVIDIA L40S 71 ms 2.81×
NVIDIA A100 91 ms 2.20×
NVIDIA RTX 4060 Ti 162 ms 1.24×
NVIDIA RTX 3070 180 ms 1.11×
NVIDIA L4 202 ms 0.99×
NVIDIA RTX 3060 Ti 207 ms 0.97×
NVIDIA RTX 4060 232 ms 0.86×

Qualitative Comparison

★ = paired audio. For AvatarForcing*, we extend the hardcoded RoPE range and increase the maximum generation length from 30 to 600 seconds so that it can process full conversations.

Quantitative Results

AVTR-1 achieves the strongest overall performance among the compared dyadic systems, leading all visual-quality metrics and most conventional listening-motion metrics while remaining competitive in lip synchronization. We evaluate on 184 speaker–listener pairs from the conversational, improvised subset of the Seamless Interaction test split.

Visual quality and lip synchronization. Best value among the dyadic systems (Paired audio = Yes) in green; talking-head baselines shown for context.

Method Paired audio Visual quality Lip synchronization
FID ↓ FVD ↓ CSIM ↑ LSE-D ↓ LSE-C ↑
SoulX Lite No 12.1 89.3 0.91 6.74 3.38
SoulX Pro No 9.4 53.6 0.93 7.35 3.09
Ditto No 16.0 113.6 0.95 7.12 3.11
FLOAT No 11.7 83.2 0.90 6.99 2.93
AvatarForcing* Yes 14.4 85.8 0.77 7.32 2.41
DyStream Yes 43.1 119.4 0.87 6.57 3.21
AVTR-1 (Ours) Yes 14.3 76.8 0.94 7.08 3.28

Conventional listening-motion metrics among dyadic systems. For listening metrics without an arrow, values closer to ground truth are better. The pose components of PFD and variance are reported in units of 10−2. Best value among the dyadic systems in green.

Method rPCC ↓ PFD ↓ SID Var
Exp Pose Exp Pose Exp Pose Exp Pose
AvatarForcing* 0.109 0.163 37.51 7.560 4.749 3.848 1.213 2.099
DyStream 0.128 0.141 40.87 7.598 4.497 3.606 1.164 1.731
AVTR-1 (Ours) 0.083 0.140 25.98 6.417 4.970 3.254 1.313 0.930
Ground truth 0.000 0.000 0.000 0.000 5.231 4.001 1.463 1.789

Reference-Based Directed Granger Gain

A listener can follow the other participant's motion without responding to their speech. The motion metrics above cannot distinguish this behavior from a speech-conditioned response. To address this limitation, the Reference-Based Directed Granger Gain (R-DGG) measures how much the speaker's speech improves prediction of the listener's current motion after accounting for the listener's own history and the speaker's motion. R-DGG passes its validation check: the ground-truth interval remains above zero, while the interval for GT × other includes zero. All three dyadic systems have intervals above zero, while the intervals for every talking-head generator include zero. The intervals of the three dyadic systems overlap, so these results do not support a reliable ranking among them.

Values in units of 10−4 nats. p = Pr(G ≤ 0). Green marks estimates whose 95% interval excludes zero.

Method Paired audio R-DGG Bootstrap percentile p
2.5% 50% 97.5%
AvatarForcing* Yes 0.32 0.07 0.32 0.56 0.008
AVTR-1 (Ours) Yes 0.57 0.19 0.57 0.98 0.002
DyStream Yes 0.48 0.06 0.49 0.93 0.011
Ditto No 0.06 −0.20 0.06 0.31 0.335
FLOAT No 0.01 −0.33 0.00 0.42 0.491
SoulX Lite No −0.22 −0.54 −0.22 0.21 0.857
SoulX Pro No 0.18 −0.34 0.17 0.58 0.239
Ground truth — 0.72 0.32 0.72 1.12 <0.001
GT × other — −0.11 −0.56 −0.11 0.45 0.659

Citation

@misc{kravtsov2026avtr1openstackrealtime,
  title         = {AVTR-1: Open Stack for Real-Time Interactive Avatars},
  author        = {Artem Kravtsov and Dmitrii Ziganshin and Vsevolod Poletaev
                   and Gleb Balitskiy and Anastasia Tikhonova and Egor Burkov
                   and Vadim Lebedev},
  year          = {2026},
  eprint        = {2609.22913},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2609.22913}
}