ECHO
Dyadic 3D facial motion generation with asymmetric deterministic articulation and stochastic reaction.
Takeaway. A structured anchor–residual formulation separates speech-constrained articulation from one-to-many listener reactions under dual-stream audio-only input.
Problem
Dyadic 3D facial motion generation must model two asymmetric behaviors. Speech articulation is relatively constrained by phonetic content, while listener reactions remain inherently diverse and one-to-many.
Method
ECHO decomposes facial motion into a deterministic anchor and a stochastic residual. The anchor preserves stable speech-related structure; the residual captures diverse interaction dynamics. A training-only Motion Memory regularizer supplies local motion priors for weakly conditioned listening windows, while semantic-group scaling controls residual injection across facial regions.
Contributions
- Asymmetric modeling: deterministic articulation for speaking and stochastic interaction dynamics for listening.
- Training-only motion prior: Motion Memory strengthens local structure without becoming an inference dependency.
- Audio-only deployment: conversational facial motion is generated from dual-stream audio at inference time.
Citation
@inproceedings{cai2026echo,
title = {ECHO: Dyadic 3D Facial Motion Generation with Asymmetric Deterministic Articulation and Stochastic Reaction},
author = {Cai, Zhuoqiang and Sun, Yujie and Niu, Chaoyue and Yu, Hongyun and Chen, Zhiwen and Lv, Chengfei and Wu, Fan},
booktitle = {Proceedings of the 34th ACM International Conference on Multimedia},
year = {2026},
doi = {10.1145/3767308.3834964}
}