ECHO

Dyadic 3D facial motion generation with asymmetric deterministic articulation and stochastic reaction.

ACM Multimedia 2026 First Author Conversational Facial Motion
ECHO method and conversational facial-motion results

Takeaway. A structured anchor–residual formulation separates speech-constrained articulation from one-to-many listener reactions under dual-stream audio-only input.

Problem

Dyadic 3D facial motion generation must model two asymmetric behaviors. Speech articulation is relatively constrained by phonetic content, while listener reactions remain inherently diverse and one-to-many.

Method

ECHO decomposes facial motion into a deterministic anchor and a stochastic residual. The anchor preserves stable speech-related structure; the residual captures diverse interaction dynamics. A training-only Motion Memory regularizer supplies local motion priors for weakly conditioned listening windows, while semantic-group scaling controls residual injection across facial regions.

Contributions

  • Asymmetric modeling: deterministic articulation for speaking and stochastic interaction dynamics for listening.
  • Training-only motion prior: Motion Memory strengthens local structure without becoming an inference dependency.
  • Audio-only deployment: conversational facial motion is generated from dual-stream audio at inference time.

Citation

@inproceedings{cai2026echo,
  title     = {ECHO: Dyadic 3D Facial Motion Generation with Asymmetric Deterministic Articulation and Stochastic Reaction},
  author    = {Cai, Zhuoqiang and Sun, Yujie and Niu, Chaoyue and Yu, Hongyun and Chen, Zhiwen and Lv, Chengfei and Wu, Fan},
  booktitle = {Proceedings of the 34th ACM International Conference on Multimedia},
  year      = {2026},
  doi       = {10.1145/3767308.3834964}
}

References