Zhuoqiang Cai

Multimodal Generative Models for Interactive Intelligence

Zhuoqiang Cai

M.S. Student in Computer Science and Technology at Shanghai Jiao Tong University

I develop generative models for expressive digital humans and context-aware intelligent agents.

  • ACM MM 2026 First Author
  • CVPR 2026 Co-author
  • CVPR 2025 Challenge 1st Place
Portrait of Zhuoqiang Cai

Research output

Selected Publications

View all publications →
  1. ACM MM
    echo-teaser.png
    ECHO: Dyadic 3D Facial Motion Generation with Asymmetric Deterministic Articulation and Stochastic Reaction
    Zhuoqiang Cai, Yujie Sun, Chaoyue Niu, Hongyun Yu, Zhiwen Chen, Chengfei Lv, and Fan Wu
    ACM MM · 2026 First Author

    A structured anchor–residual formulation that separates speech-constrained articulation from one-to-many listener reactions under dual-stream audio-only input.

    In Proceedings of the 34th ACM International Conference on Multimedia, Nov 2026
  2. CVPR
    fhavatar-teaser.png
    FHAvatar: Fast and High-Fidelity Reconstruction of Face-and-Hair Composable 3D Head Avatar from Few Casual Captures
    Yujie Sun, Zhuoqiang Cai, Chaoyue Niu, Jianchuan Chen, Zhiwen Chen, Chengfei Lv, and Fan Wu
    CVPR · 2026 Co-author

    A fast few-shot 3D Gaussian avatar framework with composable face and hair representations, supporting real-time animation and editing.

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026

Research agenda

From expressive humans to interactive intelligence

I am an M.S. student in Computer Science and Technology at Shanghai Jiao Tong University. My research develops multimodal generative models for interactive digital humans and long-horizon intelligent agents.

My recent work includes ECHO (ACM MM 2026, first author), which models the asymmetric dynamics of speaking articulation and listener reactions for dyadic 3D facial motion, and FHAvatar (CVPR 2026, co-author), which reconstructs composable face-and-hair avatars from a few casual captures. I also build local-first interactive desktop systems.

Conversational Facial Motion

Modeling speech-constrained articulation and diverse listener reactions in dyadic interactions.

ECHO · ACM MM 2026 · First Author →

Animatable 3D Avatars

Reconstructing composable face-and-hair avatars from sparse, in-the-wild observations.

FHAvatar · CVPR 2026 · Co-author →

Long-Horizon Intelligent Agents

Building agents that maintain causal context and recognize appropriately timed opportunities for assistance.

Ongoing Research →

Recognition

Selected Honors

2024

Huawei Intelligent Base Scholarship

Independent engineering

Engineering Highlight

Explore projects →

Background

Experience & Education

Experience

2024.07 — 2025.09

Intern

Taobao and Tmall Group, Alibaba Group

Hangzhou, China

Education

2025 — Present

Master’s Student, Computer Science and Technology

Shanghai Jiao Tong University

Shanghai, China

2021 — 2025

Bachelor’s, Software Engineering

Shanghai Jiao Tong University

Shanghai, China

Updates

Recent News

All news →
Aug 01, 2026 ECHO was accepted to ACM Multimedia 2026.
Mar 24, 2026 FHAvatar was accepted to CVPR 2026 and is available on arXiv.
Jun 01, 2025 Our team won 1st place in the Dynamic Novel View Synthesis Track of the CVPR 2025 NeRSemble Benchmark Challenge.

Contact

Let’s talk research.

For research discussions or collaboration, please reach me by email.