Multimodal Generative Models for Interactive Intelligence
Zhuoqiang Cai
M.S. Student in Computer Science and Technology at Shanghai Jiao Tong University
I develop generative models for expressive digital humans and context-aware intelligent agents.
- ACM MM 2026 First Author
- CVPR 2026 Co-author
- CVPR 2025 Challenge 1st Place
Research output
Selected Publications
- ACM MM
ECHO: Dyadic 3D Facial Motion Generation with Asymmetric Deterministic Articulation and Stochastic ReactionACM MM · 2026 First AuthorA structured anchor–residual formulation that separates speech-constrained articulation from one-to-many listener reactions under dual-stream audio-only input.
In Proceedings of the 34th ACM International Conference on Multimedia, Nov 2026 - CVPR
FHAvatar: Fast and High-Fidelity Reconstruction of Face-and-Hair Composable 3D Head Avatar from Few Casual CapturesCVPR · 2026 Co-authorA fast few-shot 3D Gaussian avatar framework with composable face and hair representations, supporting real-time animation and editing.
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026
Research agenda
From expressive humans to interactive intelligence
I am an M.S. student in Computer Science and Technology at Shanghai Jiao Tong University. My research develops multimodal generative models for interactive digital humans and long-horizon intelligent agents.
My recent work includes ECHO (ACM MM 2026, first author), which models the asymmetric dynamics of speaking articulation and listener reactions for dyadic 3D facial motion, and FHAvatar (CVPR 2026, co-author), which reconstructs composable face-and-hair avatars from a few casual captures. I also build local-first interactive desktop systems.
Conversational Facial Motion
Modeling speech-constrained articulation and diverse listener reactions in dyadic interactions.
ECHO · ACM MM 2026 · First Author →Animatable 3D Avatars
Reconstructing composable face-and-hair avatars from sparse, in-the-wild observations.
FHAvatar · CVPR 2026 · Co-author →Long-Horizon Intelligent Agents
Building agents that maintain causal context and recognize appropriately timed opportunities for assistance.
Ongoing Research →Recognition
Selected Honors
2025
1st Place, Dynamic Novel View Synthesis Track
CVPR 2025 NeRSemble Benchmark Challenge
2024
Huawei Intelligent Base Scholarship
Independent engineering
Engineering Highlight
Public Source · Tauri · React · Rust · Swift
Focus Pet
A local-first desktop focus companion that turns attention rhythms into a responsive virtual pet.
Background
Experience & Education
Experience
2024.07 — 2025.09
Intern
Taobao and Tmall Group, Alibaba Group
Education
2025 — Present
Master’s Student, Computer Science and Technology
Shanghai Jiao Tong University
2021 — 2025
Bachelor’s, Software Engineering
Shanghai Jiao Tong University
Updates
Recent News
| Aug 01, 2026 | ECHO was accepted to ACM Multimedia 2026. |
|---|---|
| Mar 24, 2026 | FHAvatar was accepted to CVPR 2026 and is available on arXiv. |
| Jun 01, 2025 | Our team won 1st place in the Dynamic Novel View Synthesis Track of the CVPR 2025 NeRSemble Benchmark Challenge. |
Contact
Let’s talk research.
For research discussions or collaboration, please reach me by email.