Yuxiao Yang
I am currently a Researcher at ByteDance. Previously, I received my master’s degree from Tsinghua University, advised by Prof. Haoqian Wang, and my bachelor’s degree in Information Engineering from Zhejiang University.
My research interest includes video generation, 3D generation, and multimodal learning.
📧 Email / 💻 Github / 📚 Google Scholar
💼 Experience
ByteDance|Researcher|2026.02 - Now- Contributing to the Seedance series of generative video models.
MiniMax|TopTalent Intern|2025.12 - 2026.01- Contributed to video foundation model development.
Alibaba Group|Research Intern|2025.02 - 2025.12- Worked on complex human motion video generation and unified multimodal generation.
Baidu|Research Intern|2024.12 - 2025.02- Explored auto-regressive multi-view image generation.
HKUST|Research Intern|2024.06 - 2024.11- Worked on high-fidelity 3D textured mesh generation.
📝 Publications

EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
Y Yang, H Sheng, S Cai, J Lin, J Wang, B Deng, J Lu, H Wang, J Ye
🌐 Website 📄 Paper 💻 Code 🤖 Model
- We present EchoMotion, a novel diffusion-based framework for unified human video and motion generation.

Wonder3D++: Cross-Domain Diffusion for High-Fidelity 3D Generation From a Single Image
Y Yang, X Long, Z Dou, C Lin, Y Liu, Q Yan, Y Ma, H Wang, Z Wu, W Yin
📄 Paper 💻 Code 🤖 Model
- We present Wonder3D++, a novel cross-domain diffusion framework for high-fidelity 3D generation from a single image.

AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual References
J Wang, H Sheng, S Cai, Y Yang, W Zhang, C Yan, B Deng, J Ye
🌐 Website 📄 Paper
- We present AnyID, an ultra-fidelity identity-preserving video generation framework from diverse visual references.

Auto-Regressively Generating Multi-View Consistent Images
JK Hu*, Y Yang*, J Liu, J Wu, C Zhao, Y Lu
📄 Paper 💻 Code 🤖 Model
- We propose a novel auto-regressive model for generating multi-view consistent images.

NOVA3D: Normal Aligned Video Diffusion Model for Single Image to 3D Generation
Y Yang, P Li, Y Zhang, J Lu, X He, M Qin, W Wang, H Wang
📄 Paper
- We introduce NOVA3D, a normal aligned video diffusion model for single image to 3D generation.

Mvreward: Better aligning and evaluating multi-view diffusion models with human preferences
W Wang, H Xu, Y Yang, Z Liu, J Meng, H Wang
📄 Paper
- We propose Mvreward, a novel framework for better aligning and evaluating multi-view diffusion models with human preferences.

Target-Balanced Score Distillation
Z Xu, Q Wang, Y Yang, L Zhang, Z Liang, Y Li
📄 Paper
- We introduce Target-Balanced Score Distillation, a novel method for improving the performance of score-based generative models.
