I am a Ph.D. student (2020-2026) at the State Key Lab of CAD&CG, Zhejiang
University, advised by Prof. Guofeng Zhang and Prof. Zhaopeng Cui.
During my Ph.D., I work closely with Yinda Zhang.
I used to visit Computer Vision and Geometry Group at ETH Zürich, advised by Marc Pollefeys.
My research interests are in multimodal model, video generation, world model and 3D computer vision.
We present a unified multimodal model for holistic digital human generation and understanding across text, audio, motion, semantic video, image, and video.
We scale a DiT-based mixture-of-experts video foundation model with more than 70,000 hours of embodied data for physically grounded video generation and embodied intelligence.
We introduce a one-step diffusion model that purifies sparse-view reconstructions corrupted by distractors, improving multi-view consistency and rendering quality.
We introduce a benchmark and multimodal automated evaluation framework for retrieval-augmented visually-rich generation, covering execution correctness, design quality, and content quality.
We propose an Atlanta-world guided, implicit-structured Gaussian representation for efficient and smooth surface reconstruction in indoor and urban scenes.
We present a large-scale synthetic urban dataset with diverse, controllable illumination for benchmarking intrinsic decomposition, inverse rendering, and multi-illumination reconstruction.
We extend neural mesh-based implicit fields with geometry, texture, and semantic controls for versatile, efficient, and interactive volumetric editing.
We propose a new GS with video diffusion model for novel view synthesis and surface reconstruction from extremely sparse(3-4), unposed images in unbounded 360 scenes.
We introduce a novel frequency-aware framework for view synthesis that simultaneously captures the overall scene structure and high-definition details within a single NeRF model.
We present a novel neural rendering framework, which is able to learn accurate geometry and reflection of
the mirror and support various scene manipulation applications.
We present a novel mesh-based implicit field with disentangled geometry and texture codes on mesh vertices, which facilitates a set of editing functionalities.
We present a multi-device integrated cargo loading management system with AR, which monitors cargoes by fusing perceptual information from multiple devices in real-time.