About
I am a fourth-year Ph.D. student at Peking University (expected graduation: July 2027), advised by Prof. Jian Zhang. I received my B.E. degree in Computer Science (Zhufeng Honors Program) from Sichuan University in 2022.
My research focuses on image and video generation, including text-to-image post-training, reward modeling, knowledge editing for diffusion transformers, and controllable panoramic video generation. I am familiar with data pipelines, common algorithms, and codebases for generative model post-training, and have experience with large-scale (100-GPU) training. My work has received 800+ citations with an h-index of 10.
I am currently on the job market — if you have a suitable position, please feel free to contact me via email.
News
One paper (DualDiff3D) has been accepted to ECCV 2026 🎉
One paper (OmniDrag) has been accepted to IJCV 🎉
One paper (AlignedGen) has been accepted to NeurIPS 2025 🎉
Selected Publications
View All →Pair2Point: Distilling Pairwise Reasoning into Pointwise Rewards for High-Quality Text-to-Image Generation
Qian Wang, Xinhang Leng, Yinan Li, Sujie Hu, Huaisong Zhang, Chen Li, Jian Zhang
Under Review 2027
A two-stage reward modeling paradigm that trains a pairwise comparison model and distills it into an efficient pointwise scorer, with corresponding benchmarks and preference datasets.
AgentKE: Agentic Knowledge Editing for Diffusion Transformers
Qian Wang, Shuyu Wang, Zhile Guan, Xinhua Cheng, Xiandong Meng, Jian Zhang
Under Review 2027
An agentic knowledge-editing framework for text-to-image DiTs that decomposes natural-language intents with a VLM, edits linear projection layers in the protection-set null space, and reflects on the editing results.
DualDiff3D: Dual Structure-Appearance Diffusion Priors for Reliability-Enhanced 3D Gaussian Splatting
Qian Wang, Yu Wang, Weiqi Li, Xinhua Cheng, Xiandong Meng, Ronggang Wang, Jian Zhang
European Conference on Computer Vision (ECCV) 2026
Uses two diffusion models to extract structure from low-quality novel views and appearance from reference views, with a reliability-enhanced Render-Refine-Optimize loop for sparse-view 3DGS.
360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model
Qian Wang, Weiqi Li, Chong Mou, Xinhua Cheng, Jian Zhang
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024
A controllable panoramic video generation method with the WEB360 dataset, supporting text and dense optical flow as motion conditions.
