Publications
Pair2Point: Distilling Pairwise Reasoning into Pointwise Rewards for High-Quality Text-to-Image Generation
Qian Wang, Xinhang Leng, Yinan Li, Sujie Hu, Huaisong Zhang, Chen Li, Jian Zhang
Under Review 2027
A two-stage reward modeling paradigm that trains a pairwise comparison model and distills it into an efficient pointwise scorer, with corresponding benchmarks and preference datasets.
AgentKE: Agentic Knowledge Editing for Diffusion Transformers
Qian Wang, Shuyu Wang, Zhile Guan, Xinhua Cheng, Xiandong Meng, Jian Zhang
Under Review 2027
An agentic knowledge-editing framework for text-to-image DiTs that decomposes natural-language intents with a VLM, edits linear projection layers in the protection-set null space, and reflects on the editing results.
DualDiff3D: Dual Structure-Appearance Diffusion Priors for Reliability-Enhanced 3D Gaussian Splatting
Qian Wang, Yu Wang, Weiqi Li, Xinhua Cheng, Xiandong Meng, Ronggang Wang, Jian Zhang
European Conference on Computer Vision (ECCV) 2026
Uses two diffusion models to extract structure from low-quality novel views and appearance from reference views, with a reliability-enhanced Render-Refine-Optimize loop for sparse-view 3DGS.
OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation
Weiqi Li, Shijie Zhao, Chong Mou, Xuhan Sheng, Zhenyu Zhang, Qian Wang, Junlin Li, Li Zhang, Jian Zhang
International Journal of Computer Vision (IJCV) 2026
Enables motion control for omnidirectional image-to-video generation.
Garment De-Warping for Virtual Try-On in the Wild
Yu Gu, Jiexuan Zhang, Qian Wang, Jian Zhang
2025 IEEE International Conference on Image Processing (ICIP) 2025
AlignedGen: Aligning Style Across Generated Images
Jiexuan Zhang, Yiheng Du, Qian Wang, Weiqi Li, Yu Gu, Jian Zhang
Advances in Neural Information Processing Systems (NeurIPS) 2025
Aligns visual style across a set of generated images.
OD-VAE: An Omni-Dimensional Video Compressor for Improving Latent Video Diffusion Model
Liuhan Chen, Zongjian Li, Bin Lin, Bin Zhu, Qian Wang, Shenghai Yuan, Xing Zhou, Xinhua Cheng, Li Yuan
2025 IEEE International Conference on Multimedia and Expo (ICME) 2025
Mind-Edit: MLLM Insight-Driven Editing via Language-Vision Projection
Shuyu Wang, Weiqi Li, Qian Wang, Shijie Zhao, Jian Zhang
arXiv preprint arXiv:2505.19149 2025
Prompt2Poster: Automatically Artistic Chinese Poster Creation from Prompt Only
Shaodong Wang, Yunyang Ge, Liuhan Chen, Haiyang Zhou, Qian Wang, Xinhua Cheng, Li Yuan
ACM International Conference on Multimedia (ACM MM) 2024
An automatic poster creation framework that uses an LLM to extract user intention from prompts and generate an aligned background.
360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model
Qian Wang, Weiqi Li, Chong Mou, Xinhua Cheng, Jian Zhang
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024
A controllable panoramic video generation method with the WEB360 dataset, supporting text and dense optical flow as motion conditions.
NTIRE 2023 Challenge on 360deg Omnidirectional Image and Video Super-Resolution: Datasets, Methods and Results
Mingdeng Cao, Chong Mou, Fanghua Yu, Xintao Wang, Yinqiang Zheng, Jian Zhang, Chao Dong, Gen Li, Ying Shan, Qian Wang, others
IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) 2023
A spatial-temporal two-stage model for omnidirectional image and video super-resolution in the NTIRE 2023 challenge.
Panoptic Compositional Feature Field for Editable Scene Rendering with Network-Inferred Labels via Metric Learning
Xinhua Cheng, Yanmin Wu, Mengxi Jia, Qian Wang, Jian Zhang
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2023
Uses metric learning to leverage 2D network-inferred labels for discriminating feature fields, enabling 3D segmentation and editing.
Deep Generalized Unfolding Networks for Image Restoration
Chong Mou, Qian Wang, Jian Zhang
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2022
Integrates a gradient estimation strategy into proximal gradient descent, enabling unfolding networks to handle complex real-world degradations.
More is Better: Multi-Source Dynamic Parsing Attention for Occluded Person Re-Identification
Xinhua Cheng, Mengxi Jia, Qian Wang, Jian Zhang
ACM International Conference on Multimedia (ACM MM) 2022
Introduces multi-source knowledge ensemble in occluded person re-ID to leverage external semantic cues from different domains.
A Simple Visual-Textual Baseline for Pedestrian Attribute Recognition
Xinhua Cheng, Mengxi Jia, Qian Wang, Jian Zhang
IEEE Transactions on Circuits and Systems for Video Technology (TCSVT) 2022
Models pedestrian attribute recognition as a multimodal problem and captures intra- and cross-modal correlations with a simple visual-textual baseline.
Detection Features as Attention (Defat): A Keypoint-Free Approach to Amur Tiger Re-Identification
Xinhua Cheng, Jianing Zhu, Nan Zhang, Qian Wang, Qijun Zhao
2020 IEEE International Conference on Image Processing (ICIP) 2020