Experience
Visual Algorithm Intern
2026.01 – 2026.08Tencent WXG
Led reward-model research for WeChat's in-house text-to-image model and the Xiaowei assistant T2I module, covering preference data, evaluation, training, distillation, and RL post-training.
- Proposed Pair2Point, a two-stage reward paradigm: train a pairwise reasoning model with full-parameter SFT on Qwen3-VL-8B-Instruct, then distill it into an efficient pointwise scorer. The pointwise model reached SOTA accuracy on 3 independent benchmarks.
- Built a high-quality preference dataset: 20K clustered prompts, 35K candidate pairs, 8-dimensional rubrics, and 20K high-confidence preference pairs plus 70K pointwise multi-dimensional scores.
- Applied the pointwise reward to Flow-GRPO and DiffusionNFT for RL post-training of Flux.1-dev, Z-Image, Qwen-Image, and WeChat's in-house generator, achieving SOTA on GenAI-Bench, UniGenBench, PickScore, HPSv2, and HPSv3.
Text-to-ImageReward ModelPost-Training
Visual Algorithm Intern
2024.07 – 2025.12RabbitPre
Worked on computer vision and generative AI for sponsored projects and products, spanning pose control, image editing, video understanding, and content generation.
- Built a skeleton-driven character video generation pipeline with pose control and automatic New Year video synthesis.
- Developed product-image generation and relighting with inpainting, supporting multi-scene commercial visuals.
- Built a short-video clone and remix platform that combines video parsing, visual understanding, text rewriting, and generation.
Video GenerationImage EditingMLLM