Experience

Visual Algorithm Intern

2026.01 – 2026.08

Tencent WXG

Led reward-model research for WeChat's in-house text-to-image model and the Xiaowei assistant T2I module, covering preference data, evaluation, training, distillation, and RL post-training.

  • Proposed Pair2Point, a two-stage reward paradigm: train a pairwise reasoning model with full-parameter SFT on Qwen3-VL-8B-Instruct, then distill it into an efficient pointwise scorer. The pointwise model reached SOTA accuracy on 3 independent benchmarks.
  • Built a high-quality preference dataset: 20K clustered prompts, 35K candidate pairs, 8-dimensional rubrics, and 20K high-confidence preference pairs plus 70K pointwise multi-dimensional scores.
  • Applied the pointwise reward to Flow-GRPO and DiffusionNFT for RL post-training of Flux.1-dev, Z-Image, Qwen-Image, and WeChat's in-house generator, achieving SOTA on GenAI-Bench, UniGenBench, PickScore, HPSv2, and HPSv3.
Text-to-ImageReward ModelPost-Training

Visual Algorithm Intern

2024.07 – 2025.12

RabbitPre

Worked on computer vision and generative AI for sponsored projects and products, spanning pose control, image editing, video understanding, and content generation.

  • Built a skeleton-driven character video generation pipeline with pose control and automatic New Year video synthesis.
  • Developed product-image generation and relighting with inpainting, supporting multi-scene commercial visuals.
  • Built a short-video clone and remix platform that combines video parsing, visual understanding, text rewriting, and generation.
Video GenerationImage EditingMLLM