ShopX
A model-native fulfillment framework for agentic shopping. ShopX lets large models reason, retrieve, rank, and act directly in product space—preserving intent across complex, multi-turn shopping journeys.
Hi there, I’m
Algorithm Engineer · Alibaba Inc.
I received my bachelor’s degree from Huazhong University of Science and Technology (HUST) in June 2019 and my Ph.D. from the Technical University of Denmark (DTU) in March 2025, advised by Prof. Aasa Feragen. During my Ph.D., I was also affiliated with the Pioneer Centre for AI. In 2023, I was a visiting researcher with Prof. Qi Dou’s group at the Chinese University of Hong Kong (CUHK).
My current research centers on native multimodal understanding and multimodal retrieval: tokenizing heterogeneous information into native model capabilities while keeping model decisions grounded and interpretable—particularly through mid-training and post-training.
Latest progress
Recent work at the intersection of large models, retrieval, and real-world commerce.
A model-native fulfillment framework for agentic shopping. ShopX lets large models reason, retrieve, rank, and act directly in product space—preserving intent across complex, multi-turn shopping journeys.
Research
* equal contribution · † corresponding author
Across these works, one question keeps recurring: how can heterogeneous information be represented, grounded, and ultimately fused into models that understand, retrieve, and reason over it natively?
Models, training, and evaluation across visual, spatial, temporal, and language signals—from grounded understanding to inference-time alignment.

Builds an in-the-wild benchmark for AI-generated image detection from human-curated and composite-pipeline images. MIRAGE-R1 adds reflective vision-language reasoning with adaptive thinking to balance reliability and speed.

Introduces a comprehensive ultrasound benchmark with 7,241 cases across 15 anatomical regions and eight clinically inspired tasks. Its evaluation exposes persistent gaps in spatial reasoning and clinical language generation.

Adapts a multimodal large language model from natural images to 3D prostate MRI through a two-stage training process. Clinical PI-RADS guidelines are distilled into a lightweight scoring network without extra annotations or deployment parameters.

Unifies multi-view images, point clouds, and biomedical language knowledge for operating-room scene graphs. Temporal interactions across 2D and 3D streams improve recognition of long-form surgical activity.

Steers a frozen vision-language model during decoding with dense, token-level visual preference signals instead of retraining the backbone. A lightweight reward model improves visual faithfulness across 13 benchmarks with negligible inference overhead.
From shortcut awareness and interpretable clinical concepts to label-efficient anomaly detection—making model decisions traceable to the evidence they use.

Shows how clinical annotations and dataset construction can become shortcuts in medical image segmentation. Controlled tests reveal failure modes and practical mitigation strategies for more stable generalization.

Models the expert process of seeing, conceiving, and concluding through progressive visual and property concepts. The resulting explanations are human-interpretable, intervenable, and more robust across external ultrasound datasets.

Frames fetal brain anomaly detection as an unsupervised problem, training diffusion models only on normal ultrasound scans. iNAAD aggregates inpainted reconstructions across noise levels to detect and localize rare abnormalities without anomaly labels.

Reframes feature attribution as both credit allocation along a path and principled path selection. Optimal generative transport produces more stable, structured explanations while preserving competitive faithfulness.
Turning point clouds, topology, shape, and physical scale into model-readable representations—the path toward tokenizing heterogeneous information.

Replaces fixed point-cloud neighborhoods with density-adaptive queries and learned masked attention. The resulting graph convolution is robust to noise while remaining efficient for 3D classification and scene understanding.

Learns topological consistency from data using paired lightweight U-Nets and self-supervised contrastive signals. It improves both pixel accuracy and structural continuity without requiring a predefined topology.

Combines transvaginal ultrasound with cervical segmentation predictions and pixel-spacing information to estimate spontaneous preterm birth risk. The shape- and spatially-aware model outperforms clinical and machine-learning baselines while providing interpretable, calibrated predictions.
Community
Journey
A decade-long path through different ways of representing and understanding information.
I began in 2016 with time-series signals, asking how learning systems could read patterns unfolding over time. That question moved into 3D point clouds, then medical images and explainable computer vision. I later worked across tabular data and language. Today, these threads come together in multimodal foundation models—building systems that can understand, retrieve, and reason across heterogeneous information.
Along the way: a Ph.D. at DTU, a research visit at CUHK, and now algorithm engineering at Alibaba.
Off the clock
Small rituals, growing interests, and perhaps a few notes I’ll write here someday.
Snivy is my all-time favorite.
Discovering new places slowly, one walk and one small story at a time.
Learning the craft with an espresso machine—from grinding beans and refining the extraction to practicing latte art.
Lunch breaks are for books—a quiet reset in the middle of the day.
Making something from scratch, whether it’s dinner or something fresh from the oven.
A soft spot for desserts, with a good drink alongside them.
Country-level visits and total page views.