Hi there, I’m

Manxi
Lin.

Algorithm Engineer · Alibaba Inc.

I received my bachelor’s degree from Huazhong University of Science and Technology (HUST) in June 2019 and my Ph.D. from the Technical University of Denmark (DTU) in March 2025, advised by Prof. Aasa Feragen. During my Ph.D., I was also affiliated with the Pioneer Centre for AI. In 2023, I was a visiting researcher with Prof. Qi Dou’s group at the Chinese University of Hong Kong (CUHK).

My current research centers on native multimodal understanding and multimodal retrieval: tokenizing heterogeneous information into native model capabilities while keeping model decisions grounded and interpretable—particularly through mid-training and post-training.

Manxi Lin at Shanghai Disneyland on August 30, 2026 Manxi Lin by the sea in Yantai, May 2026
Shanghai Disneyland · Aug 30, 2026

What I’m building now

Recent work at the intersection of large models, retrieval, and real-world commerce.

July · 2026
From Intent to Item

ShopX

A model-native fulfillment framework for agentic shopping. ShopX lets large models reason, retrieve, rank, and act directly in product space—preserving intent across complex, multi-turn shopping journeys.

Selected publications

* equal contribution · corresponding author

Across these works, one question keeps recurring: how can heterogeneous information be represented, grounded, and ultimately fused into models that understand, retrieve, and reason over it natively?

Topic 01

Multimodal understanding, alignment & reasoning

Models, training, and evaluation across visual, spatial, temporal, and language signals—from grounded understanding to inference-time alignment.

MIRAGE benchmark and reflective reasoning teaser
AAAI 2026Vision-language reasoning
MIRAGE: Towards AI-Generated Image Detection in the Wild

Builds an in-the-wild benchmark for AI-generated image detection from human-curated and composite-pipeline images. MIRAGE-R1 adds reflective vision-language reasoning with adaptive thinking to balance reliability and speed.

Cheng Xia*, Manxi Lin*, Jiexiang Tan*, Xiaoxiong Du, Yang Qiu, Junjun Zheng, Xiangheng Kong, Yuning Jiang, Bo Zheng
U2-BENCH tasks across diverse anatomical regions
ICLR 2026Multimodal benchmark
U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding

Introduces a comprehensive ultrasound benchmark with 7,241 cases across 15 anatomical regions and eight clinically inspired tasks. Its evaluation exposes persistent gaps in spatial reasoning and clinical language generation.

Anjie Le… Manxi Lin, Hongcheng Guo (selected authors)
Two-stage multimodal model adaptation for PI-RADS scoring
MICCAI 2024MRI · Multimodal LLM
Incorporating Clinical Guidelines through Adapting Multi-modal Large Language Model for Prostate Cancer PI-RADS Scoring

Adapts a multimodal large language model from natural images to 3D prostate MRI through a two-stage training process. Clinical PI-RADS guidelines are distilled into a lightweight scoring network without extra annotations or deployment parameters.

Tiantian Zhang*, Manxi Lin*, Hongda Guo, Xiaofan Zhang, Ka Fung Peter Chiu, Aasa Feragen, Qi Dou
TriTemp-OR image, point cloud, language, and temporal fusion framework
MICCAI 2024Image · Point cloud · Language
Tri-modal Confluence with Temporal Dynamics for Scene Graph Generation in Operating Rooms

Unifies multi-view images, point clouds, and biomedical language knowledge for operating-room scene graphs. Temporal interactions across 2D and 3D streams improve recognition of long-form surgical activity.

Diandian Guo*, Manxi Lin*, Jialun Pei, He Tang, Yueming Jin, Pheng-Ann Heng
Comparison of training-time, sequence-level, and token-level vision-language alignment
Findings of ACL 2026Inference-time alignment
Token-level Inference-Time Alignment for Vision-Language Models

Steers a frozen vision-language model during decoding with dense, token-level visual preference signals instead of retraining the backbone. A lightweight reward model improves visual faithfulness across 13 benchmarks with negligible inference overhead.

Kejia Chen, Junjun Zheng, Jiawen Zhang, Manxi Lin, Xiao Pan, Jiacong Hu, Jian Lou, Zunlei Feng, Mingli Song
Topic 02

Grounding models in evidence & concepts

From shortcut awareness and interpretable clinical concepts to label-efficient anomaly detection—making model decisions traceable to the evidence they use.

Clinical annotation and dataset construction shortcuts in medical segmentation
MICCAI 2024Robust segmentation
Shortcut Learning in Medical Image Segmentation

Shows how clinical annotations and dataset construction can become shortcuts in medical image segmentation. Controlled tests reveal failure modes and practical mitigation strategies for more stable generalization.

Manxi Lin*, Nina Weng*, Kamil Mikolaj, Zahra Bashir, Morten Bo Søndergaard Svendsen, Martin Tolsgaard, Anders Nymark Christensen, Aasa Feragen
Progressive concept bottleneck model for fetal ultrasound quality assessment
arXiv 2022Concept bottlenecks
Explainable Fetal Ultrasound Quality Assessment with Progressive Concept Bottleneck Models

Models the expert process of seeing, conceiving, and concluding through progressive visual and property concepts. The resulting explanations are human-interpretable, intervenable, and more robust across external ultrasound datasets.

Manxi Lin, Aasa Feragen, Kamil Mikolaj, Zahra Bashir, Martin Grønnebæk Tolsgaard, Anders Nymark Christensen
Diffusion-based fetal brain anomaly detection pipeline
ASMUS 2024Unsupervised anomaly detectionBest Poster Award
Unsupervised Detection of Fetal Brain Anomalies using Denoising Diffusion Models

Frames fetal brain anomaly detection as an unsupervised problem, training diffusion models only on normal ultrasound scans. iNAAD aggregates inpainted reconstructions across noise levels to detect and localize rare abnormalities without anomaly labels.

Markus Ditlev Sjøgren Olsen, Jakob Ambsdorf, Manxi Lin, Caroline Taksøe-Vester, Morten Bo Søndergaard Svendsen, Anders Nymark Christensen, Mads Nielsen, Martin Grønnebæk Tolsgaard, Aasa Feragen, Paraskevas Pegios
Comparison of attribution maps from generative-flow and baseline methods
arXiv 2026Feature attribution
From Baselines to Transport Geodesics: Axiomatic Attribution via Optimal Generative Flows

Reframes feature attribution as both credit allocation along a path and principled path selection. Optimal generative transport produces more stable, structured explanations while preserving competitive faithfulness.

Cenwei Zhang, Lin Zhu, Manxi Lin, Lei You
Topic 03

Representing structured modalities

Turning point clouds, topology, shape, and physical scale into model-readable representations—the path toward tokenizing heterogeneous information.

Adaptive neighborhoods and masked attention in diffConv
ECCV 2022Point clouds
diffConv: Analyzing Irregular Point Clouds with an Irregular View

Replaces fixed point-cloud neighborhoods with density-adaptive queries and learned masked attention. The resulting graph convolution is robust to noise while remaining efficient for 3D classification and scene understanding.

Manxi Lin, Aasa Feragen
DTU-Net texture and topology network architecture
IPMI 2023 · OralTopology
DTU-Net: Learning Topological Similarity for Curvilinear Structure Segmentation

Learns topological consistency from data using paired lightweight U-Nets and self-supervised contrastive signals. It improves both pixel accuracy and structural continuity without requiring a predefined topology.

Manxi Lin, Zahra Bashir, Martin Grønnebæk Tolsgaard, Anders Nymark Christensen, Aasa Feragen
Shape- and spatially-aware ultrasound model for preterm birth prediction
ASMUS 2023Shape · Spatial informationBest Paper Honorable Mention
Leveraging Shape and Spatial Information for Spontaneous Preterm Birth Prediction

Combines transvaginal ultrasound with cervical segmentation predictions and pixel-spacing information to estimate spontaneous preterm birth risk. The shape- and spatially-aware model outperforms clinical and machine-learning baselines while providing interpretable, calibrated predictions.

Paraskevas Pegios, Emilie Pi Fogtmann Sejer, Manxi Lin, Zahra Bashir, Morten Bo Søndergaard Svendsen, Mads Nielsen, Eike Petersen, Anders Nymark Christensen, Martin Tolsgaard, Aasa Feragen

Academic service

Conference reviewerCVPR · ECCV · MICCAI · AAAI
Journal reviewerMedical Image Analysis · KBS

Across modalities

A decade-long path through different ways of representing and understanding information.

One question kept evolving.

I began in 2016 with time-series signals, asking how learning systems could read patterns unfolding over time. That question moved into 3D point clouds, then medical images and explainable computer vision. I later worked across tabular data and language. Today, these threads come together in multimodal foundation models—building systems that can understand, retrieve, and reason across heterogeneous information.

Along the way: a Ph.D. at DTU, a research visit at CUHK, and now algorithm engineering at Alibaba.

  1. 01
    Time-series signals2016 · where the story began
  2. 02
    Point cloudsLearning from 3D geometry
  3. 03
    Medical imagesVisual understanding in high-stakes settings
  4. 04
    Explainable computer visionUnderstanding why models decide
  5. 05
    Tabular dataLearning from structured information
  6. 06
    LanguageSemantics, context, and generation
  7. 07
    Multimodal foundation modelsWhere the threads meet today

A few things I like

Small rituals, growing interests, and perhaps a few notes I’ll write here someday.

01 · FAVORITE

Pokémon

Snivy is my all-time favorite.

02

Travel

Discovering new places slowly, one walk and one small story at a time.

03 · LEARNING

Making Coffee

Learning the craft with an espresso machine—from grinding beans and refining the extraction to practicing latte art.

04 · LUNCH BREAK

Reading

Lunch breaks are for books—a quiet reset in the middle of the day.

05 · IN THE KITCHEN

Cooking & Baking

Making something from scratch, whether it’s dinner or something fresh from the oven.

06 · INDULGENCES

Sweets & Spirits

A soft spot for desserts, with a good drink alongside them.

Ideas travel farther when shared. Let’s talk.

Audience

Visitors around the world

Country-level visits and total page views.

Visitor counter showing total page views and visitor countries