Publications & Research

Selected research projects, peer-reviewed publications, and patents.

Selected Research

Auditing Data Leakage in Whole-Slide Image Multimodal Benchmarks

We audit the striking zero-shot claims of pathology vision–language models on whole-slide image VQA and find them compromised by data leakage at two levels: patient-level (slides from one case split across train and test) and institutional-level (different cases sharing staining-batch and scanner signatures through a common Tissue Source Site). Tracing slide, case, and TSS identifiers across public resources, we document 92.3–100% case-level train–test overlap on TCGA-derived benchmarks, show both leakage levels are linearly decodable from foundation-model features, and find reported accuracies concentrate on the most contaminated benchmarks — so current evaluation cannot separate genuine multimodal reasoning from nearest-neighbor retrieval over memorized artifacts. We close with concrete recommendations for contamination-free evaluation.

2026 · Under review, WACV 2026 (Datasets & Benchmarks Track) · arXiv:2607.12278

CleanSlide: A Leakage-Audited and Shortcut-Controlled Benchmark for Whole-Slide Vision–Language Models

The largest contamination-free VLM benchmark built on TCGA: no Tissue Source Site overlap, no patient-ID overlap, and cleaned question construction. Questions balance answer positions to remove position bias, use strong distractors to avoid easy-option shortcuts, guard against report-to-answer tautology, and carry patient/site isolation through to the question level. Items answerable by text-only models are filtered or balanced, and the blind (vision-free) baseline accuracy is reported as a cleanliness metric.

2026 · In preparation · Code & data

Vision–Language Models for Computational Pathology: A Survey of Paradigms, Granularities, and Benchmarks

A VLM-specific synthesis of computational pathology organized along two axes — a training-paradigm axis (contrastive dual-encoder, generative instruction-tuned, reasoning/RL-enhanced, agent-based, and VLM-augmented multiple-instance learning) and a granularity axis (patch, region, whole-slide). Its centerpiece is an evidence-graded paradigm–task suitability matrix, calibrated with controlled probe studies and head-to-head comparisons against pure-vision foundation models rather than leaderboard numbers.

2026 · In preparation · Paper list

LANTERN: Learning a Data-Efficient Pathology Foundation Model through Multimodal Disentangled Representations

A three-level LLM-based image-description framework over 15 public pathology datasets yielding the disentangled Balance9k dataset; fine-tuning DINOv3 on 9k curated samples rivals million-scale fine-tuning. Cell-Guided Attention and organ-agnostic learning improve cross-organ generalization by ~30%, and Text-Guiding Pooling mitigates LLM hallucination effects.

2025 · University of Virginia

RENet: Reliable and Efficient Cross-Modality Learning for Unsupervised VI-ReID

A hashing-based metric for reliable cross-modal similarity (+2.39% Rank-1 over common distance metrics) and a multi-directional update module that efficiently bridges the modality gap, achieving state-of-the-art performance.

2024

Brain Age Prediction from T1-MRI — Research Internship, Wan Lab, UNMC

Six-month internship at the University of Nebraska Medical Center: enhanced a local U-Net for brain age prediction with Swin Transformer blocks; code open-sourced at wan-mlab/Swin-U-NET.

2024

Modality-Aware Pedestrian Attentive Learning for VI-ReID

PedMix, a data augmentation tailored to visible-infrared re-identification (+4.67% mAP), combined with a modality feature transfer module that fuses cross-attention and convolution at minimal overhead.

2023

Lightweight Directional-Aware Network for Classification

A comprehensive review and taxonomy of non-learned-operator deep models, plus a lightweight directional-aware network reaching state-of-the-art accuracy at a fraction of the computational cost.

2023 · Artificial Intelligence Review (IF 10.7)

Publications

Patents