Hao Yin (殷皓)

Actively looking for research internships in the United States

Master’s Student,
Artificial Intelligence & Data Science,
University of Science and Technology of China (USTC).

Research: My research primarily focuses on multimodal large language models (MLLMs), with an emphasis on enhancing their perception and reasoning capabilities through reinforcement learning–based post-training strategies. Previously, I investigated methods to accelerate MLLM inference by pruning redundant tokens based on internal information flow, as well as techniques to reduce object hallucination and improve visual grounding through test-time interventions. Broadly, my work aims to develop generalizable methodologies for building more intelligent, efficient, and reliable multimodal AI systems.

Background: I am currently pursuing my Master’s degree in Artificial Intelligence & Data Science at the University of Science and Technology of China, advised by Zilei Wang in the Vision and Multimedia Research Group. I received my Bachelor’s degree in Mathematics and Applied Mathematics from China University of Mining and Technology, where I worked with Hu Shao.

news

Feb 25, 2026	Our paper REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding has been accepted to CVPR 2026! This work proposes a tool-augmented MLLM reasoning framework that enables introspective reasoning across both visual and textual modalities, significantly improving long-form video understanding.
Jan 15, 2026	Excited to share that I’ve completed my internship at Tencent and have begun a research internship at Xiaomi! My project focuses on enhancing the spatial perception capabilities of MLLMs, primarily through post-training methods, emphasizing effective reasoning frameworks and innovative training strategies.
Oct 7, 2025	Our paper The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs? has been accepted to NeurIPS 2025! This work reveals that contrastive decoding fails to genuinely mitigate hallucinations. Any apparent improvements are merely artifacts of confounding factors, not true effectiveness.
Sep 15, 2025	Excited to share that I’ve begun a research internship at Tencent! My project will involve exploring how to boost the common sense reasoning abilities of MLLMs. This will primarily be achieved through post-training approaches, focusing on efficient data construction and innovative training strategy design.
Feb 28, 2025	Our paper Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference has been accepted to CVPR 2025! This work investigates the internal visual information flow patterns in MLLMs and proposes a novel training-free inference acceleration method based on our findings.
Feb 28, 2025	Our paper ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large Language Models has been accepted to CVPR 2025! In this work, we present a method to mitigate object hallucination in MLLMs by strengthening attention to visual input.

selected publications

CVPR 2026
REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding

Jiaze Li^*, Hao Yin^*, Wenhui Tan^*, Jingyang Chen^*, Boshen Xu, and 5 more authors

In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Feb 2026

Abs Bib PDF

Self-reflection mechanisms that rely on purely text-based rethinking processes perform well in most multimodal tasks. However, when directly applied to long-form video understanding scenarios, they exhibit clear limitations. The fundamental reasons for this lie in two points: (1)long-form video understanding involves richer and more dynamic visual input, meaning rethinking only the text information is insufficient and necessitates a further rethinking process specifically targeting visual information; (2) purely text-based reflection mechanisms lack cross-modal interaction capabilities, preventing them from fully integrating visual information during reflection. Motivated by these insights, we propose REVISOR (REflective VIsual Segment Oriented Reasoning), a novel framework for tool-augmented multimodal reflection. REVISOR enables MLLMs to collaboratively construct introspective reflection processes across textual and visual modalities, significantly enhancing their reasoning capability for long-form video understanding. To ensure that REVISOR can learn to accurately review video segments highly relevant to the question during reinforcement learning, we designed the Dual Attribution Decoupled Reward (DADR) mechanism. Integrated into the GRPO training strategy, this mechanism enforces causal alignment between the model’s reasoning and the selected video evidence. Notably, the REVISOR framework significantly enhances long-form video understanding capability of MLLMs without requiring supplementary supervised fine-tuning or external models, achieving impressive results on four benchmarks including VideoMME, LongVideoBench, MLVU, and LVBench.
@inproceedings{li2025revisor, title = {REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding}, author = {Li, Jiaze and Yin, Hao and Tan, Wenhui and Chen, Jingyang and Xu, Boshen and Qu, Yuxun and Chen, Yijing and Ju, Jianzhong and Luo, Zhenbo and Luan, Jian}, year = {2026}, month = feb, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition}, }
NeurIPS 2025
The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs

Hao Yin, Guangzong Si, and Zilei Wang

In Advances in Neural Information Processing Systems, Dec 2025

Abs Bib PDF Code

Contrastive decoding strategies are widely used to reduce hallucinations in multimodal large language models (MLLMs). These methods work by constructing contrastive samples to induce hallucinations and then suppressing them in the output distribution. However, this paper demonstrates that such approaches fail to effectively mitigate the hallucination problem. The performance improvements observed on POPE Benchmark are largely driven by two misleading factors: (1) crude, unidirectional adjustments to the model’s output distribution and (2) the adaptive plausibility constraint, which reduces the sampling strategy to greedy search. To further illustrate these issues, we introduce a series of spurious improvement methods and evaluate their performance against contrastive decoding techniques. Experimental results reveal that the observed performance gains in contrastive decoding are entirely unrelated to its intended goal of mitigating hallucinations. Our findings challenge common assumptions about the effectiveness of contrastive decoding strategies and pave the way for developing genuinely effective solutions to hallucinations in MLLMs.
@inproceedings{yin2025mirageperformancegainscontrastive, title = {The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs}, author = {Yin, Hao and Si, Guangzong and Wang, Zilei}, year = {2025}, month = dec, booktitle = {Advances in Neural Information Processing Systems}, }
CVPR 2025
ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models

Hao Yin, Guangzong Si, and Zilei Wang

In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun 2025

Abs Bib PDF Code

Contrastive decoding strategies are widely used to mitigate object hallucinations in multimodal large language models (MLLMs). By reducing over-reliance on language priors, these strategies ensure that generated content remains closely grounded in visual inputs, producing contextually accurate outputs. Since contrastive decoding requires no additional training or external tools, it offers both computational efficiency and versatility, making it highly attractive. However, these methods present two main limitations: (1) bluntly suppressing language priors can compromise coherence and accuracy of generated content, and (2) processing contrastive inputs adds computational load, significantly slowing inference speed. To address these challenges, we propose Visual Amplification Fusion (VAF), a plug-and-play technique that enhances attention to visual signals within the model’s middle layers, where modality fusion predominantly occurs. This approach enables more effective capture of visual features, reducing the model’s bias toward language modality. Experimental results demonstrate that VAF significantly reduces hallucinations across various MLLMs without affecting inference speed, while maintaining coherence and accuracy in generated outputs.
@inproceedings{yin2025clearsightvisualsignalenhancement, title = {ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models}, author = {Yin, Hao and Si, Guangzong and Wang, Zilei}, year = {2025}, month = jun, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition}, }
CVPR 2025
Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference

Hao Yin, Guangzong Si, and Zilei Wang

In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun 2025

Abs Bib PDF Code

Multimodal large language models (MLLMs) improve performance on vision-language tasks by integrating visual features from pre-trained vision encoders into large language models (LLMs). However, how MLLMs process and utilize visual information remains unclear. In this paper, a shift in the dominant flow of visual information is uncovered: (1) in shallow layers, strong interactions are observed between image tokens and instruction tokens, where most visual information is injected into instruction tokens to form cross-modal semantic representations; (2) in deeper layers, image tokens primarily interact with each other, aggregating the remaining visual information to optimize semantic representations within visual modality. Based on these insights, we propose Hierarchical Modality-Aware Pruning (HiMAP), a plug-and-play inference acceleration method that dynamically prunes image tokens at specific layers, reducing computational costs by approximately 65% without sacrificing performance. Our findings offer a new understanding of visual information processing in MLLMs and provide a state-of-the-art solution for efficient inference.
@inproceedings{yin2025liftingveilvisualinformation, title = {Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference}, author = {Yin, Hao and Si, Guangzong and Wang, Zilei}, year = {2025}, month = jun, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition}, }