Observe
Find structure in continuous visual streams
Treat video as events unfolding in time rather than isolated frames. Semantic boundaries, event anchors, and long-term memory determine what the model actually sees.
About
Ph.D. student in AI at Xiamen University MAC Lab
I study how long-horizon visual information should enter multimodal models, and how generation can help a model understand the world. My current questions are concrete: how to preserve event structure in continuous video, generate inspectable candidate explanations, and return to the original evidence to revise them.

Research vision
I want to build multimodal systems that can keep seeing, organizing, and understanding information over real-world time scales. The path starts with event structure in long video, then asks how generation can propose testable hypotheses and how a model can accumulate evidence without losing it as a visual stream continues.
Visual experience in the real world rarely arrives as a tidy stack of independent images. An experiment, a journey, or a match can last for hours, while the evidence that changes a judgment may appear only briefly. A model has to notice when an event changes and remember how those changes connect. This motivates my first research thread: organizing semantic boundaries, event anchors, and query-relevant information so that a finite input still retains a coherent account of what happened.
Generation adds another possibility. A model can propose an event description, a missing state, or a possible future, then return to its observations to look for support and contradiction. These generations are candidates rather than evidence. Their value lies in turning an ambiguous internal state into something that can be inspected. If a candidate cannot be grounded in time, entities, or causal order, it should be revised or rejected.
The system I ultimately want to build should maintain a working account of events while video keeps arriving. It should know what has been confirmed, what remains hypothetical, and when earlier evidence needs to be revisited. An answer is only one output of this process. More important is an event representation that can keep changing with new information while preserving where its evidence came from.
Research
Start from event structure in continuous video, let generation propose candidates, and let evidence test them.
Find structure in continuous visual streams
Treat video as events unfolding in time rather than isolated frames. Semantic boundaries, event anchors, and long-term memory determine what the model actually sees.
Turn an uncertain reading into inspectable candidates
Generate event descriptions, missing states, or possible futures so an unfinished judgment becomes an intermediate object that can be localized, compared, and falsified.
Let understanding keep changing with evidence
Place each candidate back on the timeline and against the original visual evidence. Keep what is supported, revise conflicts, and update event memory over long horizons.
Publications
These projects preserve events, boundaries, and query cues at different scales, providing the evidence base for what comes next.
2026CVPR 2026
Detects semantic boundaries from a noisy query-frame relevance signal, then allocates frame budgets by segment importance. The training-free method improves VideoMME, MLVU, and LongVideoBench by 5.5, 9.5, and 6.2 points.
2026AAAI 2026 · Equal contribution
Allocates visual tokens to frames before cross-modal interaction. CoT turns a complex query into scorable clues, improving six video benchmarks by 3.2 points on average under the same token budget.
2026arXiv preprint · Equal contribution
Partitions video into coherent events, localizes a query-relevant anchor in each event, and applies adaptive MMR for global refinement across coverage, relevance, and diversity.
Ongoing
These directions are in progress. For now, I state the questions and leave the claims to experiments.
Studying how generative objectives, internal features, and candidate hypotheses can expose what a model has not yet understood, then using observations to ground, test, and revise them.
Exploring hierarchical memory, event updates, and evidence revisiting so models can sustain long-term understanding under low-latency constraints.
News
Beginning the Ph.D. stage in Artificial Intelligence at Xiamen University.
Started an internship at AMap, Alibaba Group, working on multimodal and video understanding.
WFS-SB was accepted to CVPR 2026 and the code was released.
QuoTA appeared at AAAI 2026.
Blog
An English reader's map to the full Chinese audit of video codecs, OneVision-Encoder, codec-video-prep, LLaVA-OneVision-2, and Mage-VL.
Why does the same foundation model behave so differently across agent systems? This long-form essay follows execution loops, context, durable state, tools, evaluation, and rollback through ACE, MCE, Meta-Harness, ADAS, AFlow, STOP, and Self-Harness.
Timeline
From Sep 2026
MAC Lab, advised by Prof. Liujuan Cao and Assoc. Prof. Xiawu Zheng
Sep 2024 - Aug 2026
Research in long-video understanding and multimodal reasoning
May 2026 - Present
Multimodal and video understanding
Sep 2020 - Jun 2024
Began research in generative vision and facial aesthetics