About

Wang Chen陈旺

Ph.D. student in AI at Xiamen University MAC Lab

I study how long-horizon visual information should enter multimodal models, and how generation can help a model understand the world. My current questions are concrete: how to preserve event structure in continuous video, generate inspectable candidate explanations, and return to the original evidence to revise them.

Xiamen, ChinaResearch intern at AMap, Alibaba Group, since May 2026I also occasionally run the WeChat account AI Hacker
Wang Chen by a mountain lake
Wang ChenResearch & life

Research vision

Understanding a world in motion

I want to build multimodal systems that can keep seeing, organizing, and understanding information over real-world time scales. The path starts with event structure in long video, then asks how generation can propose testable hypotheses and how a model can accumulate evidence without losing it as a visual stream continues.

01

Visual experience in the real world rarely arrives as a tidy stack of independent images. An experiment, a journey, or a match can last for hours, while the evidence that changes a judgment may appear only briefly. A model has to notice when an event changes and remember how those changes connect. This motivates my first research thread: organizing semantic boundaries, event anchors, and query-relevant information so that a finite input still retains a coherent account of what happened.

02

Generation adds another possibility. A model can propose an event description, a missing state, or a possible future, then return to its observations to look for support and contradiction. These generations are candidates rather than evidence. Their value lies in turning an ambiguous internal state into something that can be inspected. If a candidate cannot be grounded in time, entities, or causal order, it should be revised or rejected.

03

The system I ultimately want to build should maintain a working account of events while video keeps arriving. It should know what has been confirmed, what remains hypothetical, and when earlier evidence needs to be revisited. An answer is only one output of this process. More important is an event representation that can keep changing with new information while preserving where its evidence came from.

Research

A research path taking shape

Start from event structure in continuous video, let generation propose candidates, and let evidence test them.

01

Observe

Find structure in continuous visual streams

Treat video as events unfolding in time rather than isolated frames. Semantic boundaries, event anchors, and long-term memory determine what the model actually sees.

EFSWFS-SB
02

Propose

Turn an uncertain reading into inspectable candidates

Generate event descriptions, missing states, or possible futures so an unfinished judgment becomes an intermediate object that can be localized, compared, and falsified.

Generative understandingHypothesis grounding
03

Verify

Let understanding keep changing with evidence

Place each candidate back on the timeline and against the original visual evidence. Keep what is supported, revise conflicts, and update event memory over long horizons.

Evidence tracingStreaming video understanding
Explore the research path

Publications

Selected work

These projects preserve events, boundaries, and query cues at different scales, providing the evidence base for what comes next.

WFS-SB diagram showing query relevance signal, wavelet transform, and semantic boundaries2026

CVPR 2026

Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding

Wang Chen, Yuhui Zeng, Yongdong Luo, Tianyu Xie, Luojun Lin, Jiayi Ji, Yan Zhang, Xiawu Zheng

Detects semantic boundaries from a noisy query-frame relevance signal, then allocates frame budgets by segment importance. The training-free method improves VideoMME, MLVU, and LongVideoBench by 5.5, 9.5, and 6.2 points.

QuoTA performance curves under different visual token budgets2026

AAAI 2026 · Equal contribution

QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension

Yongdong Luo*, Wang Chen*, Weizhong Huang, Shukang Yin, Haojia Lin, Jinfa Huang, Chaoyou Fu, Jiayi Ji, Xiawu Zheng, Jiebo Luo

Allocates visual tokens to frames before cross-modal interaction. CoT turns a complex query into scorable clues, improving six video benchmarks by 3.2 points on average under the same token budget.

EFS pipeline for event partitioning, anchor localization, and global refinement2026

arXiv preprint · Equal contribution

Event-Anchored Frame Selection for Effective Long-Video Understanding

Wang Chen*, Yongdong Luo*, Yuhui Zeng, Luojun Lin, Tianyu Xie, Fei Chao, Rongrong Ji, Xiawu Zheng

Partitions video into coherent events, localizes a query-relevant anchor in each event, and applies adaptive MMR for global refinement across coverage, relevance, and diversity.

View all publications

Ongoing

Two questions I am pursuing

These directions are in progress. For now, I state the questions and leave the claims to experiments.

A

How generation can help understanding

Studying how generative objectives, internal features, and candidate hypotheses can expose what a model has not yet understood, then using observations to ground, test, and revise them.

B

Long-horizon streaming video understanding

Exploring hierarchical memory, event updates, and evidence revisiting so models can sustain long-term understanding under low-latency constraints.

News

Updates

  1. Beginning the Ph.D. stage in Artificial Intelligence at Xiamen University.

  2. Started an internship at AMap, Alibaba Group, working on multimodal and video understanding.

  3. WFS-SB was accepted to CVPR 2026 and the code was released.

  4. QuoTA appeared at AAAI 2026.

Blog

Recent notes

Browse all notes

Timeline

Experience

From Sep 2026

Xiamen University · AI · Ph.D. stage

MAC Lab, advised by Prof. Liujuan Cao and Assoc. Prof. Xiawu Zheng

Sep 2024 - Aug 2026

Xiamen University · AI · M.S.-Ph.D. track

Research in long-video understanding and multimodal reasoning

May 2026 - Present

AMap · Alibaba Group · Internship

Multimodal and video understanding

Sep 2020 - Jun 2024

Fuzhou University · AI · B.Eng.

Began research in generative vision and facial aesthetics