From Video Compression to Codec-Native Multimodal Models
An English reader's map to the full Chinese audit of video codecs, OneVision-Encoder, codec-video-prep, LLaVA-OneVision-2, and Mage-VL.
Blog
Technical details worth explaining beyond papers, with sources, versions, and open questions kept visible.
An English reader's map to the full Chinese audit of video codecs, OneVision-Encoder, codec-video-prep, LLaVA-OneVision-2, and Mage-VL.
Why does the same foundation model behave so differently across agent systems? This long-form essay follows execution loops, context, durable state, tools, evaluation, and rollback through ACE, MCE, Meta-Harness, ADAS, AFlow, STOP, and Self-Harness.