VLX-Flow: streaming video understanding model
Streaming video-understanding model: it processes video in chunks and answers from updated memory. The repo has docs only; weights are not released.
Project facts
GitHub Ecosystem- License
- Apache-2.0
- Stars
- 188
- Data checked
- 2026-10-10
Snapshot figures reflect the check date and may change over time.
Most video-understanding models treat a video as a file: load the whole clip, analyze it once, return an answer. Camera feeds, robots and drones don’t work that way. The input keeps arriving, and questions can come at any moment. VLX-Flow, released by OmAI Lab in June 2026, splits the input into consecutive chunks, encodes each new chunk as it arrives, and updates the model’s internal memory. When someone asks a question, the model answers from that memory instead of reprocessing the full history.

The overview figure in the README shows that pipeline: chunked input, two memory layers, and low-latency interaction.
Core features
- Streaming chunks: the video is split into short, consecutive chunks of a few frames each and processed in temporal order. Earlier history stays in compressed form inside the model state instead of being appended as raw frames again and again.
- Two-layer memory: a visual cache holds recent frame-level detail for immediate answers and event detection. Semantic memory holds higher-level context built from the stream and the conversation, including streaming descriptions, earlier observations, and user questions and answers, which keeps longer narratives coherent.
- Linear Attention: the language model includes Linear Attention components. Standard self-attention needs a KV cache that grows with the sequence. Linear Attention carries history in a recurrent state that updates as each chunk arrives, so memory grows more smoothly.
- Low-latency answers: when a question arrives, the model answers from the state it already maintains instead of recomputing the full history.

The README describes this chart qualitatively. Full Attention rises as history grows. SlideWindow resets its window and produces a rise-reset-rise sawtooth. VLX-Flow compresses history through its two-layer memory and keeps time to first token low and stable. The chart has no millisecond values, and the README gives no answer-quality scores, so read it as a statement of design direction, not a measured result.
Typical use cases
- Continuous captioning: someone watching a production line or a warehouse wants to know what just happened, at any time. The README’s “streaming captioning” targets exactly this.
- Ask while watching: a robot or drone receives frames while it moves and answers questions like “is that box still there?” as they come up.
- Event alerts: the README says the model can be extended toward alerts when its maintained state meets a condition. That is a stated direction; the repo contains no code for it.
Quick start
There is nothing to install from this repository yet. The root holds only the README, LICENSE, CONTRIBUTING.md and an assets folder. There is no model code, and the README’s Release section reads “Checkpoints: Coming soon”, so the weights are not out. To see the model work today, the README points to the “Try VLX” platform at om-agent.com, which is the hosted route. Come back for local setup once the Release section lists checkpoints.
Summary
Worth watching for anyone following real-time video understanding, robotics, or security cameras, or tracking how streaming multimodal models develop. Skip it if you want to run something locally this week. The repo has no code, the weights are marked Coming soon, and the README publishes no benchmark scores, parameter count, or training data. The repository is licensed Apache-2.0. Its last push was on September 24, 2026, and the star count of 188 was checked on October 10, 2026. If you need a video tool you can use today, see video-use, which lets a coding agent edit footage directly.