Commentary#Video generation#Coding agent

When the Model Won’t Paint Pixels: Opus 5.5’s Video Week

Anthropic never advertised video for Opus 5.5. Within a week users made it a category: 1,401 videos, four public bills, one fight over whether it is video.

Earthrise as photographed by Apollo 8 in 1968: Earth hanging above the gray lunar horizon (photo by NASA, public domain)

Opus 5.5 cannot paint a single pixel — the official docs say so outright: text and images in, text out. And yet it became one of the most prolific video makers on the internet this week. Deedydas, a partner at Menlo Ventures, cut a five-minute documentary on 250 years of American history with it; a group-theory promo made from one Chinese sentence pulled hundreds of thousands of views on Bilibili; and one developer’s corpus dedicated to these clips collected 1,401 deduped videos in three days. None of it is generation. It is all programs. A week of public post-mortems has been enough to map the mechanism, the limits and the price — worth doing the accounting in the open.

A video movement, grown in a week

The timeline first. Opus 5.5 arrived in launch week on September 22 with no video mention in any official material. Within three days, the corpus built by a developer called atheremeroy held 1,401 deduped videos, with the daily count climbing from 121 on the 22nd to 537 on the 24th, then falling back to 213 the next day (a dissection of the corpus posted by the X account baboonAI4S; snapshot frozen around September 26). The launch thread on Hacker News took 1,806 points and 1,129 comments — almost all model comparison; video came up in exactly one subthread. Nobody official pushed this. Users backed into it.

Once they had, the infrastructure grew in a week: atheremeroy’s corpus updates daily with per-video sourcing; zhuyansen’s collection lists 986 works and 259 prompts; the editing tool video-use (from the browser-use team) climbed to 27,613 stars; shipvideo, the engine behind launchvideo.io, open-sourced its whole pipeline on September 24. The Chinese scene kept pace: about 228k views for the group-theory promo and 165k for the same-prompt Opus-vs-GPT 6 Astra comparison on Bilibili (as of September 30), while des13, a Taichung web shop, put the model straight into client work — four brand animations with prompts, timings and token bills published.

One genre detail: the top category in the corpus, at 21%, is “AI praising AI.” Product promos take 20%, interactive games 19%, explainers 12%. The first subject of this video wave is the AI itself.

It doesn’t paint pixels

The mechanism in one sentence: the model writes a program that renders every frame, then hands the frames to ffmpeg. zhuermu read the source of five finished videos from five independent runs and found nearly the same pipeline each time — an HTML/Canvas/WebGL page with a renderFrame(t) where every frame is a function of time, headless Chrome capturing frames, ffmpeg encoding and mixing, then the model sampling its own frames into a contact sheet, spotting what’s wrong, fixing the code and re-rendering. Deedydas’s version is blunter: “no video gen model was used here, no additional libraries. Just JavaScript, playwright and ffmpeg. Claude coded this website animation and took a video of it.”

That makes it a different species from Sora-class models. Sora guesses pixels; Opus writes a renderer. Guessed footage blurs text, grows extra fingers, melts chart labels. Written footage reruns after a one-line change. The artifact is code — readable, editable, reproducible. That is the line between engineering and pulling a slot machine.

Where the edges are

The corpus numbers draw the capability map: code-drawn motion graphics 33%, 3D rendering 29%, photorealism 2.8% (Gemini auto-labeled; read as direction). What it does best is what programmers do daily; what it does worst is the home turf of conventional video models. The best work in this ecosystem comes overwhelmingly from people who can code — the LHC proton collision was code-driven Blender, the house that grows from a sketch is Three.js, and one 80-second short is a single dependency-free HTML file.

Tighter than the visual edge is the interpretive one. zhuermu’s explainer on agent self-evolution shipped a gorgeous v1 where all 35 arXiv IDs were real — and the whole timeline sat askew, context engineering labeled 2024 though it went mainstream in June 2025. His phrase: “Every part was correct; the story assembled from them was tidier than history.” The ceiling is not how well it draws. It is what it takes your request to mean.

Its ears are fake too. It cannot listen to music, so the music-video workaround is an engineer’s: decode the song in Python, compute per-frame energy, bass and a beat grid into a table, and have every visual read the table — “It could not hear the song, so it measured it.”

The pushback arrived on schedule. Under a 427-point Hacker News thread (“Opus 5.5 is good at explainer videos”), top comments are unsentimental: this is “a presentation skill turned into video”; this genre was slop before AI; “my first reaction is to ask for the text version.” None of that is unfair. With photorealism at 2.8%, true live-action texture still means Sora or Veo paints and Opus directs — which is exactly how deedydas’s documentary was made (full flow promised this week).

The economics

After one week the price list is public, all self-reported: $0.90 to cut a 2-minute take (video-use’s own demo); $2.96 for an 84-second brand film; about $15 for an 8-minute-plus explainer with one rework; about $100 for a 40-plus-minute tutorial (cosine, roughly 20% of his subscription quota); and the top end — shneural’s 15-second motion-graphics showreel at 900 frames, 1h32m, an $81 API bill, with the model having commandeered the machine’s Blender unprompted. Tony Dinh’s before-and-after is the whole trend in one line: $1,000+ for this kind of promo a year ago, under 30 minutes now.

The real economics are structural, though. zhuermu set Opus against the whiteboard-video pipeline he built himself two months ago: the homegrown pipeline is a fixed script — near-zero marginal cost per video, one style forever; Opus needs no tooling, but rewrites the tool for every video and throws it away. And one line hides in the footnotes: prompt caching makes a revision cheaper than a video model’s re-roll — des13’s measured 4 to 12 minutes per natural-language change. Cheap iteration, not cheap video, is the actual wedge into creative workflows.

The water in the wave

The bubble surfaced the same week. baboonAI4S’s counts (auto-labeled corpus) put 282 videos as not-made-by-Opus, 39 exact-duplicate groups, and a rival video company — higgsfield — slipping in 22 posts; someone re-borrowed another creator’s 1,750-word prompt within 24 hours at 87% five-gram overlap — one documented case, anecdotal by definition, and still proof that prompts have become assets worth stealing. The daily curve peaked September 24 and halved the next day — one day inside a three-day snapshot; too early to call the tide out, naive to call it endless.

The curious part: the method itself survives reproduction. zhuermu re-ran three viral hits from their original prompts, one shot each — “launch demos are cherry-picked” did not survive contact. The water is in the corpus wall, not the method. And the emptiest shelf says the most: data visualization is 1.3% of 1,401 videos. The wave churns toward spectacle and self-reference; the work left is in the places nobody is filming.

The ceiling moved to the brief

Set the week’s post-mortems side by side and the closest thing to a conclusion is zhuermu’s: what decides the result is increasingly the brief you hand it, and the material behind it. Both his explainers were wrong in v1 and right in v2; the only difference was a fact-checked brief. cosine’s line — “the simpler the prompt, the more it invents, and the better” — doesn’t contradict it: a one-liner buys imagination; a deliverable buys verification. Two purchases, two situations.

The durable lesson isn’t any single video. It’s how the capability arrived: Anthropic never advertised video, and it was almost certainly not on the roadmap — users made it a category inside a week. The next one will probably not show up in a keynote either. It will show up in what people find while spending $3 to try. If that’s you today, we published a build-it-yourself playbook alongside this piece — the routes, the costs and the traps, from one line to a finished cut.