Playbooks#Video editing#Coding agent#Prompts
Making Videos with Opus 5.5: From One Line to a Finished Cut
Opus 5.5 can’t paint a pixel — it writes the program that renders every frame. The routes, real costs and pitfalls from a week of public post-mortems.

About $3 and half an hour buys you an 84-second, 1080p cut with a soundtrack that lands on every edit point — text that never blurs, numbers that never melt, and every frame fixable by changing a line of code. The program that made it was written by Opus 5.5.
Fact-checked as of 2026-09-30. Every load-bearing number carries its source. Claims come in three grades: verified (official docs or first-hand confirmed material), creator-reported (the creator’s own account, not independently re-checked), and inferred (beyond what was disclosed; flagged wherever it appears). Unmarked claims are creator-reported.
Not a video model
Anthropic’s platform docs state it plainly: Opus 5.5 takes text and images in, and outputs text (platform docs). The model page positions it around agentic coding, computer use and knowledge work — no video anywhere, and our launch coverage didn’t mention it either. Every clip on your X timeline is a by-product of the model orchestrating external tools.
The blogger zhuermu paid for five videos out of pocket on the API, then read the source of all five. Five independent runs converged on nearly the same pipeline (creator-reported, blog): write an HTML/Canvas/WebGL page with a renderFrame(t) where every frame is a function of time; headless Chrome captures frames in parallel; ffmpeg encodes, concatenates and mixes; the model samples frames into a contact sheet, inspects, fixes the code, re-renders. In his words:
Opus 5.5 is not a video model like Sora. It does not paint pixels. It writes a program that draws every frame in a headless browser, then hands the frames to ffmpeg.
So this ecosystem has exactly three mechanisms: write a program that draws every frame; edit external footage; or direct an outside video model. This guide follows that split — routes 1 and 2 are all-code rendering, route 3 splits into live-action and generated footage. One rule of language throughout: we never write “it drew a frame” or “it generated a video,” only “the program it wrote rendered the frames” and “the footage it edited.”
Photorealism has a number. The X account baboonAI4S ran a dissect-and-count over the 1,401 deduped videos in atheremeroy’s corpus (the tweet never pins down which three days): code-drawn motion graphics 33%, 3D rendering 29%, photorealistic just 2.8%, the rest scattered across other styles (creator-reported, thread). The categories come from Gemini 3.8 Flash auto-labeling (1,119 “yes/likely” files), not human review — read them as direction, not census. If you want photoreal humans, skip to route 3: let a video model paint, let Opus direct.
Environment and budget
Three tools: Claude Code (Pro or 5x subscription, or metered API — zhuermu used the API), local ffmpeg, and headless Chrome (with Playwright, npx playwright install; or reuse your machine’s Chrome in headless mode, which is what zhuermu did). Voice-over, pick one: edge-tts — free, stable, flat (cosine’s choice); MiniMax (zhuermu’s; “not metered separately” was his exact wording, and the blog never clarifies whether that means free or just unitemized); Kokoro-82m (what deedydas — Menlo Ventures partner and Anthropic investor — used on his 3b1b-style paper explainer). Headless mode has one nasty trap; a dedicated lesson covers it below.
On cost, every figure is the author’s own account (creator-reported, not independently re-checked; zhuermu’s test was self-funded, no vendor credits):
| Scenario | Figure | Source |
|---|---|---|
| deedy’s 1-minute promo | “about $2” by his own claim; API or subscription, he doesn’t say | tweet |
| Third-party reproductions | $3.21 and $4 (OpenRouter, secondhand) | Traictory roundup |
| zhuermu’s five videos | at least $55.41 across eight runs — one timed-out run left no cost record, so this is a floor; $2.96–$14.79 each | blog |
| zhuermu’s rule of thumb | shorts $3–9; an 8-minute-plus explainer with one rework about $15, interruptions extra | same |
| cosine’s 40+ minute tutorial | about $100, roughly 20% of his subscription quota (tier unstated) | blog |
| Subscription math | a 1-minute finished video ≈ 1% of a 5x plan’s weekly quota; ~400 minutes if you burn the whole quota (i陆三金’s personal estimate — order of magnitude only) |
Every row is an n=1 self-report: deedy doesn’t state his metering, the reproductions are secondhand, cosine’s tier is unstated. Budget to an order of magnitude, no better. Three cost controls do have sources: make the agent quote before any expensive step (koldo2k’s rule — one line in the prompt); subscriptions carry a hard weekly cap, and if you hit it mid-render, a handover doc lets a fresh session pick up (cosine’s step 8; zhuermu’s interrupted runs survived exactly this way); on the API, /cost in Claude Code shows the running bill — Ctrl-C, shrink scope, rerun.
Wall-clock: an 84-second video ran about 24 minutes and 48 turns (zhuermu); a 15-second vertical with a second-by-second storyboard, 23 minutes plus 4 per natural-language revision (des13, a Taichung web shop); a 6-minute English lesson, about 45 minutes per round (0x0funky). cosine never published total time for his 40+ minute tutorial. Prompt caching makes “one more revision” cheaper than a video model’s re-roll — that’s i陆三金’s judgment call, no comparison numbers attached, weight it yourself; des13’s 4-minute revision is the same logic made concrete. His stack advice is a good default: HyperFrames or Remotion plus TTS as the base, then Gemini image gen, ElevenLabs or 3D if you want flash.
Before building anything, check the shelf: HyperFrames (setup is one command, npx hyperframes skills update; des13 shipped four brand animations with it and never opened an editor; v2code has a seven-step tutorial), buildwithhanif/claude-animation-skill, ClaudeAnimationBase (a p5.js starter), Lemo-Opuscar (39 styles), Papermotion (paper-cut animation engine), video-use (live-action editing; 27,613 stars when checked on 09-30, that number moves), shipvideo (behind launchvideo.io; one agent file plus three tools, self-reported ~90k input / 15k output tokens per video). Install one, ship one video, then decide whether to build your own pipeline.
zhuermu also keeps the honest scorecard on build-vs-rent: a homegrown pipeline is a fixed script — near-zero marginal cost, one style forever. With Opus there is no tool to build, but every video rewrites the tool from scratch, a few dollars to a dozen-plus each, in exchange for per-chapter visual design. His conclusion is the title of the lessons below: what decides the result is increasingly the brief you hand it, and the material behind it.
Shorts from a single line
The shortest route. Two ways in.
Bare: one prompt plus a delivery-constraints note. zhuermu attached the same note to all five videos; copy it as-is:
1080p 30fps; every visual generated by code; except the voice-over script
provided, no image/video/music/speech generation service; you may sample
frames to check your own work.
The “no speech generation service” line bans the agent from calling paid services on its own. Two ways to handle voice: synthesize audio yourself with edge-tts or MiniMax and hand the files over (zhuermu’s way), or amend the line to “you may synthesize voice-over locally with edge-tts” and let it run on your machine — cosine’s pipeline works this way (edge-tts is free and local, no paid service touched). One-line fix; don’t let this note stall your first day.
With HyperFrames: follow v2code’s seven-step tutorial — install Node and FFmpeg, write HTML with data-* attributes (the tutorial has full examples), preview, render.
Shortest first run (bare route): make an empty directory with two files — brief.md holding your goal and every fact you can source, voiceover.txt holding the narration line by line (or mp3 segments you already made). Start a Claude Code session in that directory and paste the template at the end of this piece. When it returns a storyboard and a 2–4 second animation sample of the hardest shot, reply in the same session: “storyboard approved, render the 0:04 shot first,” or “change the opening to …” — replying is the whole approval mechanism; no commands needed. Before accepting the final cut, ask for the contact sheet and actually look at it.
Before any of that, internalize the three-stage acceptance from athemeroy’s production guide: storyboard first; then an animated sample of the hardest 2–4 seconds; only after both pass, render the full piece and deliver a reproducible project. Skip the first two stages and every rework dollar is waste.
Two proofs come from zhuermu’s reproductions (creator-reported):
The Wasmer launch clip. Everything it got was one line from HN — “Make a modern slick and punchy video for this announcement:” — plus the blog post. 84 seconds, $2.96, ~24 minutes, 48 turns, one shot. The soundtrack came from numpy oscillators at 120 BPM, cut to every edit point. Before delivering, it sampled its own frames and caught three bugs: transparent title letters, two captions overlapping, an ffmpeg command overflowing the phone screen. One more detail: the blog said “about 5% slower than native” without giving the native number, so it labeled the native bar “baseline” instead of inventing a figure.
Earthrise. The entire input: “Recreate the moment Apollo 8 photographed Earthrise.” It found NASA Goddard’s reconstruction timeline (SVS 4129), LOLA lunar terrain and the 1968 ephemerides on its own, started at 16:38:13 UT, advanced in real time, and landed all three shutter releases on their historical timestamps. NASA’s official version works backwards from the photograph; this one replays the event along NASA’s timeline. zhuermu’s conclusion: rendering quality stopped being the bottleneck long ago — what it takes your request to mean is the bottleneck.
The third proof is from cosine’s collection: the group-theory promo (Bilibili UP sadssxa; the prompt is in the comments). One line of Chinese — “make me a promo film for abstract algebra, a few minutes, visually rich, not too abstract, mainly show the beauty of abstract algebra, rich and striking visuals, you know what I mean” — and no skills at all. It decided on its own to synthesize an original soundtrack in Python (D minor, 90 BPM, scene changes locked to bars), then wrote a 1,200-line 1080p animation page. About 228k views on Bilibili (BV12Zhm6oE6m).
Chinese-scene counterpoints from des13: one English prompt, a 40-second horizontal piece in 25 minutes; revisions in plain language — “the pigeon blocks the text, line him up with the other three,” “stop scrolling down” — 4 to 12 minutes per new version. Token bills published: 45k–83k output per video, 9.43m–12.96m total across four (cache reads included, and dominant).
A single line is not the only shape of constraint. riku720720 wrote a 3,459-character Japanese spec sheet for a pixel runner: Canvas 2D, 160×90, no external assets — spec the details down and the room for interpretation shrinks (thread). 3D works too: the Clearwater shallow-water sim is open source, and AxtonLiu’s pelican-cinema piece ran on a hand-written GPU raymarcher with procedural music (thread). Between one line and a spec sheet, pick your point on the cost-control curve.
One-line prompts probe the ceiling; deliverable work needs the reading pinned down. zhuermu reproduced three viral hits — Wasmer, Earthrise, P(doom) — one shot each, so “launch demos are cherry-picked” did not survive those three tests. Three samples, one operator: not a general law. His two explainers both failed v1 and passed v2 (see the lessons). For determinism, keep reading.
The long-explainer pipeline
The blogger cosine (余弦 — no relation to the Genie company) used this flow for a 40+ minute tutorial: illustrations, voice-over, sound, progress bar, episode splits, loudness normalization, cover — all arranged by the model (creator-reported, blog). Eight steps:
Steps 1–2, material and script. Feed it everything first: official docs, product screenshots, reference videos, dropped straight into the project directory (PDF, images, mp4). One hard rule: every claim in the narration must trace to something in the material — if it can’t be sourced, it doesn’t get said. Write the narration sentence by sentence, one sentence per visual. Get a few stills or a short preview and approve it before any full render; that’s where rework cost dies.
Steps 3–4, voice and captions. Synthesize voice one sentence at a time, cache each, redo only what changed. Captions cut on word-level timestamps; visuals follow the words. Sound effects generated in code; BGM auto-ducks under speech; loudness normalized per episode. edge-tts is free and stable — and flat.
Steps 5–6, packaging and speed. Progress bar, chapter timeline, episode splits, covers, show notes — all generated from one master timeline; change the content, rerun, everything syncs. If rendering drags, say “rendering is too slow, fix it” and it finds ways: parallel workers, one reused browser instance, slow-changing layers rendered at a third of the frames. cosine’s numbers: collection render 21 → 9 minutes, the progress-bar layer 15 → 2, a single episode 6 → 4.
Steps 7–8, acceptance and handover. After rendering, it self-checks: frame counts, A/V alignment, loudness, frame samples for caption collisions. The final watch-through is yours — its “looks fine” is not a publish button. When the context gets long, have it write a handover doc — stack, file structure, how to change things — so a fresh session can pick up.
The Chinese scene has an even lazier sample: Bilibili’s @美丫姐 fed it an article PDF, went to sleep, and woke up to a finished video and a cover — voice acting split by character, electronic timbre for the AI characters, warning music on cue, zero skills, zero keyframes (column). Her own yardstick: in August, AI editing still meant scripting keyframe descriptions by hand. Two months, one whole workflow apart.
Variants: dotey’s 12-minute Transformer explainer (thread); kimmonismus’s 3-minute AI history in roughly 7,400 lines of Remotion React/TS; Tony Dinh’s comparison (as transcribed by cosine) — $1,000+ a year ago, under 30 minutes now, brand assets scraped by the model itself, units unstated; Addy Osmani’s 40-second “how browsers work,” every frame drawn in JS; emollick re-cutting the same content as an anime short in one pass (last three via cosine); deedy’s 3b1b-style paper explainer on manim + Kokoro-82m + ffmpeg — “I didn’t specify any specific tools” (tweet).
cosine’s counter-intuitive observation: “the simpler the prompt, the more it invents — and the better the inventions.” Don’t misread it; the material has to be there first. An empty corpus plus a vague prompt equals free improv.
External footage routes
⚠️ Status: deedy’s full flow was not published as of 2026-09-30. In his 09-29 thread he wrote “I’m going to share my entire flow this week” (original) — “this week” means the week starting 09-29. This section rests on his published seven-step summary and the technical disclosures in his 09-24 tweets; everything beyond that is community inference. We re-checked the original thread on publish day and will update when it drops. Timeline: Opus 5.5 launch week 09-22; zhuermu’s post 09-26; cosine 09-27; baboonAI4S’s dissection 09-29; deedy’s thread 09-29.
Set expectations first: this section is a menu, not a recipe. deedy’s flow is unpublished, and the only runnable item — video-use — ships with results but no manual.
Live-action and archive. deedy’s “America: a 250 year tribute” is a five-minute documentary, “completely orchestrated by opus” in his words (dotey’s Chinese retelling circulated widely). The seven published steps: write the full script → gather CC-licensed footage from YouTube and archives → trim and crop → upscale and colorize → assemble with voice-over and captions → score the soundtrack → critique the edit and iterate. The one hard technical disclosure came earlier: the 09-24 promo, “no video gen model was used here, no additional libraries. Just JavaScript, playwright and ffmpeg” — and video generation appears nowhere in the documentary’s seven steps either. Which upscaler, which colorizer, how the footage was vetted: unsaid, so treat every thirdhand description on the market as inference.
Smaller verified cases exist in live-action editing. gregpr07 shot 13 takes and handed them to video-use, which read the captions, picked the best take, cut, color-graded and burned subtitles — two minutes, $0.90 (thread; he is Gregor Zunic, video-use’s maintainer — weigh that affiliation). AxtonLiu turned an 83-second live-action clip into line art, self-reported zero manual work, about 29 minutes (thread). And an example of honest bookkeeping: sab8a’s talking-head fine cut with Opus 5.5 plus OpenEdit took about 1h51m — $23 in tokens and $49 on fal, reported as two separate lines (thread). Bill external services separately in any hybrid pipeline and month-end holds no surprises.
Generated footage. あばねちゃん’s chain: GPT sol (her wording; product unspecified) generates three dance pose stills, Grok turns each into a 6-second dance clip, Opus edits the three into a 15-second showreel (thread). Her Japanese prompt, verbatim:
あなたがどれほど素晴らしいモーションデザイナーかを示す、ダイナミックな15秒のモーショングラフィックスビデオを作成してください。タイトルは「億万株姫☆あばねちゃん」です。履歴書用のショーリールのような感じで、全力でやって。
Opus as director. gavinpurcell wired it to Runway MCP so a video model shoots the scenes (thread); OriSilver drives Seedance from Blender blockouts (thread); koldo2k went full orchestration — Magnific, Seedream 5 Pro, GPT 2.5 paper-craft, Kling 2.5, Lyria 3 — for 1,500 credits on some service (which one, or what that is in dollars, the tweet doesn’t say; treat as expensive), with the agent required to quote before any costly step (thread). That last rule ports into any hybrid pipeline: if the bill isn’t yours to watch, the approval has to be.
Division of labor, per apiyi’s guide: humans and photoreal texture go to Sora/Veo-class models, Opus handles orchestration and packaging; for still photos, “invoke, don’t generate” — code plus a Ken Burns move and a mask is enough.
Six lessons that hold
Five come from zhuermu’s post-mortems; the sixth is style control:
1. Fact-check first, brief second. RSI v1 was gorgeous and every arXiv ID checked out — and it drew “prompting → context → harness → self-evolution” as a four-leg relay with every date too early: context engineering went mainstream in June 2025, labeled 2024; harness engineering dates to February 2026, labeled 2025. zhuermu’s line: “Every part was correct; the story assembled from them was tidier than history.” RSI v2 fixed it: 431 seconds, 20 chapters, 35 real arXiv IDs. Values v1 was worse — eight factual problems; his blog lists four (preprint numbers for journal versions, “28–62 percentage points” stated more absolutely than the paper, an unsourced claim about Western values, specs/constitutions drawn as a fourth training stage — the rest are on his blog). Values v2 fixed all eight. The only difference between v1 and v2: a fact-checked brief.
2. Say which reading you want. Given one line, it picks the safest interpretation. Earthrise replaying NASA’s timeline was a happy accident; pin the reading in the brief when the result matters.
3. Make it watch its own output. All five videos caught bugs by sampling frames. The contact sheet is standard procedure — require it before delivery. Eric Buess’s summary of the loop (via cosine): the model writes code, renders, looks at the frames it rendered, and fixes what’s off. The whole method’s reliability lives in that loop.
4. Don’t let long jobs die with the turn. In headless mode, the moment the agent’s reply ends, background render processes die with it. Two fixes: render in the foreground in segments, or state in the prompt that it must not stop until rendering is done.
5. Picture follows sound. Synthesize the narration, measure each line’s real duration, then lay out visuals. Stretching audio to fit the picture drifted more than ten seconds when zhuermu tried it.
6. Route style and rhythm around the model. When you can’t describe a style, show references — images or clips (i陆三金). Switching reasoning tier plus brush library switches look: the hand-drawn MV dissected by 小互 (an AI-interpretation site) came out Flash-style on Medium reasoning and watercolor-picture-book on xhigh plus p5.brush — note he changed two variables at once, so attribution is open. Rhythm is where the model most often fails (ClaudeAnimationBase guide): validate pacing on a 2–4 second sample before rendering the full piece.
Troubleshooting and templates
| Symptom | Fix |
|---|---|
| Overlapping captions, transparent titles, ffmpeg command overflows the phone screen | Frame-sample self-check (contact sheet) |
| Wrong dates or facts | Redo with a fact-checked brief (lesson 1) |
| Audio drifts out of sync | Voice first, measure durations, then lay out visuals |
| Background render dies mid-run | Foreground segmented renders; prompt says “don’t stop until rendering finishes” |
| Pacing feels off | Sample the hardest 2–4 seconds first |
| Wrong art style | Reference images/clips; swap brush library; swap reasoning tier |
Three template shapes on record. One-liner: the top viral hit in des13’s ranking (13.7m views, his tally) is one sentence — 「製作一支 2 分鐘的 [主題] 動畫影片,每一幀都用程式碼繪製」 (“make a 2-minute animated film on [topic], every frame drawn in code”) — plus the delivery note from route 1. Storyboard-grade: des13’s 15-second vertical with second-by-second storyboards, revised in natural language. Commission-grade: donaldjewkes’s MV brief runs 9,554 characters with a reference MP4, the song and source code, allows image generation, Seedance 2.5 and ElevenLabs, and kept the agent working 12 hours straight (thread) — the shape to imitate when you need fine control. We also keep a copyable motion-showreel prompt tested on GLM-5.3 in the same genre.
The final template is copy-ready; four modules map to patterns multiple creators converged on — goal, material boundaries, delivery spec, self-check license:
Make a 2-minute animated short on [topic], every frame drawn in code.
Material boundaries:
- Voice-over script in voiceover.txt, synthesized locally with edge-tts;
apart from that, no image/video/music generation services.
- Check every fact in the narration (dates, numbers, names) against the
material I provided; cut what you can't source, or rephrase to what can.
Delivery spec:
- 1080p 30fps MP4; all visuals rendered as HTML/Canvas; score synthesized in code.
- Synthesize voice-over first, measure each line's real duration, and lay
out visuals to the real timings; if stretching audio would exceed 10
seconds, re-layout the visuals instead of stretching.
Self-check and process:
- You may sample frames into contact sheets at any time.
- Deliver the storyboard and an animated sample of the hardest 2–4 seconds
first; I confirm before you render everything.
- Do not stop while rendering is incomplete; before a turn ends, write
progress and a resume plan into a handover doc.
Corpora and further reading
Three prompt corpora: athemeroy’s 88-case prompt/workflow matrix; zhuyansen’s 96 full-text prompts, playable at jasonzhu.ai; and X-RayLuan’s verbatim-prompt repo (the whole repo funnels to EasyVeo and its picks skew commercial — cherry-pick). We’ve previously dissected one gallery of Opus 5.5 prompts and a game-VFX prompt on this site.
The corpora carry water of their own: baboonAI4S counted 282 videos labeled not-made-by-Opus, 39 duplicate groups, and a rival video company — higgsfield — slipping in 22 posts of its own. Don’t take a hot repo’s numbers on faith. Prompt lengths span 172 characters to 17,000 — a hundredfold — and someone re-borrowed another’s 1,750-word prompt within 24 hours at 87% 5-gram overlap (all baboonAI4S’s counts). Read the source before the sentence.
Two last traps. The model cannot hear audio — to hit a beat, have it measure the music first: the P(doom) MV decoded the song in Python, computed per-frame energy, bass and a beat grid into a table, and made every visual read that table — “It could not hear the song, so it measured it.” And when deedy finally publishes his full documentary flow, the external-footage section above needs a rewrite — we checked on publish day (09-30); it hadn’t dropped yet.