ffmpeg-skill: makes your agent probe first and re-encode last
kajisho5's ffmpeg-skill makes an agent edit video and audio with local FFmpeg in a fixed order: probe, stay lossless, verify. 42 scripts, Python stdlib only.
Skill details
npx skills add kajisho5/ffmpeg-skill --skill ffmpeg-skillAn agent that “knows FFmpeg” still guesses: it assumes a frame rate, re-encodes a file that only needed a stream copy, and reports “done” without opening the result. ffmpeg-skill is kajisho5’s open-source repository (MIT, about 1.9k stars as of 2026-10-08, latest release v2.5.1 on October 5, 2026). It pairs one SKILL.md with 42 Python scripts that turn editing audio and video into a fixed workflow, calling ffmpeg and ffprobe locally with no cloud service and no API key.
The order it works in
The SKILL.md gives every job a fixed sequence whose spine is measure, edit, check:
- Probe first. Run
probe.pyon each input you plan from and read duration, frame rate, resolution, codecs and channels, instead of assuming from the file name. - Stay lossless. If the request can be met without re-encoding, don’t.
cut.pyandloudness.pycopy the video stream by default. - Dry-run, then execute. Every writing script takes
--dry-run --json, so the plan comes before the encode. - Chain in a sensible order. Colour, cut, join, silence, fit, captions and overlays, sync, audio, loudness, export. Three or more steps go into one
project.jsonthatrender.pyruns. - Check the deliverable.
check.py --platformtests the file against the destination, andlook.pymakes a contact sheet so the agent actually looks at the picture.

The clip above is one of the README demos: input on the left, the output of fit.py --aspect 9:16 --fit crop on the right. The project says every demo is generated from synthetic footage and can be rebuilt with python3 demos/build.py.
What it constrains
- No raw ffmpeg commands. Each operation is a script with typed arguments, nothing goes through a shell, and no filter graph is accepted from the caller. If a request needs something none of the 42 scripts exposes, the agent has to say so rather than improvise.
- Originals stay put. An existing output path is refused;
--overwriteis the only way to say yes, replace it. - Reports carry numbers. Every job ends in five labelled lines (
Done,Steps,Check,Look,Notes) with duration, resolution and frame rate taken from a probe. A non-zero exit or an empty file is reported as a failure. - Judgement stays with a person. The skill executes an edit and doesn’t decide one.
scenes.py --highlightsranks by loudness or duration and gives candidates only, and caption text is burned as written, never shortened to fit. - Changed pictures must be looked at. After captions, overlays, crops or colour conversion, the contact sheet has to be inspected; an agent with no image view must write that pixels weren’t inspected.
The 42 scripts cover cutting and joining, silence and filler-word removal, reframing to 9:16 or square, captions with karaoke highlight, LUTs and HDR to SDR, multicam and external-mic sync, loudness, platform exports and compliance checks. Templates ship for TikTok, Reels, Shorts, YouTube, X, LinkedIn, Facebook and podcasts, so render.py --template tiktok clip.mp4 is one command.
Who it’s for
It suits anyone asking Claude Code, Cursor or Codex to work through local footage: cutting silence from a talking-head recording, turning landscape into vertical, normalising a folder of files, exporting to platform specs. It edits existing material. For building animation from code, look at hyperframes or remotion-best-practices, which render from HTML or React.
Know the catches. It needs FFmpeg 5.0+ and Python 3.9+, and on macOS the README asks for brew install ffmpeg-full, because the plain ffmpeg formula lacks the subtitles, drawtext and zscale filters and the caption scripts won’t run. Automatic transcription needs a local whisper; without one you supply the text or a timed file. It won’t judge taste either: highlight picks, framing and grading go back to the calling agent or you. And the SKILL.md is about 30 KB that takes up context once loaded, which is the price of the steady workflow.