Productivity#Infographic

answer-me-with-html: agent answers become one page you can actually read

An MIT-licensed agent skill by QingYunA: the model writes only a Markdown draft and the bundled CLI builds the page, cutting output tokens to about 1/8.

Skill details

Install
npx skills add QingYunA/answer-me-with-html --skill answer-me-with-html

A TCP three-way-handshake explainer page generated by answer-me-with-html: sequence diagram, thread-model table and state timeline in panels

Ask your agent to explain how something works and you get six paragraphs you have to read twice; ask it to draw the picture and you watch it hand-type SVG coordinates for half a minute. answer-me-with-html, open-sourced by developer QingYunA on October 2, 2026 (MIT), splits the job: the model writes only a short Markdown draft, and a CLI bundled inside the skill does the layout, colors and diagram geometry. The result is a single HTML file with no CDN references, readable offline. The WeChat account 开源日记 featured it on October 7; as of October 8, 2026 the six-day-old repo has 2,038 stars.

answer-me-with-html comparison: six paragraphs of terminal text on the left, the rendered one-page HTML panel layout on the right

The split it imposes

The SKILL.md opens with a rule: do not hand-write HTML, CSS or SVG. The model’s job narrows to two things — decide whether the question deserves a page at all (≥ 3 interrelated concepts, a flow or sequence, or a comparison across ≥ 3 dimensions), then write the content draft. Template choice, panel layout and coordinate math all belong to am, a single-file CLI with no dependencies beyond Node.js 20+. Why split it this way? The author counted the tokens in 9 pages the model wrote by hand — the actual content is only 21% of what it types:

Part of the pageToken shareWith this skill
SVG diagrams (coordinates and paths)47%CLI writes it
CSS15%CLI writes it
HTML tags17%CLI writes it
Text21%Model writes it, as Markdown

On the same questions the measured medians (3 topics × 3 runs, Claude Sonnet 5.5, October 7, 2026 benchmark) drop from 4,893 output tokens to 612 — about 8× — and from 31 seconds to 12. Components follow the shape of the information: flow for architecture and branches, sequence for messages between actors, tree for hierarchies, timeline for history, plain Markdown tables for comparisons where ok/no/warn in a cell renders as ✓ ✗ ! badges. Coordinates are computed, so arrows never point at empty space and labels never clip. Three themes ship (blueprint, shadcn, paper) with light and dark modes, and the draft language follows the question — English, Simplified and Traditional Chinese, Japanese.

Controlled writing rules

The prose inside the panels gets its own constraints, modeled on ASD-STE100 — the controlled English originally written for aircraft maintenance manuals: one sentence says one thing, active voice, one word one meaning, and steps capped at 20 words in English or 35 characters in Chinese. The CLI checks on every render and flags long sentences, passive voice and filler; the Chinese rules build on the open Simplified Technical Chinese list and even catch light verbs (进行优化 → 优化), typos (登陆 → 登录) and vague quantities (尽快, 若干). Warnings advise by default; style: strict refuses to render a failing draft, and style: off disables the check.

STE controlled-writing check: overlong sentences and passive voice flagged line by line, with suggested fixes

Explainer videos in one line

am video is the video half of answer-me-with-html: say “make a 3b1b-style video of the TCP handshake” and the agent writes a script whose every narration line maps to one frame; the am video command turns it into a player page where diagrams appear beat by beat and the camera follows the narration. The official comparison (one TCP topic, measured October 5, 2026): a hand-written video page costs 27,839 output tokens and 202 seconds, am video costs 1,566 tokens and 17 seconds — about 17.8× fewer. Voice-over prefers ElevenLabs and falls back to system speech; --mp4 exports 1080p and needs Chrome, ffmpeg and Node.js 22+.

An explainer-video player page generated by am video: scene diagrams synced with narration captions

Who it’s for

answer-me-with-html suits anyone who regularly asks an agent to explain architectures, compare options or investigate a failure, this pays off immediately: ask how Redis and Memcached compare and you get a comparison table plus a one-line conclusion instead of five paragraphs of maybe. It shares its core idea with image-blaster — the model writes content and declarations, a program does the heavy lifting — and with threejs-architecture-effects, which takes 3D coordinates out of the model’s hands. Know these before installing:

  • The bill drops far less than the token count: the official benchmark shows about 15% savings, because every turn still reads the system prompt, your question and the conversation; the skill instructions themselves are about 4,000 tokens (cached, so later pages in a session don’t pay it again).
  • One-line answers stay text: the SKILL.md explicitly skips small talk and trivial questions — in a large context a page costs more, and the README states these boundaries plainly.
  • The repo is six days old: created October 2, 2026, 2,038 stars already, fast iteration — expect behavior to keep changing.

You need Node.js 20+ and there is no npm install step. Claude Code installs it from the plugin marketplace; every other agent takes one npx skills add QingYunA/answer-me-with-html (the full command is in the facts panel; the vercel-labs/skills installer supports 70+ agents). After installing, the README recommends adding the always-on rule to CLAUDE.md or AGENTS.md so every conclusion comes with a page.