Models#Code generation#Next.js#Evals#Anonymous model#Vercel
Pixel Canary ties GPT 6 Astra on Next.js coding evals — at 4× the runtime
Stealth model Pixel Canary went free on Vercel AI Gateway on September 25, tying GPT 6 Astra on Next.js coding evals — at nearly 17 minutes per task.

Another stealth model is being stress-tested in public, and this one brought receipts. On September 25, a model under the codename Pixel Canary appeared on Vercel AI Gateway as stealth/pixel-canary, free while it stays anonymous; the same day’s Next.js Agent Evals leaderboard puts it at a 90.3% baseline success rate, tying GPT 6 Astra (high) — at an average of 1,015.80 seconds per task, the slowest on the board.
The facts
- Availability and positioning: per Vercel’s changelog, the model is aimed at coding — building and refactoring applications — and is “well suited to frontend development and mobile app design,” from responsive layouts to navigation and interactive components. The developer is undisclosed, and the model page ships no detailed capability metadata. Per InfoQ’s report, Cline has also announced free access on its side.
- Eval results: Next.js Agent Evals (leaderboard updated September 25, 2026) scores pass@4 — a task passes if any of up to four attempts succeeds. Pixel Canary passes 28 of 31 tasks at baseline (90.3%), tying GPT 6 Astra (high); with Next.js docs bundled via AGENTS.md it passes 30 of 31 (96.8%), tying the top score in that setting. Passing tasks cover App Router migrations, data fetching, image and font optimization, caching and view transitions.
| Model | Agent | Avg time (s) | Baseline success | With AGENTS.md |
|---|---|---|---|---|
| Claude Opus 5.5 (high) | Claude Code | 269.62 | 97% | 97% |
| GPT 6 Sol (high) | Codex | 285.70 | 97% | 97% |
| Pixel Canary | OpenCode | 1015.80 | 90% | 97% |
| GPT 6 Astra (high) | Codex | 251.89 | 90% | 97% |
| Kimi K3 | OpenCode | 352.34 | 84% | 97% |
Numbers come from Vercel’s Next.js Agent Evals (September 25, 2026 leaderboard); success rate is pass@4.
- Speed is the catch: 1,015.80 seconds per task on average — about 4× GPT 6 Astra (251.89s) and 3× Kimi K3 (352.34s). Board-topping Claude Opus 5.5 finishes in 269.62 seconds, with a success rate 7 points higher.
- Specs and data policy: the model page lists a 262,144-token context window, 131,072 max output tokens and adjustable reasoning effort. The changelog states ZDR is not available for this model: prompts and responses may be used for training and model improvement.
- Community reports and identity guesses: per InfoQ’s roundup, users measured single-digit tokens per second and tasks running five to eight hours, some ending in provider errors. Identity guesses range from Qwen 3.8 (citing Thai and Hindi tokenization patterns) to Google (the Pixel and Canary naming) to Kimi K3.1 — none confirmed.
Our take
Stealth launches are load tests before an announcement — last week’s Space Bunny topped the usage charts on vibes alone. Pixel Canary is the first stealth model to bring hard eval numbers that place it in the top tier, so “worth trying” finally has evidence behind it. The slowness is priced in plainly: at 17 minutes a task, it fits long background jobs — migrations, refactors — not interactive coding.
Two things to weigh before pointing company code at it: without ZDR, your prompts may be kept as training data, and the free ride lasts only while the model stays anonymous. Once the identity and pricing land, compare them with GPT-6 Sol and Luna to see what the stealth period was really subsidizing.