Models#Code generation#Next.js#Evals#Anonymous model#Vercel

Pixel Canary ties GPT 6 Astra on Next.js coding evals — at 4× the runtime

Stealth model Pixel Canary went free on Vercel AI Gateway on September 25, tying GPT 6 Astra on Next.js coding evals — at nearly 17 minutes per task.

Black Vercel banner announcing free stealth availability of Pixel Canary on AI Gateway

Another stealth model is being stress-tested in public, and this one brought receipts. On September 25, a model under the codename Pixel Canary appeared on Vercel AI Gateway as stealth/pixel-canary, free while it stays anonymous; the same day’s Next.js Agent Evals leaderboard puts it at a 90.3% baseline success rate, tying GPT 6 Astra (high) — at an average of 1,015.80 seconds per task, the slowest on the board.

The facts

  • Availability and positioning: per Vercel’s changelog, the model is aimed at coding — building and refactoring applications — and is “well suited to frontend development and mobile app design,” from responsive layouts to navigation and interactive components. The developer is undisclosed, and the model page ships no detailed capability metadata. Per InfoQ’s report, Cline has also announced free access on its side.
  • Eval results: Next.js Agent Evals (leaderboard updated September 25, 2026) scores pass@4 — a task passes if any of up to four attempts succeeds. Pixel Canary passes 28 of 31 tasks at baseline (90.3%), tying GPT 6 Astra (high); with Next.js docs bundled via AGENTS.md it passes 30 of 31 (96.8%), tying the top score in that setting. Passing tasks cover App Router migrations, data fetching, image and font optimization, caching and view transitions.
Model Agent Avg time (s) Baseline success With AGENTS.md
Claude Opus 5.5 (high) Claude Code 269.62 97% 97%
GPT 6 Sol (high) Codex 285.70 97% 97%
Pixel Canary OpenCode 1015.80 90% 97%
GPT 6 Astra (high) Codex 251.89 90% 97%
Kimi K3 OpenCode 352.34 84% 97%

Numbers come from Vercel’s Next.js Agent Evals (September 25, 2026 leaderboard); success rate is pass@4.

  • Speed is the catch: 1,015.80 seconds per task on average — about 4× GPT 6 Astra (251.89s) and 3× Kimi K3 (352.34s). Board-topping Claude Opus 5.5 finishes in 269.62 seconds, with a success rate 7 points higher.
  • Specs and data policy: the model page lists a 262,144-token context window, 131,072 max output tokens and adjustable reasoning effort. The changelog states ZDR is not available for this model: prompts and responses may be used for training and model improvement.
  • Community reports and identity guesses: per InfoQ’s roundup, users measured single-digit tokens per second and tasks running five to eight hours, some ending in provider errors. Identity guesses range from Qwen 3.8 (citing Thai and Hindi tokenization patterns) to Google (the Pixel and Canary naming) to Kimi K3.1 — none confirmed.

Our take

Stealth launches are load tests before an announcement — last week’s Space Bunny topped the usage charts on vibes alone. Pixel Canary is the first stealth model to bring hard eval numbers that place it in the top tier, so “worth trying” finally has evidence behind it. The slowness is priced in plainly: at 17 minutes a task, it fits long background jobs — migrations, refactors — not interactive coding.

Two things to weigh before pointing company code at it: without ZDR, your prompts may be kept as training data, and the free ride lasts only while the model stays anonymous. Once the identity and pricing land, compare them with GPT-6 Sol and Luna to see what the stealth period was really subsidizing.