Models

MiniMax ships M3.1-Flash preview: 1M context on a Flash budget

MiniMax launched M3.1-Flash-Preview on Sept 27: 1M context, five thinking tiers, native multimodal input — but only via its Token Plan, no benchmarks or price.

Orange and gold fireworks bursting in a night sky

On September 27, MiniMax quietly put a new model online — no launch event, no benchmark table, not even a pay-as-you-go price. The M3.1-Flash-Preview is reachable only through the company’s Token Plan subscription and its MiniMax Code product. During China’s Golden Week holiday the preview found its audience anyway: a full-day hands-on published October 3 by the WeChat account 小互AI argues the model is the first at Flash tier to take on the “make videos with code” workload that Anthropic’s Opus 5.5 turned into a genre, while burning just 2% of a weekly quota.

Key points

  • M3.1-Flash-Preview went live on September 27, 2026, available only via Token Plan and MiniMax Code — no per-token price, no official benchmark, no open weights.
  • The official guide lists a 1M-token context, five reasoning-effort tiers (low/medium/high/xhigh/max, default max, cannot be disabled), text/image/video input, and a MoE architecture with 428B total and roughly 23B active parameters at about 150 tokens per second.
  • Independent coverage confirms the release shape and the specs, but as of October 3 no independent benchmark exists — every “it can compete” claim traces back to experience reports.

Background

The lineage matters: M2 → M2.1 → M2.5 → M2.7 all stopped at 204,800-token contexts. This year’s M3 pushed the window to 1M with an official “frontier multimodal coding model” label and a pay-as-you-go price of $0.30 per million input and $1.20 per million output tokens. M3.1-Flash is the first “Flash” branding on this line — historically the fast, cheap tier.

The timing lands in the middle of the code-as-video wave Opus 5.5 set off, where front-end and motion work became the public showcase for model capability; we covered the trend and a practical playbook at the end of September (the wave, the playbook). On the Chinese side, GLM 5.3’s Flash tier already occupies the same ecological niche, which is exactly the comparison the community reached for first.

The facts

Verifiable specifications from the official guide (agent.minimax.io) and platform docs:

Item Spec
Model ID MiniMax-M3.1-Flash-Preview
Context window 1,000,000 tokens (128K output cap per test fixtures in the MiniMax Code repo)
Reasoning effort low / medium / high / xhigh / max, defaulting to max; cannot be disabled
Input/output Text, image and video input; text output
Architecture MoE, 428B total / ~23B active parameters, sparse attention, native vision encoder
Speed ~150 tokens per second (vendor figure, load-dependent)
Availability Token Plan and MiniMax Code only; same model ID on Anthropic-compatible and OpenAI-compatible endpoints
Caching Prompt caching supported

The gaps are equally factual: no benchmark published by October 3 and no entry in third-party model indexes; no pay-as-you-go price (M3’s $0.30/$1.20 is the nearest existing reference); no downloadable weights; and the Token Plan pricing page’s coverage footnote has not yet been updated to include M3.1-Flash, contradicting the model pages — an inconsistency two third parties flagged on launch day.

What others say

The most detailed Chinese-language account is 小互AI’s October 3 day-long review (all first-person claims): the model synthesized a soundtrack in pure Python for a 15-second motion-graphics showreel; a 140-second six-act fireworks show took about two hours; for a Lego assembly animation the model wrote its own OBB-plus-SAT geometry validator to eliminate floating pieces; a roamable 16-minute virtual world of roughly 2.48 square kilometers took four hours. The author’s verdict — best-in-class among Chinese Flash tiers, ahead of GLM 5.3 Flash — is experience, not a benchmark.

orcarouter, an AI routing service, called the release mechanism itself the story on launch day: the newest model bundled into a subscription rather than metered billing, with price, benchmarks and weights all blank, which suspends any horizontal comparison. The ai-on-mac fact-check confirmed the 512K/1M context and 128K output fixtures in the MiniMax Code repository while cautioning that “a test fixture is not the same thing as a public API guarantee”; same-day field reports from the TRAE community and Linux DO (as aggregated there) confirm the preview label appearing for some accounts.

Our take

What is genuinely new for developers is the tier strategy: a 1M context, native multimodality and top-tier reasoning effort under a Flash name, distributed by subscription. The benchmark gap is its own story — until one exists, the release mechanism is the entire information content. If the experience holds, the “flagship defines the ceiling, Flash does the work” cadence just accelerated again; for metered-billing users, the absence of a price is itself the gate.

Honest framing: every “competes with Opus 5.5” claim currently rests on experience reports — one account, no controlled tasks, no independent replication. Treat the 2% quota figure and the 2.48 km² world as upper-bound anecdotes. There is also a structural point: reasoning cannot be disabled and defaults to max, so a fixed share of every call’s cost goes to thinking — which explains why MiniMax dares to bundle it into a subscription, and hints that metered pricing, once it arrives, will not be cheap. What to watch: the official benchmark and pricing — precedent from the M2 line suggests both arrive when the preview graduates — and how MiniMax splits capacity between Token Plan and metered access, which decides whether this is a subscription perk or a product.

How to try it

Token Plan holders can pick the model inside MiniMax Code; the API model ID is MiniMax-M3.1-Flash-Preview on both Anthropic-compatible and OpenAI-compatible endpoints. Without a subscription there is currently no official path — wait for the pricing footnote to catch up.