Amazon open-sources Strands Decider 2B as decision models pile up
AWS Strands Labs open-sourced Strands Decider 2B on October 1: a free, locally-run decision model on a Qwen3.5-2B base — the third such release in three weeks.

AWS Strands Labs open-sourced Strands Decider 2B on October 1: a “decision model” running on a Qwen3.5-2B base, built for cheap, deterministic judgments inside agent workflows — free and runnable locally, authored by Amazon distinguished engineer Marc Brooker. The third such entry in three weeks.
The facts
- Positioning: a decider writes no essays and performs no reasoning theater — it emits structured verdicts (allow/block/defer) for the high-frequency small decisions in an agent loop.
- Timeline: TypeSafe’s Jev launched September 18 (named for the economist), OpenAI shipped a similar capability September 30, Amazon open-sourced on October 1.
- Cost logic: the comparison TechCrunch cites is $2.94 per monitored action for Jev versus $372 on a frontier model — a two-order-of-magnitude gap once decisions spill onto a dedicated small model.
- The splash of cold water: TypeSafe CEO Diogo Almeida says the current batch looks “more like ML people wanting to implement a cool architecture than a team deeply dedicated to making intelligence useful.”
- Open-source detail: Decider 2B is free and local; license and benchmarks live on strandsagents.com and Brooker’s blog.
The third layer of the agent stack
An agent stack now has three tiers: frontier models plan and write, decision models judge the flow, and sandbox plus orchestration enforce the boundaries — Google AX owns scheduling, OpenClaw Enterprise owns governance, deciders own the savings.
Editorial take
The Jevons paradox is the naming’s punchline: the cheaper each judgment gets, the more judgments agents dare to make, and total consumption rises. Agent builders should wire a decider in, but not on faith — Almeida’s criticism (no hard benchmarks) currently stands, so measure the misjudgment rate on your own workflow before replacing the rule engine.
What a decider actually replaces
Every production agent already runs a decision layer — it is just made of regexes, thresholds and an LLM call with “reply in JSON.” Deciders formalize that layer into a trained model: faster than a frontier call, more flexible than rules, and cheap enough (in Strands’ framing) to put between every tool call. The open question is error budget: a rule engine fails visibly, a 2B model fails confidently, and agent loops amplify confident errors.
Fit in the Strands line
Strands Agents is AWS’s agent SDK; a first-party decider completes the stack the same way it does for TypeSafe and OpenAI — the vendor that owns your harness wants to own your judgment layer too. Brooker’s post (September 28) frames it as infrastructure work, which fits his distributed-systems background; the Qwen base rather than an in-house model says Amazon is optimizing for cost and openness, not frontier capability.
The benchmark vacuum
Nobody shipping a decider has published rigorous misjudgment-rate numbers against their customers’ actual workflows — Almeida’s critique in full. Until someone does, adoption is an act of faith with a rollback plan. The measurable version of the pitch: take your agent’s last week of logs, replay every decision point through the decider, count disagreements against what shipped, and price the differences. That test costs an afternoon and decides the question better than any launch post.
The open-sourcing choice also sets up an interesting race: if Decider 2B’s weights become the default decider inside popular open harnesses, AWS gains ecosystem gravity without selling a single token — the same playbook that made Linux distributions a strategic asset. Watch which harnesses integrate it first and on whose terms.
Free weights, unproven judgment.
The decider layer is real; the evidence for any specific decider is not — yet.
Until then, the cheapest correct judgment is still the one your rules engine already makes.
That rule changes the day a decider ships with reproducible misjudgment benchmarks — the first one to publish them converts skepticism into standard-setting.