ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
Strands Decider 2B Goes Open Source: A 2B Model That Only Answers Choices, With a Confidence Score in Milliseconds

Strands Decider 2B Goes Open Source: A 2B Model That Only Answers Choices, With a Confidence Score in Milliseconds

AI information • Admin • • 5 views

Strands Decider 2B was released as open source by the Strands Agents team on October 7, 2026. It does not write essays or chat. It does one thing: pick one option from a given set and report how confident it is. With 2 billion parameters, local latency in the tens of milliseconds, and weights plus training data fully public, this class of small decision models is claiming a slice of agent workflows that used to belong to large models.

How it differs from an ordinary LLM

A regular LLM answers by generating text token by token, with no bounded answer space. A decision model flips that around: the candidate options are fixed first, a single parallel pass scores every option, and the output is just the choice and its scores. The trade-offs are explicit. It cannot generate text, it is markedly weaker at complex reasoning than reasoning models, and it is useless for coding, summarization or chat. What it buys in return is just as clear: the answer always lands inside the given options, it is fast, and every judgment carries a calibrated reliability score — an estimate of how likely the call is to be right — which frontier model APIs usually do not expose directly.

An architecture built by subtraction

Strands Decider 2B starts from a Qwen3.5-2B torso, removes the language modeling head that generates text, and replaces it with a pointer head of just over a million parameters. That head scores each option by comparing the internal state at the option's position against the state at an answer marker. The torso itself is fine-tuned with a rank-16 LoRA adapter. The team says this release is the 19th iteration, and the repository documents what changed in each version, including why an earlier slot-head design that performed clearly worse was abandoned.

Scores and speed

On the public JevBench set it ranks 3rd of 33 models in the 2B class, and 1st of 30 once models slightly above 2B are excluded; calibration quality, measured with the Brier score, is published alongside accuracy. Latency is a median of about 115 ms on a local RTX 3090 and about 153 ms for small tasks on an M3 MacBook, growing roughly linearly with task size. For workflows that need a judgment before every tool call, that order of latency can actually fit inside the call chain, instead of paying for a full large-model inference for each small decision.

What it actually does inside an agent

Listed uses include model routing, tool selection, evaluations, guardrails, memory, context management and policy classification. The bundled example shows the pattern well: an agent rushes to call a weather tool even though the user never named a city, and before the call executes, Decider answers two yes-no questions — are the argument values grounded in anything the user actually said, and is it too early to call this tool? A few lines of code then turn those answers into proceed, deny, ask a human, or hand the turn back to the model with feedback. The broader pattern is hybrid division of labor: leave the hardest judgments to a large model and give the many repetitive, clearly scoped ones to a decision model, cutting cost and latency together.

For teams building agents, the signal is that the division of labor is getting finer: not every judgment deserves a frontier model. Its boundary is equally clear — the options must be defined by a person or an upstream system, and Decider only chooses. If the options themselves are wrong, choosing accurately does not help. It installs with pip install strands-decider, the weights are on Hugging Face, and the training data and scripts are public, giving teams that want to train their own small decider a complete starting point.

Recommended Tools

More