AWS Releases Open Strands Decider 2B for Agent Decisions
AWS has open-sourced a 2-billion-parameter decision model that returns calibrated choices in roughly 115 milliseconds for agent workflows.
A model built to choose
AWS Strands Labs has released Strands Decider 2B, an open-weight model designed for bounded decisions rather than conversational text generation. Given a state and typed questions, it can select among options, answer yes-or-no prompts, or assign scores with probabilities. The weights, training data and training scripts are available under Apache 2.0, with local use supported on CPUs, consumer GPUs and Apple silicon.
The model is built from a Qwen3.5-2B torso whose language-model head has been replaced with a pointer head. That design removes the need to generate token sequences. Instead, the model compares hidden representations for the available answers and produces a structured result. AWS says the released version is the 19th major iteration of the architecture.
Why agents are the target
Strands Decider 2B is intended to sit inside an agent workflow as a low-latency control layer. It could decide whether an agent should call a tool, route a support request, classify an input, or escalate an uncertain case to a person. AWS reports a median local latency of about 115 milliseconds on an RTX 3090, while small tasks on an M3 MacBook averaged roughly 153 milliseconds. The model ranked third among 33 models in its 2B class on JevBench, and first when models slightly above two billion parameters were excluded.
The limitations are equally important: it cannot write arbitrary text, solve open-ended problems, or replace a general reasoning model. Its value depends on developers defining the decision space correctly and validating the confidence scores outside AWS’s own testing.
Why it matters
The release gives developers a permissively licensed, locally deployable component for making agent systems faster and more predictable. If decision models become a standard companion to generative models, agent architectures may shift from asking one large model to do everything toward using specialized models for generation, control and escalation.