Compare / JevModel
Strands Decider 2B: what it does and how to compare it with Jev
Understand AWS Strands Labs’ experimental open decision model, checkpoint versions, local serving and a practical Jev evaluation plan.
Last updated:
The short answer
- Strands Decider is an experimental Strands Labs decision-model project with Apache-2.0 code and published checkpoints.
- The launch used hobson-v19; the current repository documents hobson-v21. Pin a version rather than treating every “2B” result as the same model.
- It supports local inference; that does not establish a managed AWS endpoint or production service guarantee.
- Evaluate its decision quality and local operating cost against hosted Jev on the same labeled workload.
What “AWS Strands Decider 2B” actually refers to
The official Strands announcement introduces an experimental project from Strands Labs for bounded decisions. The 2B name identifies the base-model size family, not a hosted subscription or a promise that the complete runtime occupies two gigabytes. The repository separates a decision checkpoint from its base model and inference implementation. This matters when researching the popular AWS query: an AWS-associated project can still require you to download, host and operate the model yourself. Start with the official repository and the checkpoint’s model card. Record the license and deployment path before assuming that it is available through your existing cloud account.
Structured Decision Models for Autonomous AgentsIs Jev open source?Pin the checkpoint, base model and runtime together
The October 1 launch refers to hobson-v19, while the repository checked on October 7 uses StrandsAgents/strands-decider-2B-hobson-v21 in its examples. A comparison that mixes those revisions cannot explain whether an improvement came from the model or from the application. The documented loader obtains Qwen/Qwen3.5-2B-Base separately, using the revision specified by the checkpoint configuration. Budget for that base download as well as the decision files. Save the checkpoint identifier, base revision, package or Git revision, device and precision in the evaluation report. If you reproduce an older benchmark, use its older checkpoint deliberately and label the result accordingly.
Use the documented CLI for a first local question
The command below follows the current repository’s ask interface. It is a starting example, not a recorded result from our hardware. Install the runtime from the official instructions for your device first and verify that the checkpoint is available. The CLI can select CUDA, MPS or CPU; alternative backends and installation extras have their own release requirements. If an extra is documented only for a forthcoming package release, do not assume that an older published package already includes it. Begin with one small text case, inspect the full answer, and only then test realistic input sizes and concurrency.
Choice, Score, and NoulYour first Jev decisionstrands-decider ask StrandsAgents/strands-decider-2B-hobson-v21 --state "The customer requests a duplicate invoice correction." --choice "queue=billing,technical,review"A System One route is not sufficient proof of compatibility
The documented local server provides /v1/systemone and /health. JevBench can evaluate a TypeSafe-style adapter, but that benchmark integration does not verify every Jev client contract. Before swapping a URL, compare authentication, request fields, supported question types, response nesting, error status and missing-answer behavior. Write down which fields your application actually reads. The local server binds to loopback by default and does not provide authentication; keep it local for the pilot. A network-accessible service needs a separately reviewed access layer and operating plan. JevModel’s own API keys, quotas and envelope remain specific to this website.
JevModel API: request and responseJev vs OpenJevCompare decision errors before comparing token prices
Use a frozen set of independently labeled examples with the same allowed outcomes for Strands Decider and Jev. Include ordinary requests, mixed intent, empty evidence and unsupported languages relevant to your product. Keep tuning examples separate from the final report. Measure correct automated routes, wrong automated routes and human-review volume. A model that sends every difficult case to review may look accurate among accepted cases while accomplishing little automation. Report coverage alongside error rate. Inspect score rubrics separately from categorical choices, because a numerically close score can still cross a business threshold and trigger the wrong action.
Evaluate a Jev workflowJev probability, confidence and human-review thresholdsMeasure local operation at the concurrency you need
Published weights remove one dependency on a hosted model endpoint, but replace it with hardware and maintenance responsibilities. Measure startup, warm requests, memory, queueing and recovery after a process restart. Include idle capacity in cost estimates if a device must stay available all day for a small request volume. If you share a GPU with another workload, record that contention during the test. Compare the complete local service cost with a hosted Jev integration, including engineering effort and review labor. Do not infer production latency from one CLI response or claim an offline privacy guarantee while the first run still downloads assets.
JevModel pricing and free Jev runsClef vs Jev: choosing a decision model for your workflowChoose a bounded pilot and keep an exit path
Strands Decider is worth a local pilot when you want to inspect or adapt the model and can operate the runtime. Hosted Jev is worth testing when an existing decision API and fewer local serving responsibilities fit your project. Those are deployment preferences, not evidence that either model will classify your data better. For the first pilot, freeze one queue definition, preserve the existing route and run the candidate in shadow mode. Review disagreements, version every change and agree on a rollback criterion. This site currently serves Jev through its configured provider; reading this guide does not enable Strands Decider inside the Playground.
Route an agent’s next stepJev vs LayaJevModel is independent and not affiliated with TypeSafe AI.