Compare / JevModel
Valen vs Jev in 2026: Which Decision Model Fits Your Agent?
Compare Valen and Jev for agent decisions: visual inputs, open weights, published Sokoban benchmarks, self-hosting steps and the limits of each approach.
Last updated:
TL;DR: Valen or Jev?
- Relationship: Valen is an independent, Jev-inspired multimodal model, not an official Jev release.
- Choose a hosted Jev path for text-based, typed decisions when you want to avoid operating model infrastructure.
- Consider Valen when a decision needs image or video evidence and you can operate and evaluate its open weights.
- Benchmark limit: the preview reports 87.60% on a selected 500-case Sokoban single-step set; it does not establish superiority over Jev.
Related through an idea, built as a separate model
Valen’s own README describes a Jev-inspired model that adds visual input to structured decision-making. TypeSafe’s Jev supplies typed answers to questions about a state; Valen applies a similar interface to text, images and video. Valen uses Qwen3.5 and its own trained decision head. Calling it an official Jev vision edition would misrepresent the project. For an agent developer, the useful question is whether the next decision needs pixel evidence and whether operating a separate model is justified. This guide is independent editorial research: JevModel is not affiliated with either project, and our current workbench does not serve Valen.
Structured Decision Models for Autonomous AgentsWhat changes when the state includes pixels?
A screenshot can contain evidence that a text-only record omits: which button is disabled, where an obstacle sits, or whether an object is visible. Valen scores supplied candidates against that evidence and the question instructions. The Qwen path reads hidden states through a shared decision head and does not generate answer tokens. Its reference implementation still runs backbone computation per branch: Choice and Noul use one branch per question, while Score uses one branch per level. A three-level Score plus one Choice and one Noul therefore requires five backbone forwards. Typed outputs constrain the answer space; they do not make a wrong observation or an incomplete candidate list correct.
A practical comparison
These are differences in the documented interfaces and operating choices, not an accuracy ranking. A Jev-style question shape does not establish SDK, preprocessing or response compatibility. Keep each integration separate until you have checked it against the actual endpoint or local implementation.
| Decision factor | Hosted Jev / this site | Valen |
|---|---|---|
| Relationship | Jev is TypeSafe’s model; this site provides a separate service. | Independent implementation inspired by Jev’s decision interface. |
| Input evidence | Our workbench accepts text / JSON state, without media uploads. | Project supports text, images and video. |
| Answers | Choice, Score and Noul through the documented API. | Choice, Noul and Score through the project’s decision interface. |
| Operations | Managed model access; request data goes to the providers. | Download matching checkpoint and base weights; own inference and maintenance. |
| Evidence for selection | Measure your task against the actual hosted model and workflow. | Published preview is Sokoban-trained; test your own visual domain. |
Watch a decision become an action
The project’s Sokoban replay shows a useful application pattern: inspect the board, select a direction, apply a move, then inspect the new board. In the displayed level, the authors report nine decisions and 1.13 seconds of cumulative decision latency. That is a selected demonstration, not a measured speedup over hosted Jev or a guarantee for browser automation. Open the animation below to inspect it; closing the disclosure hides the motion. The source identifies the thinking-model replay as accelerated, so the visual timing is not a stopwatch comparison.
Show / hide the project animation

Read the preview scores at their actual scope
The Preview model card reports 87.60% accuracy on 500 single-step Sokoban questions, and 38 completed games out of 100 with a 200-move cap. These measure different outcomes: selecting one action and finishing a sequence. The card also discloses that the 100 simple levels retained successful 2B RLCD cases before filling the set, conditioning evaluation on model outcomes. It is not an unbiased full benchmark. The General visual-question results in the family overview use separately trained checkpoints; they are not measurements of this Sokoban preview. None of these scores supplies a head-to-head Valen-versus-Jev result. A production decision needs a test set selected independently of either model’s successes.
Confidence is a signal to evaluate
Valen’s evaluation documentation defines confidence as probability concentration. For Choice it rescales the largest probability against a uniform baseline; Noul returns the probability of true without a separate confidence field. The blur demonstration shows concentration changing as visual information deteriorates. It does not establish calibrated correctness across other tasks. Record reliability on held-out examples before choosing an execution threshold. Add an explicit unknown or review candidate when appropriate, and let code enforce permissions even when confidence is high.
Show / hide the blur experiment

A screenshot workflow to pilot
Consider a hypothetical internal tool that routes screenshots of failed jobs to retry, inspect_logs or human_review. Supply a cropped screenshot and a precise description of each permitted action. Keep retry budgets, access checks and the final execution in ordinary code. A partial screenshot or an unfamiliar error should have a review path. Start in shadow mode: an operator chooses the action while the model’s choice is logged for comparison. The engineering owner reviews errors and handoffs; the security owner decides which screenshots may enter the pipeline. This is a proposed evaluation scenario, not a reported Valen deployment or a promise that the Sokoban preview understands your software UI.
Self-hosting: inspect before installing
The project’s reference setup specifies Linux, Python 3.10+ and an NVIDIA GPU. The Preview is a checkpoint plus config, not a standalone AutoModel repository: it requires matching Qwen3.5-2B base weights and the trained vision_top stage. Follow the model card’s prepare_model script to verify the base revision rather than mixing arbitrary downloads. The commands below are documented setup instructions; this article has not run them or downloaded model weights. A separate vLLM Jev project lists Valen serving on Linux and Apple Silicon, including experimental short-video input. Treat that as another runtime to validate, with its own installation and request rules.
git clone https://github.com/Liuziyu77/Valen.git Valen
cd Valen
bash scripts/setup/bootstrap.sh
source .venv/bin/activate
hf download Valen-Team/Valen-Preview-0923 --local-dir models/Valen-Preview-0923
python scripts/setup/prepare_model.py
python -m valen.inference \
--checkpoint models/Valen-Preview-0923 \
--data data/smoke/train.jsonl \
--output predictions.jsonlThe cost and license decision
The Valen repository releases its code under Apache 2.0; the Preview card also carries an Apache 2.0 license. Base models and source datasets retain their own terms, so review the whole dependency chain before redistribution or commercial training. Open weights remove a hosted-model dependency but leave GPU time, idle capacity, media preprocessing, updates and review work to the operator. Compare cost per successfully completed workflow, including failed actions and human review, rather than token price alone. Self-hosting can keep inference local only when the full media, logging and monitoring pipeline is configured that way; downloading open weights does not by itself establish a privacy guarantee.
Choose with an independent test
If text or JSON already contains the evidence for a bounded judgment, evaluate the hosted Jev path first. If original pixels are essential and your team can own deployment or task-specific training, Valen is a candidate for a controlled pilot. Freeze representative examples before tuning; include blurred inputs, missing evidence, misleading option names and cases where no action is suitable. Use separate calibration and test sets. Measure wrong-action rate, review rate, end-to-end p50/p95 latency and complete-workflow success, then assign an owner to review regressions after checkpoint or preprocessing changes. The deciding question is which errors and operating responsibilities your application can accept, not which selected replay looks fastest.
JevModel is independent and not affiliated with TypeSafe AI.