Developers / JevModel

Jev probability, confidence and human-review thresholds

Understand Choice confidence, Noul probability and Score rubrics. Evaluate thresholds on labeled data before automating a decision.

Last updated:

Is confidence the same as probability?

No. In Jev, probabilities describe the distribution of a typed answer, while confidence summarizes its concentration. For a Choice with n labels and top probability p, TypeSafe documents confidence = (p - 1/n) / (1 - 1/n). An illustrative three-label answer with probabilities 0.90, 0.06 and 0.04 has confidence 0.85. Confidence is not an independently measured accuracy rate. JevModel passes through the upstream answer; validate actual fields before use.

How to classify text with the JevModel APIChoice, Score, and Noul

How do Noul and Score differ?

Noul returns the probability of yes in the noul field and has no separate upstream confidence field. A value near 0.5 suggests uncertainty about that binary question. Score evaluates ordered levels and can return a fractional value; its numeric answer is not a probability. Interpret a Score against its rubric, and inspect the returned distribution and confidence where available.

Score leads on a useful rubricChoice, Score, and Noul

What should make a classifier abstain?

Send ambiguous inputs, unsupported topics, malformed responses and service errors to a review queue. Add an explicit review label where appropriate, but do not rely on the model to always choose it. Gate automated routes using validated labels and a threshold tested for that question. Close leading probabilities can signal ambiguity. Keep deterministic permissions and human confirmation for sensitive actions even when the model is confident.

Route customer support with JevAdd a decision gate before tool useAI Decision Flows with Jev and human review

How do you choose a threshold?

Separate instruction tuning, threshold selection and a final held-out test. On labeled inputs, report errors among automatically accepted cases together with coverage: how many cases remain automated after review. Check each class and rare high-cost mistakes. Reliability diagrams and Brier scores can assess probability calibration; evaluate probabilities rather than treating confidence as a correctness probability. Recheck after changing labels, instructions or upstream model behavior. No universal threshold is guaranteed here.

Evaluate a Jev workflowBatch text classification with the JevModel API

JevModel is independent and not affiliated with TypeSafe AI.