Developers / JevModel

Cloudflare Clef and Clef-flash: a practical integration guide

Learn how Cloudflare Clef takes typed questions, where to call its API, and what to validate before using its answers in an agent or review workflow.

Last updated:

Quick answer: a provider-specific decision endpoint

  • Clef and Clef-flash are separate Cloudflare models with published weights.
  • Use the Workers AI model route or binding; JevModel credentials do not authorize that service.
  • Start with a text-only question, then validate media handling and the response envelope separately.
  • Examples here are documentation-based and were not executed against a paid endpoint.

What is Cloudflare Clef?

Clef is a decision model for questions with predefined outputs. The model cards describe a 27B Clef and a 9B Clef-flash, each mapping supplied evidence to option probabilities rather than writing a free-form response. That makes them candidates for queue selection, rubric scoring and bounded agent steps. It does not make them action executors. A returned queue still needs a program to route the ticket, and a proposed tool still needs permission to run. Keep that distinction in the product design: the model supplies a judgment; ordinary code decides whether the judgment is usable and whether the user is authorized. The comparison guide discusses when this path deserves a trial beside hosted Jev.

Clef vs Jev: choosing a decision model for your workflowStructured Decision Models for Autonomous Agents

Choose a serving path before writing the client

Cloudflare publishes Workers AI routes named @cf/cloudflare/clef and @cf/cloudflare/clef-flash. A Workers binding and an account-scoped REST call are two ways to reach that platform. Independently operating the weights is a separate deployment project. Decide which boundary you want before copying a request example: it determines credentials, region and operational responsibility. Keep account tokens in a server environment, never a public page bundle. Give the application only the permissions it needs and use a separate test credential when your team’s policy permits. A successful model lookup is not a successful production integration; you still need to inspect the response, timeout behavior and usage of a real request.

A minimal text request to Workers AI

The following is an original documentation-based REST example, not a JevModel request. Replace the account identifier through your server environment and provide your own authorized Cloudflare token. It asks only for a routing suggestion and includes a review outcome. Keep the state short enough to inspect manually while you verify the transport. First check HTTP status, then Cloudflare’s success/errors envelope, then the returned result and answer for queue. Do not hard-code an example probability or assume any successful JSON object contains an actionable answer. This command may incur provider charges if you execute it. We have not run it, and no output below is presented as a live result.

Choice, Score, and NoulJevModel API: request and response
curl --fail-with-body "https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef" \
  -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"model":"clef","state":"The export fails, and the customer also asks about a duplicate invoice.","questions":{"queue":{"type":"choice","instructions":"Choose the next review queue from the evidence.","criteria":{"billing":"Only a payment or invoice issue needs review.","technical":"Only a software error needs review.","review":"Multiple issues or insufficient evidence need human review."}}}}'

Write criteria that distinguish actions

A small answer set is easier to debug than an elaborate taxonomy. Begin with categories that imply distinct, reversible next steps. If billing and technical teams both handle the same cases, resolve that policy overlap before expecting a model to resolve it. Describe when each option applies rather than repeating the label. For an ordered score, explain the difference between neighboring levels with observable evidence. For a yes/no question, decide how missing information should be handled. Add that decision to your application policy rather than pretending a binary output can express every uncertainty. Keep the initial question version under review with the service owner; a label change can alter operational behavior even when no application code changes.

Choice, Score, and NoulRoute customer support with Jev

Treat media as a separate validation step

The current Workers AI parameter documentation describes an optional images array containing embedded PNG, JPEG or WebP data, with at most four images, per-image and total decoded-size limits, and no remote URLs. The model description also mentions video, but that headline should not be treated as a complete video-upload recipe for this REST request. Follow the serving path’s actual documented schema. Begin with a non-sensitive test image for which a reviewer can state the expected decision. Check file decoding, orientation and what happens when evidence is missing. If you transcribe an image into text first, measure that preprocessing stage too; an OCR error can explain a wrong decision that otherwise looks like a model failure.

Clef vs Jev: choosing a decision model for your workflow

Inspect answer distributions before automating

A selected option is not enough for a consequential route. Read the full returned answer and record how probabilities change on clear, ambiguous and out-of-scope cases. An even split, an unavailable result and a confidently wrong answer are three different conditions. Your program should preserve those distinctions. Choose thresholds on independently labeled examples, then review performance by category rather than averaging every case together. A good overall score can hide repeated errors on account-access requests or a minority language. Keep authentication and account recovery outside the classifier. For a support demo, route uncertain cases to review and show the reviewer the evidence and question version that produced the answer.

Jev probability, confidence and human-review thresholdsAdd a decision gate before tool use

Open weights require their own runtime verification

Use the actual model card’s decision-serving instructions when evaluating local Clef. Generic hosting widgets can display a chat-generation command even though the decision model needs its own loading and scoring path. Before downloading weights, record the model revision, runtime version, device, precision and expected memory budget. Verify a single request and its answer schema before testing concurrency. Include startup time, cold requests, model loading and process restarts in the operational assessment. If the deployment needs custom model code, review that code under your normal software policy. Published weights give you more control, but they do not supply monitoring, access control, backups or a maintained on-call service.

Is Jev open source?

Separate provider prices from service budgets

As checked on October 7, 2026, the Cloudflare model pages list input rates of $0.24 per million tokens for Clef and $0.09 for Clef-flash. Those rates are direct Workers AI model prices, not JevModel plans or a promise about the total application bill. Measure usage on your own payload and verify current pricing before committing a budget. Include transport, image preparation, retries and human review where they apply. Choose the flash variant only after its errors satisfy the same workflow requirements, not merely because its unit rate is lower. The useful outcome of a trial is a measured cost for an acceptable decision, not the cheapest advertised million tokens.

JevModel pricing and free Jev runsEvaluate a Jev workflow

Make the pilot produce a useful handoff

A practical first pilot should leave a saved question definition, a small independently labeled dataset, a list of failed cases and a response-contract note. Agree who owns each artifact: the service team owns category meanings, engineering owns transport and validation, and the workflow owner approves the risk threshold. Run in shadow mode before a model-selected label can change a customer record. Review disagreements and record which failures come from the input, taxonomy, parser or model. Expand to another task only when the first task has an accepted error budget and a rollback. You can practice question design in JevModel, but this website does not serve Clef, configure Cloudflare for you or convert its key into a Workers AI token.

Your first Jev decisionClef vs Jev: choosing a decision model for your workflowRoute an agent’s next step

JevModel is independent and not affiliated with TypeSafe AI.