What Is Jev AI and How Does It Make Decisions Without Text
Jev is a model from TypeSafe AI, built to read text and answer typed questions about it. TypeSafe calls it the first of a new category: System One models. Jev doesn't generate text. It won't draft a reply, explain its choice, or think out loud. Every answer comes back as a typed value with probabilities attached, which is what makes it usable as a step in your code.

The Problem: Text Generators Doing Decision Work
What your code needs from a model
Most decisions inside software are small. Is this ticket urgent? Which team owns it? Does this refund request match policy? Your code needs an answer it can branch on, plus a number telling it how far to trust that answer.
Why generated text gets in the way
A language model answers by writing, so even a one-word verdict arrives as a sentence someone has to parse. Structured output settings fix the shape of the reply, not whether it's right. And a confidence the model states in words isn't checked against how often it's actually correct.
Meet Jev and System One
What Jev is
Jev is a model from TypeSafe AI, built to read text and answer typed questions about it. TypeSafe calls it the first of a new category: System One models.
What "System One" means
The name borrows from Daniel Kahneman's split between fast, intuitive judgment and slow, deliberate reasoning. Think of the reflex that tells you a message is urgent before you've finished reading it. A System One model handles that fast kind: classifying, routing, ranking, verifying. Reasoning models cover the slow kind; Jev isn't built for it.
What it deliberately doesn't do
Jev doesn't generate text. It won't draft a reply, explain its choice, or think out loud. Every answer comes back as a typed value with probabilities attached, which is what makes it usable as a step in your code.
Anatomy of a Request
State: what you send in
The state is the material Jev judges: a customer message, an order record, a support thread, a block of JSON. Jev accepts text only, whether plain prose or JSON, and reads it as given. Questions are answered against it, so whatever a question depends on needs to be in there.
Questions: three types
Each question has a type, and the type decides what comes back. Between them they cover three shapes a decision usually takes: which one, how much, and whether.
A Choice asks which one of several options you define applies. You get a probability for every option, plus a confidence value. For example: is this ticket about billing, technical, or account?
A Score places something on ordered levels you define, such as low, medium, or high frustration. You get a score along those levels with the probabilities behind it, plus a confidence value.
A Noul, TypeSafe's name for the yes/no question, returns a single probability that the answer is yes. It has no separate confidence value; we'll see why when confidence comes up.
One call, many questions
You can send several questions about the same state in a single request. Jev evaluates them in parallel, each in isolation, so the answer to one can't influence another. You get back one set of typed answers, ready for your code to use.
Why Its Probabilities Are Meant to Be Trusted
Three ways to train a model after pretraining
After pretraining, what a model becomes depends on what it's rewarded for. Reinforcement learning from human feedback (RLHF) rewards answers people prefer, which is how raw models became conversational assistants. Reinforcement learning from verifiable rewards (RLVR) rewards answers that can be checked, such as a math result, and gave us reasoning models, which think longer and cost more. Jev takes a third path, RLCD: reinforcement learning for calibrated decisions. Its target is probabilities that match how often things actually happen.
What "calibrated" means
Weather forecasts are the familiar example. If a forecaster says 70% chance of rain on a hundred separate days, it should rain on about seventy of them. Calibrated probabilities work the same way: across all the cases where Jev says 0.8, roughly eight in ten should turn out yes.
What calibration does not promise
It describes crowds, not individuals. A 0.9 can still be wrong on a given case; calibration says that should happen about once in ten. TypeSafe's own documentation says as much: calibration holds across groups of predictions, not for any single answer. Treat calibration as the design target, and your own results as the proof.
Using Confidence to Decide What Happens Next
From probabilities to confidence
A Choice or Score returns a spread of probabilities, but a decision needs a single number to gate on. That number is confidence. For a Choice with n options, it measures how far the top probability sits above chance, scaled so 0 means a coin flip across all options and 1 means certainty: (top − 1/n) ÷ (1 − 1/n). With billing, technical, and account as options, a top probability of 0.9 gives a confidence of 0.85, and 0.5 gives 0.25. A Noul has no such value: with only two outcomes the formula would merely rescale the probability, so the probability already does the job.
Three bands: act, verify, hand off
Sort every result into one of three outcomes. High confidence: act automatically. Middle: verify first, using a rule in code, a second question, or a slower reasoning model. Low: hand the case to a person. For a Noul, read the probability directly: near 0 or 1 you can act, and the stretch around 0.5 is the uncertain zone.
Setting thresholds by the cost of being wrong
Where the bands begin is your call, and it should track what a mistake costs. Mis-tagging a newsletter is cheap, so a loose cutoff is fine. Approving a refund or locking an account is not, so the act band should start high. As an illustration only, you might act at a confidence of 0.6 for tagging and 0.9 for money. Start stricter than seems necessary, log the outcomes, and loosen a cutoff only when that log shows it holds up.
Putting It Together: A Refund-Request Walkthrough
Split one judgment into small questions
"hould we refund this customer?" sounds like one question, but it's three: did they ask for a refund, does the evidence point to a duplicate charge, and does the policy allow a refund in that case? TypeSafe's own example splits it this way, and its documentation advises against bundling several judgments into one question. The state holds what a human agent would read: the customer's message, the recent transactions, and the refund policy. Each question becomes a Noul.
Combine the answers with ordinary code
Check each probability against its own cutoff instead of multiplying them. The three questions describe the same case, so their errors can move together, and a product would imply more precision than you have. Exact checks stay in your code, such as whether the amount is within the refund limit or the charge falls inside the allowed window.
Route the case using probability bands
With cutoffs you've chosen (the numbers here are illustrative): all three above 0.9 and the amount within limit, approve automatically. Any answer in the uncertain middle, verify against the payment processor's record. Confident answers that disagree, such as a clear duplicate charge but a clear policy no, go to a person. A clear no that nothing contradicts gets the standard decline. Jev made three fast judgments; your code made the decision
Where Jev Struggles
Literal reading and multi-step indirection
Jev answers the question as written, not the one you meant. A vague question gets a vague answer, so phrase each one the way you'd brief a careful new colleague. Answers that need a chain of steps to reach are a known weak spot.
Numbers, counting, and dates
Arithmetic isn't what this model is for. Counting items, summing amounts, and comparing dates are unreliable, so do that work in code and give Jev the result. Instead of two timestamps, put "customer has waited 9 days" in the state.
Noisy input, adversarial text, and option order
Accuracy slips when a large state buries the relevant detail among irrelevant ones, so send only what the questions need. Text inside the state that is written to steer the answer can steer it, so treat untrusted input with care. With Choice questions, Jev also leans toward options listed first, so test your option order instead of assuming it's neutral.
Cost, Limits, and the Verdict
Price, limits, and access
TypeSafe lists input at $0.042 per million tokens, with no charge for output. Context is 64,000 tokens, with the state and your longest question limited to 32,000 of them. Through Cloudflare's catalog, the listed window is 32,000 tokens, so confirm limits for your access route.
Good fits and poor fits
Jev suits frequent, fast judgments with a defined set of answers: routing tickets, classifying content, ranking candidates, checking claims against a policy. It's the wrong tool when the output must be words, or the judgment needs real reasoning.
The verdict
Jev isn't a smarter chatbot; it's a different instrument. It gives code a fast judgment with probabilities to act on, and leaves writing and reasoning to other models. Its strongest claims, calibration included, come from TypeSafe's own documentation, so test it on your data before trusting it with money or customers.