IT Buddy
← Back to blog
AI Implementation 5 min read

An AI that does not write. It only decides.

Uros Vujic 24. september 2026

The missing piece

In our previous post we wrote about the loop: that this autumn's AI breakthroughs did not come from smarter models, but from the scaffolding around them. The conclusion was that the hard part is not building the loop itself. The hard part is knowing when it is finished.

The Firefox project had an answer nobody could argue with: the program crashed, or it did not. Most business processes have nothing like that. Ask a language model whether an invoice is coded correctly and you get a paragraph back. Someone has to read it and interpret it.

The same month, a model arrived that goes straight at that problem.

What Jev is

Jev is made by TypeSafe AI, a San Francisco company founded by Diogo Almeida. According to TechCrunch, he is a former OpenAI researcher who helped build ChatGPT. The model was released in limited early access on 15 September.

What sets it apart from everything you have used is that it does not write text.

You decide in advance which answers are possible. For example: coded correctly, coded incorrectly, needs manual review. Jev picks one of them and attaches a probability. The answer is meant to be read by a program, not a person.

Typical language model Jev
Answers with Free text A predefined choice and a probability
Produces the answer Word by word All at once
Response time Seconds to minutes 70 to 500 milliseconds, according to TypeSafe
Built for People Software

TypeSafe calls these "System One models", after Daniel Kahneman's term for fast, intuitive thinking. It is not built for conversation or for writing anything creative. It is built to make many small decisions quickly.

Why it is interesting

Think back to the loop from the previous post. One of its four steps was that the loop has to decide whether the task is actually done. With an ordinary language model that step is weak, because the answer is text that has to be interpreted.

With Jev, that step becomes a number. "Coded incorrectly, 94 per cent sure" is something a program can act on without anyone reading anything.

Anthony Maio, who has written one of the most thorough reviews, thinks the model fits best as a control layer around other agents. It can pick which tool to use, judge whether a result is good enough, notice that a loop is going round in circles, and decide when a person needs to step in.

It is not a new chatbot. It is a new kind of building block.

The claim that does not hold

Several write-ups say Jev never hallucinates. That is not true, and TypeSafe more or less says so itself. The zero per cent in its own charts is, in its own words, not measured. It only means the model never returns an answer outside the format you defined.

It can still pick the wrong option. And it can do so with high confidence.

That difference matters. A language model that is wrong is often wrong in a way you can see. A program that gets back "coded correctly, 97 per cent" asks no questions.

What nobody knows yet

There is a lot we do not know about Jev, and it is worth saying plainly:

How it is built has not been published. The architecture, the training and how the probabilities are calibrated are all undisclosed.

Nobody has tested it systematically from the outside. TechCrunch spoke to two developers who had tried it. One got faster and more accurate answers than from the model he compared it with. The other found Gemini slightly more accurate, but 10 to 20 times more expensive.

TypeSafe's own evaluations were built by its own team. It says so itself, and also notes that its simplest demos show the model in a favourable light.

None of that makes Jev uninteresting. But it does mean this is a new category to understand, not a product to buy tomorrow.

The hard question: what is sure enough?

Here is the point that holds whichever vendor wins.

Armin Ronacher, CTO of Earendil, put it precisely to TechCrunch: if the answer comes back at 50 per cent, somebody has to decide whether that is really just a coin toss.

The vendor cannot answer that question. You have to.

Should everything below 80 per cent confidence go to a person? Below 95? Is the threshold the same for a 500 kroner invoice and a 500,000 kroner one? Who owns that threshold, and who changes it?

And the threshold is not something you set once. Maio points out that the probabilities can become less trustworthy when conditions change: new routines, new kinds of customers, new patterns the model has not seen. A threshold that was right in March can be wrong in September. Someone has to keep watching.

That is not a technical question. It is a question of accountability and control, and it is exactly the kind of question that has to be answered before the technology goes into use.

What you can do now

You do not need Jev to get started. You need to know which decisions in your business are really a choice between fixed options.

Many are. Is the invoice coded correctly. Does the enquiry belong with support or sales. Is the document current. Should the case be escalated.

For each one, you can write down three things:

Which answers are possible? Write them down. If the right answer is missing from the list, any model will put its probability on a wrong one.

What is correct, and how do you know? Without a reference answer, nobody can measure whether a model gets it right.

How sure do you need to be before nobody has to look at it? And who decides that?

That work is useful whether you end up with Jev, a language model or simply a better routine. It is the structure. The technology comes after.

UV

Uros Vujic

Daglig leder, IT Buddy AS

Uros hjelper norske virksomheter med å innføre AI på en kontrollert og bærekraftig måte. Bakgrunn fra IT-infrastruktur i bank og finans, med spesialisering i AI governance, RBAC og GDPR-compliant implementering.

Ready for the next step?

Take our AI Ready assessment and find out where your business stands.

Take AI Ready Assessment