Skip to main content

Chat or Decide: Which Kind of AI Your Office Actually Needs

Most of what a small business needs from AI is not a conversation. A model released this month makes that distinction concrete enough to argue about.

/10 min read

  • A chatbot writes you a paragraph. Most of the work inside a small business is not a paragraph, it is a small decision made dozens of times a week: is this junk or real, does this go to the art team or straight through, does this wait until Tuesday.
  • TypeSafe released a model called Jev on September 15, 2026 that does only the second thing. It cannot write at all. You give it a fixed list of possible answers and it returns one, plus a confidence number, in well under a second and for fractions of a cent.
  • What it genuinely fixes: it cannot hand your software something unreadable, and it is cheap enough that cost stops deciding whether a step is worth automating.
  • What it does not fix: it cannot tell you why it decided anything. No explanation, ever. That rules it out of most medical, legal and safety decisions on its own.
  • Do not buy anything yet. It is days old, behind a waitlist, and the vendor says the price may be subsidised. What changed today is the question to ask at the next sales call: am I being sold a conversation or a decision.
  • Where to start, free: count for one week how often your team sorts, routes, flags or prioritises something. That list tells you whether any of this is for you, before a vendor does.

A permit record lands in a folder overnight. Somebody has to decide whether it is real work or junk. An order comes off a tablet at a trade show. Somebody has to decide whether it goes straight to production or sits with the art team first. A maintenance request arrives at eleven at night. Somebody has to decide whether it waits until Tuesday.

None of those is a conversation. All of them happen dozens of times a week, and in a small business they are made by a person who is also doing four other things.

For three years the industry has been selling small offices a way to have a conversation.

The shape of the work

I wrote in August about the chat bubble in the corner of a website, and why most small offices who bought one actually needed an intake system. That note stopped at the front door. This one is about everything behind it, because the same mistake repeats further inside the business and costs more when it does.

A language model generates text. You ask it something, it writes back a paragraph. That is genuinely valuable when a paragraph is what you want: a first draft of an email, a summary of a long document, an explanation of something unfamiliar. Those are real uses and they are not going anywhere.

It is the wrong shape when what you need is for a thing to end up in the right place. Here is the tell. If anyone has ever built you a system that asks an AI a question and then writes additional code to dig the answer back out of the sentence the AI wrote, you have watched the wrong tool get used. That middle step, the digging, is where these systems quietly break. The model phrases its answer slightly differently one Tuesday, the digging code fails to find it, and something lands in the wrong pile without anyone being told.

What changed on September 15

A company called TypeSafe released a model called Jev, in early access behind a waitlist. I have joined the waitlist. I am not building anything on it, and I will come back to why.

Jev does not write. That is the entire point of it. It cannot produce a paragraph, a sentence, or an explanation, because it has no ability to generate language at all. You give it a situation and a question with a fixed list of possible answers, and it hands back one of them plus a number for how confident it is.

The plainest way to put it: the difference between asking a colleague to write you a memo about which pile a document belongs in, and having a colleague who simply puts it in the pile.

TypeSafe is calling the category System One models, borrowing the name for the fast automatic kind of thinking rather than the slow deliberate kind. The model is named after William Stanley Jevons, a nineteenth century economist. The company was founded by Diogo Almeida, who worked on the instruction-following research at OpenAI that became ChatGPT.

What it is genuinely good at

It is fast in a way that changes what is possible. Answers come back in 70 to 500 milliseconds, against seconds or minutes for a full language model on the same task. That is the difference between a decision you can make inside a process while somebody waits, and one that has to happen on a schedule overnight.

It is cheap enough to stop counting. Pricing is $0.042 per million tokens of input, roughly 750,000 words, with the answers themselves free. For the kind of small repeated sorting decision described above, that is fractions of a cent each. Cost stops being the thing that decides whether a step is worth automating.

It cannot hand back something unreadable. Because the possible answers are fixed before the question is asked, the model cannot invent a fifth category when you gave it four, and it cannot return something your software fails to parse. This removes an entire class of failure that has plagued every AI integration built in the last three years.

What it is bad at, which no vendor page will lead with

It cannot tell you why. You get a decision and a confidence number, and that is all you will ever get. A confidence number is not a reason. If you are in a business where a decision might one day have to be explained to a regulator, an auditor, an insurer, or an attorney, then "the model was eighty four percent confident" is not a defense, and this tool is the wrong tool for that decision. That rules it out of a great deal of medical, legal, and safety work immediately.

It can still be wrong. The phrase going around is that it cannot hallucinate, and that phrase is being repeated more loosely than it deserves. It cannot invent an option outside the list you gave it. It can absolutely choose the wrong option from inside that list. Two different claims, and only the first one is true.

The published numbers do not agree with each other. Accuracy figures being quoted for it by different write-ups differ by roughly eight points on the same benchmark, and that benchmark belongs to the vendor. When the people writing about a product cannot agree on how well it scored on the test its own maker designed, the honest reading is that nobody outside the company knows yet.

The price is probably not the real price. TypeSafe has said the pricing may be subsidized. Early access pricing is a marketing number. Building a process whose economics depend on it is a decision you may get to make twice.

Where this actually fits a small business

Not at the front door, where everyone keeps putting AI. In the middle, where the work already is.

A field services company receives records overnight, mixed in with a large volume of irrelevant ones, and somebody spends the first hour of every morning separating them. An apparel manufacturer takes orders on a tablet at a trade show, and every order has to be routed, standard straight through, custom artwork to the art team first. A property management office gets maintenance requests by email and by text and has to sort them by urgency before anyone gets dispatched.

In every one of those, the AI is not talking to anybody. It is standing inside a process that already exists and making one small call, quickly and cheaply, with a person reviewing the ones it is unsure about.

That last part is the design that matters, and it is the reason the confidence number is more interesting than the speed. You are not choosing between full automation and none. You set a threshold. Anything the model is sure about flows through, and anything it is uncertain about goes to a human, who now spends their morning on the twelve genuinely ambiguous cases instead of the four hundred obvious ones.

What I would tell a client this week

Nothing, yet. It is four days old, it is behind a waitlist, and the vendor has told you the price may not hold. I will spend time with it when access arrives, and if it turns out to be useful for a specific client bottleneck, that client will hear about it then, with numbers from the actual work rather than from a launch post.

What has already changed is the question you should ask at the next sales call. It is no longer whether the AI is any good. It is whether you are being sold a conversation or a decision, and which one your actual bottleneck needs. Those are two different purchases, they cost different amounts to get wrong, and a surprising number of vendors will not be able to tell you clearly which one they are selling.

Where to start

The same way as last time, by counting rather than by shopping.

For one week, have everyone write down each time they sort, route, flag, or prioritize something. Not the difficult judgment calls. The small repeated ones, where the answer is one of a handful of options and somebody experienced gets it right in about two seconds without thinking hard.

If that list is short by Friday, you do not have this problem, and anyone selling you a solution to it is selling you something you will not use. If it runs to thirty lines, you have just found the part of your business that this class of tool is genuinely for, and you found it yourself, which means the next person who tries to sell you a chat bubble will have a harder conversation.

Read and checked on September 19, 2026. Vendor claims are marked as such. Where published figures disagreed, that is said in the text rather than resolved silently.

Questions I get asked on this

Can I use this in my business right now?

No. Jev went into early access on September 15, 2026, behind a waitlist, and it is the first release of its kind. I joined the waitlist to evaluate it, not to build on it. Nobody should be rearchitecting a working process around a model that is four days old and priced at a level the vendor has said may be subsidized. What is useful today is the distinction it makes obvious, not the product.

What is the difference between a chatbot and a decision model?

A chatbot produces language. You ask it something and it writes back, and then either a person reads that writing or your software has to dig the answer back out of it. A decision model skips the writing entirely. You hand it a situation and a fixed set of possible answers, and it returns one of them plus a number for how confident it is. A chatbot is for when you want a paragraph. A decision model is for when you want the thing to go in the right pile.

Does it really never make things up?

That claim is circulating loosely and it is worth splitting in half. Because the possible answers are fixed before you ask, the model cannot invent a category that does not exist or hand back something your software cannot read. That part is true and it is genuinely useful. It can still pick the wrong answer from the list. Those are two different claims and only the first one holds.

Is this going to replace my staff?

It replaces a step, not a person. The steps worth handing over are the small repeated sorting calls that an experienced member of staff makes correctly in two seconds and resents making forty times a day. The judgment calls that actually need a human are the ones this class of tool is worst at, because it cannot explain its reasoning. In practice you route the uncertain cases to a person, which is why the confidence number matters more than the speed.

How do I know if my business has decisions worth automating?

Count for a week. Write down every time somebody sorts, routes, flags, or prioritizes something where the answer is one of a handful of options. Not the hard calls, the small repeated ones. If the list is short by Friday, you do not have this problem and nobody should sell you a solution to it. If it runs to thirty lines, you have found the part of your business this is for, and you found it before a vendor told you.

Run the count.

Spend a week writing down the small repeated calls your team makes, then send me the list. I will tell you which of them are worth handing to software and which ones should stay with a person. If the honest answer is that none of them are ready, that is what you will hear.

Prefer email? [email protected]