On September 15, @typesafeai AI released a model called Jev. 150,000 people joined the waitlist in a day. Vercel said it was the fastest first-day adoption they have ever seen.
It can't write a single sentence.
I've spent the last four days reading the docs, the launch post, the critics, and every project I could find. This is what I'd want someone to hand me on day one: what Jev is, why people care this much, and where the hype gets ahead of the facts.
We hired a novelist to do paperwork
Almost every AI model you've used is a writer. ChatGPT, Claude, Gemini. You give it a prompt and it produces words one at a time until it's done.
That's what you want when you need an email or a function. But most of what software does isn't writing. It's deciding.
Is this ticket about billing or a bug? Should the agent click this button or that one? Is this shell command safe to run? Does this search result actually answer the question?
A careful person makes each of these calls in two seconds. But for the last few years, the way we got software to make them was to ask a writer to act like a clerk. We told it "respond only in JSON." Then we parsed the JSON. Then we wrote retry logic for when it wasn't JSON. Then we waited three seconds and paid for tokens we threw away.
I've written that code more times than I want to admit. If you build with LLMs, you have too.
Now put that inside an agent. An agent is a loop, and every turn of the loop has small decisions in it. Which tool? Did that work? Are we done? Is this risky? Each one is another full LLM call.
Jev is built for the clerk's job and nothing else.

What Jev is
You send Jev two things.
The first is the state: whatever your program already has. A support ticket, an email, a JSON object, a log line, a web page, a game's memory. It can be messy.
The second is a set of questions, and you decide in advance what the answers are allowed to be. There are only three kinds.
Choice picks one option from a list you write, up to 255 of them. "Is this billing, technical, or sales?"
Score places the state on a scale you define. "How frustrated is this customer: calm, frustrated, or furious?"
Noul is TypeSafe's word for yes or no. But it doesn't say yes. It says 0.94. It gives you the probability that the statement is true.
Here's what that looks like in practice. The state is a customer message: "I've been trying to connect Stripe for 3 days and it keeps failing. I'm losing sales. Please help ASAP." You ask which team owns it, how frustrated the customer is, and whether it's urgent. You get back something like: technical (91% sure), frustrated, and urgent with a probability of 0.999. No paragraph. Three values your code can put straight into an if.

You can ask many questions in one request, and they all run at the same time. Responses come back in 70 to 500 milliseconds. It costs $0.042 per million input tokens, and output is free, because there's almost no output.
What it can't do matters just as much. Jev can't chat, write code, summarize, or explain its reasoning. If the answer you need is a sentence, you still need an LLM. If the answer is one of a few things you can name in advance, Jev is the tool that was missing.
The best one-line description I've found: it's an if statement that can read English.
The number matters more than the answer
This is the part most people skim past, and it's the part I think matters most.
Take a ticket that says "I got charged twice after the app crashed." Jev might pick billing. That tells you what won. It doesn't tell you how close the race was. With Jev, you can see that. If billing got 0.52 and technical got 0.46, billing technically won, but routing that ticket automatically would be reckless. It was a coin flip.

An LLM hides that. It says "billing" in the same confident voice whether it was sure or guessing. Jev puts the uncertainty in a number, and a number is something code can act on.
That turns uncertainty into part of your architecture. Very sure, act automatically. Somewhat sure, ask a stronger model. Not sure, send it to a person. You set those thresholds in code, where they can be reviewed and changed. A dashboard tag can live with 0.7. Deleting a customer's data should need a lot more.
How it's different under the hood
TypeSafe hasn't published a paper or the weights, but they've shared enough to understand the idea.
It doesn't generate text. An LLM writes word 17 after word 16. Jev reads the state once and answers every question in one parallel pass. That's why ten questions cost barely more time than one.
It's trained to be honest about how sure it is. TypeSafe calls the method RLCD, Reinforcement Learning for Calibrated Decisions. RLHF, the method behind ChatGPT, trained models to give answers people like. RLCD trains for probabilities that mean something: if Jev says 0.8 across a thousand answers, about 800 should be right. That's the central claim. Nobody outside the company has verified it yet.
You can't fine-tune it. Everyone uses the same model. You adapt it through the state you send and the criteria you write for each answer. It's very literal. It answers the question you typed, not the one you meant.
Who built it and why the names are strange
TypeSafe's founder is Diogo Almeida. At OpenAI he worked on RLHF and InstructGPT, the research that taught models to follow instructions and became the base of ChatGPT. So one of the people who built the chat era is now saying chat is the wrong shape for most automation.
His launch post starts with a question: models have been superhuman at chat for years, so where is all the automation? His answer is that chat models were built with a human on the other side. Software needs something it can depend on, the way it depends on a function.
TypeSafe spent two years in stealth and raised a $40M seed round led by DCVC.
The names are worth knowing because they explain the whole bet.
"System One" comes from Daniel Kahneman's Thinking, Fast and Slow. System 1 is fast, instinctive thinking. System 2 is slow and deliberate. LLMs are System 2 machines, and we've been renting them for System 1 work.
"Jev" is after William Stanley Jevons, a 19th century economist. He noticed that when steam engines became more efficient, Britain burned more coal, not less. Cheaper energy created more uses for energy. TypeSafe thinks intelligence works the same way. Make a judgment cost a fraction of a cent and a third of a second, and people will put judgments in thousands of places they'd never put ChatGPT.
Why it blew up
Part of it was distribution. The launch thread did tens of millions of views. Within 48 hours Jev was on Vercel's AI Gateway, OpenRouter, Cloudflare, and LangChain.
But the post that made it click for a lot of people was a small one. Mike Taylor at Every gave Jev 37 documents and asked 21 questions about each. That's 777 judgments. They came back in under 0.7 seconds and cost about a quarter of a cent.

That number changes how you think. Normally you decide which items are important enough to send to a model. At that price, you judge all of them.
The deeper reason is that everyone building agents already felt the problem. For two years we've been taping classifiers onto chat models. Jev gave that pain a name and a fix on the same day.
What people are building
In four days, builders have used it to drive browser agents (one demo ran a Zurich to London flight search in 7.1 seconds for under half a cent), triage email at scale, gate risky actions in coding agents, and run a trading bot that makes one decision per blockchain block, about every 300 milliseconds. LangChain already ships middleware that uses Jev to pick which model handles a request and to block dangerous tool calls before they run. Different projects, same program underneath: state and questions in, probabilities out, code acts when the number is high enough.
The most useful result so far
The project that taught me the most wasn't a flashy demo. It was a safety gate for AI agents.
The builder first asked Jev one broad question: "Is this dangerous?" It blocked 39% of perfectly legitimate requests.
Then they broke the danger into specific risks. Does it delete files outside the project? Does it send data to the internet? Does it spend money? One yes or no per risk, combined in code.
Blocked legitimate requests went to zero.

Same model. Better questions. That's the real skill with Jev, and it's the opposite of how we prompt LLMs. No personas, no long preambles. One judgment per question. Clear criteria for what each answer means, especially at the edges. And always an escape option like "other" or "not enough information," because if the right answer isn't on your list, Jev still has to pick something.
Where it fits
Jev doesn't replace your LLM. It sits around it.
The LLM still does the work that needs words or deep reasoning: planning, writing, coding, explaining. Jev handles the small judgments around that work.
In an agent, that looks like a loop. Jev decides whether a request is easy or hard, so the right model takes it. Before any tool runs, Jev checks whether the action is safe. After it runs, Jev checks whether the task is actually done or the agent needs another pass. The LLM does the thinking. Jev makes the small calls around it.

A two-second test
Here's how I'd decide if a problem is a Jev problem.
Can you list the possible answers ahead of time? Could a careful person decide in about two seconds? Does it happen often enough that speed or cost matters? If all three are yes, try Jev.
And a few cases where it's the wrong tool. If a plain if statement already works, keep it; code is faster and easier to test than any model. If the task needs math, counting, or dates, do that in code, because Jev is weak at all three. If the answer needs several hidden reasoning steps, split it into smaller questions or use a reasoning model. If you need to extract a value you can't list in advance, find the candidates first and let Jev choose among them.
What the hype leaves out
TypeSafe says Jev "can't hallucinate." What that means is it can't return anything outside your schema. It won't invent a department called "vibes." That's real, and it ends JSON parsing errors for good.
It can still pick the wrong option with 0.91 confidence. And a wrong answer that looks like clean data is more dangerous than a messy paragraph. Your code reads a paragraph with suspicion. It trusts a field.
The benchmarks need care too. The "up to 200x faster" and "400x cheaper" numbers come from four workflows TypeSafe's own team built. Jev's answers were scored against the average answers of other frontier models, not against labels checked by people. TypeSafe itself says those numbers are probably near the top of what you'll see. The real advantage survives the caveat. It's just smaller than the headline.
The rest is what you'd expect from something four days old. It's text only, strongest in English, closed weights, early access with a waitlist, and there's no independent calibration data yet.
None of that makes it bad. It makes it a tool with edges, which is what every good tool is.
How to start
The fastest way is the public playground at jevtypesafeai.com. No account. Paste in a real message from your work, write three questions, and watch the probabilities move. Five minutes there will teach you more than this article.
When you want to build, join the waitlist at typesafe.ai, or use it today through Vercel's AI Gateway, OpenRouter, or LangChain's langchain-typesafe package.
Don't rebuild anything around it yet. Pick one small, low-risk decision you currently send to an LLM. Run Jev next to it without letting it change anything. Compare the answers, check whether its 0.9s are actually right nine times out of ten, and only then let it own that one branch. If it earns that, give it the next one.
What I think
For a few years we treated every problem like a conversation. Some are. Most of what software does is make small decisions, over and over, and hope each one is right.
Jev isn't trying to be smarter than ChatGPT. It's trying to be the fast clerk we never had. It reads the file, answers the questions you drew, and tells you how sure it is.
I think the hype is mostly a lot of people realizing at once that they've been paying a novelist to do paperwork.
So what's the smallest decision in your product that still goes through an LLM?
