“Civilization advances by extending the number of important operations which we can perform without thinking about them.”

Alfred North Whitehead, An Introduction to Mathematics, 1911

Most job ads that reach a person in Kitsuno have travelled a long road. A crawler finds them. A model reads each one and pulls out the facts. Another model scores it against the person’s own record. If it fits, a third model can later write a CV and a letter for it, and Kitso, our concierge, can talk about it with the person.

Most ads never reach the person, and that is the point of the road. But when I looked at the numbers this month, the order of the road was wrong. In the 30 days before we changed it, 78 percent of the ads our extractor had read were thrown away later, and so were about 70 percent of the ads our scorer had scored. We spent language-model calls to read, judge and describe jobs, and many of them could have been stopped at the door by a simple question.

For a while that was affordable. The extractor and the scorer run on free or sponsored models. That time is ending. This week only Cloudflare and Hetzner were still truly free for us, Mistral was capped at about €8.50 a month, and our promotional credit at OpenRouter was almost gone. Every call we do not need eats into a budget that all our free users share.

From the first day, I built Kitsuno on four words: deterministic by default, AI by exception. On 20 September I started to look at a model that made me add a third term.

A model that only decides

Jev comes from TypeSafe, a company that came out of about two years of stealth on 15 September. Jev reads text and writes none. You send it a state (the text to judge) and your questions, and it returns probabilities. A Noul gives the probability that a statement is true. A Choice gives a probability for each option you define. A Score places the state on ordered levels that you describe in words.

It costs USD 0.042 per million input tokens, and output is free. In production, a call takes about 0.6 seconds for us. It is English first, its weights are closed, and the company sits in the United States.

The next day, Nate B. Jones published Jev and the Missing Primitive. He calls this kind of work semi-determinism: the possible outcomes are known in advance, but choosing one takes judgment. Our code was full of it, and we had been asking writing models to do it. (Nate’s writing has sent me somewhere before: an article of his in February was the trigger for Kitsuno itself.)

We looked for the places, and found old bugs

Our first step was a survey of the main places where Kitsuno calls a language model, with the real call volume of each route and the data most of them send. Nate published a companion guide the same day, Find the Jev-shaped problems in your software. When we read it two days later, through the library connector he offers, it gave us a sharper test and a few more places to look. The test has four parts, and all of them must be true: the possible answers are known, the volume is high, code acts on the answer, and there is a place to send the cases where the model is unsure.

We could probably have done this analysis without Nate’s guide, but I like that our work links to his. Jev was the real trigger: it forced us to walk every flow again, step by step. While we walked and measured, we found defects that had cost us calls and quality for weeks or months:

  • The extractor copied the whole job ad back into its answer and wrote the copy over the original text. Almost 4 percent of its calls hit the token limit, so their answers were cut off.
  • Our country resolver read “Ruhstorf, BY, DE” as Belarus and Germany. BY is Bavaria.
  • The employment filter trusted the extractor. The extractor set a working time where the ad named none, and our own extraction rule made it read “80 to 100 percent” as part time. About 40 percent of a sample of the filter’s rejections were wrong.
  • The location from the crawl never reached the job record before the extraction. Every location check at that point had worked without it since the first commit in March.
  • A table that a model had filled in June mapped “remote” to the United States. For three months, remote jobs were rejected for people who exclude the US or have no right to work there.

None of these is a Jev bug. They surfaced because we had to ask, at every step, which decision is made here, on which data, and by whom. If you want to try this on your own product, start where we started: list every place where you call a language model, with its real volume, and write down the decision each call makes. Then measure how much of the output you throw away. That list told us more than any benchmark.

Where Jev sits on the road

We made our first real Jev calls on the evening of 22 September, and the rebuild took about two and a half days. At the end of it, Jev makes decisions in six places.

  1. At the door. Before any language model reads an ad, one Jev call asks which languages the job requires, what the work mode and the employment type are, and whether the job matches one of the person’s target roles. Code does the rest of the gate: duplicates, budgets, countries, distances.
  2. After the extraction. Jev judges how well the person’s record matches the role, from “none” to “all”, and how close the domain is. Code turns that into the fit score. A language model writes the explanation only for the jobs that stay.
  3. In Be Found, where job ads find the person through our open agent-to-agent protocol. A code filter on the role family runs first, then a Jev role check, and only then a language model.
  4. In front of Kitso, the AI concierge that people chat with inside Kitsuno. Jev sorts every message into one of seven kinds of reply. Four of them (status, action, product and problem) are answered by code, from the live state of the account and our help texts. Only three go to a language model.
  5. Behind Kitso. Jev checks every sentence that the language model wrote: is it a claim, and do the facts support it?
  6. In the writer. Jev checks the claims in every CV line and letter sentence against the person’s own records.

The pattern has a name: a cascade. Cheap stages go first and reject what they can. Each stage passes on only what it cannot rule out, so the most expensive tool sees only what survives. Face detection has worked this way since Paul Viola and Michael Jones in 2001, and FrugalGPT applied the idea to language models in 2023 as an “LLM cascade”. Our cascade runs along the road. Code decides what it can. Jev decides the closed questions, as early as the data allows. Code does the arithmetic. A language model writes only what a person will read, and only for what survived. At the end, Jev checks the claims in that text. The chat and the writer follow the cascade from end to end today. In the job pipeline and in Be Found, a language model still does part of the deciding.

What changed in the calls

For crawled jobs, we compared the week before the first gate with the first day on which all the checks at the door ran. Language-model calls per 100 new jobs went from 150 to 39, about three quarters fewer. (That count leaves out a separate classifier for social posts, which does not pass through these gates.) An earlier check of a single night crawl gave 175 and 38. Across all our agents, language-model calls fell by about 70 percent a day. We now make more Jev calls a day than language-model calls.

TypeSafe named the model after William Stanley Jevons, who observed that more efficient use of coal led to more coal being burned. I expect a version of that here. Once a judgment costs a fraction of a cent, you start to find more places to ask for one.

In Be Found, 63 percent of the language-model checks had gone to pairs where the role did not fit, and only about 4 to 5 percent ended as strong matches. With the code filter, the Jev check and a rule that free plans match only full ads, strong matches per 100 language-model calls went from 4.9 to 15.7.

There, a Jev check costs about USD 0.00006, and the language-model check it replaces about USD 0.0001. Our logged production Jev calls have cost cents so far, and each of our larger test runs cost less than a dollar. The money matters less than the free quota: every call we do not make is one that the free providers do not have to carry.

There is a price, and I want to name it. Fewer jobs reach the pipeline now: about half as many per new job on that first day, although fewer profiles were active, so the mix differs. In Be Found we planned for about a third fewer strong matches, mostly on the short ad previews that some job boards send. The first day came out a little above that. For the job pipeline, the early evidence says that most of what we lose is noise. Before the gates, users acted on about 0.1 percent of the visible jobs with a very low role score, and on 1.6 percent of the others. In a backtest on my own profile, the role check, set more cautiously than it runs now, removed seven visible jobs that the old scorer had rated between 74 and 94. I looked at all seven, and all seven were noise to me. The live setting removes a few more that I have not reviewed. That is one person’s review, not proof, and for Be Found we do not know yet. We will watch both for weeks.

The numbers taught us one more thing: savings do not add up. On 23 September we built a cache, so that two people who get the same ad share one extraction. It has saved almost nothing, because the gates in front of it now stop most jobs before the extraction, repeats included.

What got better

Toward zero false claims

Our tenth principle promises that if we say the AI does not invent your credentials, we have a validator that catches it. Jev put numbers behind that promise. It asks two questions about each line: is this a claim about the person, and do the person’s records support it? On a fixed set of real job ads, measured with the same two questions, unsupported claims per draft went from 2.60 to 0.07, and the few that were left were borderline. The harder test was a slow one. We checked 30 lines where Jev and our old LLM judge disagreed, one by one against my records. The old judge was wrong on 20 of them. Jev’s two-question design agreed with the records on 29 of 30.

The same analysis found why the writer had invented. A prompt change in August had told it to write the document backwards from the job’s requirements. So the job ad’s own project came back as the person’s past work.

When the truth made the letters worse

Then I drafted real letters for real jobs, and they got worse. First, a routing bug let the letters ignore my writer settings. When we fixed that, the truth check flagged exactly the sentences that the settings ask for: why this job, what I learned, a confident close. Each flag became a repair, a cut or a warning. My words at the time were blunt: the letters got worse after the CV got truthful. The letters lost their voice, and the warning panel turned into wallpaper.

Underneath was an uncomfortable finding. The older letters read well partly because they invented bridges. Four phrases from one August letter had no match anywhere in my records.

The fix was a distinction, not a stronger model. Every sentence of a letter belongs to one of three classes. A record claim is a fact of my past: a role, a result, a number, a name. Jev checks it against my records. A stance is a view, a value, a motivation or a lesson. It is mine, and nothing checks it. An employer line speaks about the company, and Jev checks it against the job ad. A record claim that fails is held and shown to me. It is never cut and never rewritten. The CV summary now follows the same rules for claims and stances. I also asked for a system without repair calls, because the repairs made the letters more expensive, slower and more awkward.

At the end I get a short report: each line that the checks flag, with the reason. I decide whether I keep the line as it is or rewrite it. Today I cannot trust any AI one hundred percent. That is why every user can edit the CV and the letter in Kitsuno before they download the PDF or a Word file. The checks propose, and the person decides.

On our calibration set, the classes separated cleanly. On Jev’s “is this a record claim?” question, stances scored between 0.03 and 0.18, and record claims between 0.58 and 0.99, so the threshold sits in a real gap. One edge is still open: a sentence that names skills or tools can pass as a stance, and then nothing checks it. A letter now takes two writing calls and two or three Jev calls, where the version before took three to five and three to seven. Counted under the final rules, six of the last eight test letters needed no hold at all.

The chat that told a story

In parallel we had an incident. A new user asked Kitso for a scan. A fixed reply said that the scan had started, before any check ran. The server refused the scan. The language model then continued the story it found in the history and reported progress on a scan that never ran.

We took the incident as an urgent use case. Jev now sorts every message first. Four of the seven kinds of reply come from code, live data and our help texts, so no model writes them. For the other three, a Jev gate checks every sentence of the reply against the facts. A sentence that fails gets one rewrite, and if it fails again, it goes. On 620 test messages, most of them replayed real chats with identity data removed, the router picked the right kind of reply 93.5 percent of the time when it answered. Our bar was 95 percent. German lagged at 81.5 percent, mostly on short answers like “ja” to offers from our old chat, and on questions where even the right label was debatable. The gate catches sentences like “your pipeline is currently empty” when the pipeline is full, and it adds about a second to a reply. The quality of the replies is far better now.

A score that stays put

Our scorer once gave the same job 58 in one run and 90 in the next, with the same input. We then tested eleven job ads near the threshold, nine runs each. The median spread was 33 points, and six of the eleven passed and failed by chance. The temperature was 0.3, and no record shows that anybody chose it. It was a class default.

On a comparable question, Jev’s probability moved by 0.006 on average between two runs. Now Jev judges role and domain, code does the arithmetic, and on two days of my own matches no decision changed between two runs. Nate’s guide warns that level answers rank well but make bad pass marks, and we saw it ourselves. When we mapped Jev’s levels to plain shares (0, 25, 50, 75 and 100 percent), almost every job fell below the threshold. We now use the middle of each band of our scoring rubric, so the threshold keeps its meaning. That is a workaround. The cleaner design is a calibrated yes-or-no gate, and it is still to do.

Deterministic, semi-deterministic, AI

For as long as Kitsuno has existed, I have said: deterministic by default, AI by exception. Deterministic first, AI last. Now there is a layer in the middle, and it changes the questions I ask. It does not reduce complexity, at least not at first. There are more layers, more switches, more fallbacks and more tests.

The hardest part was the question at every step: do we want lower cost, fewer tokens or better quality here, and how do we get all three while the system gets stronger? We answered it three ways in two and a half days.

In Be Found we chose cost first, because Be Found is free for our users for now. A plain code filter on the role family removed a third of the language-model calls and lost 3.4 percent of the strong matches. Jev came next, with a stricter threshold wherever code sees no sign of a match.

In the scorer we first wanted Jev for stability. But with the whole candidate record in the state, a Jev call cost about as much as the language-model call it replaced. So my rule became: no Jev unless it also removes language-model calls. We merged the two Jev calls at the door into one and used the freed call after the extraction.

In the writer we chose quality first. Jev checks every line, because one false claim in a document that carries a person’s name is one too many.

Each step showed us the next one. Our shadow log now says that for a third of the jobs that reach the extractor, code and Jev can already fill every field before it runs. For a fifth, they also agree with the extractor on every field.

Ten rules we build by now

  1. One narrow judgment per question. Put the edge cases in the criteria.
  2. Code owns policy, arithmetic, distances, dates, weights and thresholds.
  3. Select, do not generate. Code finds the candidates, and Jev picks one.
  4. Put the text to judge inside the question. Pointers like “line 13 of the list” made Jev mix up neighbouring lines.
  5. Keep the state small. Input is what you pay for.
  6. A Score ranks. A yes-or-no question with clear criteria makes a gate.
  7. When Jev is unsure or fails, a safe path runs: the old path, a clarifying question or a hold. “Not specified” never rejects anything.
  8. Pin the model version. Your thresholds depend on it.
  9. Decide as early as the data allows, before any writing model runs.
  10. Check claims, never stances. In letters and summaries, hold and show; never remove anything in silence.

Most of these come from TypeSafe’s documentation and Nate’s guide. Rules 4, 9 and 10, and the “not specified” part of rule 7, we learned from our own mistakes. One more lesson does not fit in a rule: test on today’s data. Our staging tests replayed older jobs that already had a location, so they could not show the location bug. A gold set from June could not tell us how the gates behave on September traffic.

The danger, and the way back

When I first published this piece, I wrote that Jev was alone in its category. That was wrong, as the update at the end of this section shows. TypeSafe built Jev in stealth for two years, and that head start is real. The company is in the United States, the weights are closed, and zero data retention is only for enterprise customers.

Two of Kitsuno’s ten principles point the other way. The eighth says “open-source models for AI”. The ninth says “data stays in European data centers”. Our data is stored in Europe and we look for European providers where we can, but Jev is not the first US provider in our model routes, and that question stays open for us. I made the exception for Jev with open eyes. Closed weights are acceptable where Jev does the job more reliably than anything else we tested. The checks at the door run for every user, because they are the only way to protect the free quota. Our privacy page names TypeSafe.

Using a US model also made us look hard at what every model receives. The identity data we hold does not reach Jev. Our prompts do not carry the person’s residence or citizenship, and before every model call, whichever the provider, we remove the name, contact data, links, street address and birth date. The only exceptions are tasks where identity is the job, such as reading an uploaded CV. That added complexity, and it made Kitsuno more private with every provider, not only with Jev. We also wrote to TypeSafe about data retention and privacy. We have not had an answer yet.

So we build with a way back. Every Jev call has a fallback. When Jev does not answer, the job pipeline and Be Found go back to the language model, the chat falls back to its old classifier or to a short reply that states only live facts, and the writer holds the draft for review. Most gates also switch off without a deploy, with a row in the database or a setting. Every decision is in our session log, with its reason and its date. If the terms change, we know what to switch off and why it was on.

I would love to see European or open-source models that do this job. Competition would be good here, because it is the best way I know to reduce dependency.

Update, 25 September. One exists already, and I did not know it when I published this piece. Laya comes from Nandakishor Mukkunnoth and his company, ConvAI Innovations. Like Jev, it reads text, writes none and returns typed decisions with probabilities. It is built for more than 100 languages. Its weights are open under Apache 2.0, and its server accepts the same requests as Jev. Nandakishor has published research on models that decide instead of write since March 2025. With Laya, he has made this kind of model open to everyone: the code, the weights and a server that you run yourself.

We tested that it works in general. It ran on an ordinary CPU and answered our production questions in the same request format that we send to Jev. Now we will set it up on our own server and see if we can make it work for our questions. Being able to train it ourselves opens up much more. And on our own server in Europe, it costs nothing per call and fits our eighth and ninth principles. Laya could be the way back that I hoped for, and the credit for it goes to Nandakishor.

Built with Claude

I build Kitsuno with Claude as my engineering partner. For most of the Jev work we used Claude Opus 5.5, released on 22 September, at maximum effort, because this work needed deep reads of our own code and data more than speed. Those reads found most of the old bugs while we redesigned the flows. That was the unplanned gain of the week: redesigning with a new frontier model was also an audit of every flow we touched.

What comes next

  • The extractor becomes the place for doubt: it runs only when code and Jev cannot fill a field.
  • The classifier that reads social posts for job ads gives only closed answers. It is the next candidate.
  • Yako, our tender and grant desk, gets a Jev shadow run next to its current judge.
  • Laya goes onto our own server: we train it for our questions and then run it in shadow next to Jev.
  • The writer’s Jev state gets smaller, and every Jev call gets logged where we can see it.
  • We collect at least two days of real data before we change a threshold, and more for the chat gate.

Whitehead saw progress in the operations that no longer need thought. It took us two and a half days of hard thinking to move a few of ours there: small decisions, made early and cheaply, each with a switch next to it. At the end of the road there is still a person. They read the letter, and they decide whether to send it. That part stays theirs.

For builders: what a Jev call looks like

Here is a simplified version of our check at the door. The question texts and criteria are the ones we run in production; the ad and the role are made up. The state is the ad, with e-mail addresses and phone numbers replaced. There is one yes-or-no question per target role, and one per language that the job might require.

{
  "model": "jev-1.13.0",
  "state": {
    "ad": "Title: Senior Data Platform Engineer\nLocation: Munich\n\n[ad text, e-mail addresses and phone numbers replaced]"
  },
  "questions": {
    "role_0": {
      "type": "noul",
      "instructions": {
        "question": "Is the main work of the job in `ad` the same kind of work as `role`? Ignore the level of seniority.",
        "role": "Data engineering: data engineer, analytics engineer"
      },
      "criteria": {
        "true": "The core duties of the job belong to the work area that `role` describes. A different job title with the same core duties counts.",
        "false": "The core duties belong to a different work area, even if the ad shares some words, tools or the sector with `role`."
      }
    },
    "lang_de": {
      "type": "noul",
      "instructions": {
        "question": "Does the person in the job in `ad` need to work in `language`?",
        "language": "German"
      }
    }
  }
}

Code decides what the probabilities mean:

# Role: the best match over all target roles of the person.
p_role = max(answers[q]["noul"] for q in role_questions)
if p_role < 0.10:
    prune(job, reason=f"role_prejudge:p={p_role:.2f}")   # no extraction, no score

# Language: only languages the person does not work in are asked at all.
for lang in candidate_languages:
    if answers[f"lang_{lang}"]["noul"] >= 0.75:
        prune(job, reason=f"language_intake:{lang}")

After the extraction, a Choice with five levels becomes the role part of the fit score. The level shares are the middles of our rubric bands, and code takes the expected value:

BAND_MIDDLE = {"none": 0.15, "little": 0.45, "some": 0.65, "most": 0.82, "all": 0.95}
role_share = sum(p * BAND_MIDDLE[level] for level, p in answers["role"]["probabilities"].items())
role_points = role_weight * role_share   # location and seniority points come from code rules

In the chat, Jev gives two probabilities per sentence, and one rule in code decides:

def flagged(claim, support, number_not_in_facts):
    return ((claim >= 0.2 and support < 0.15)
            or (claim >= 0.5 and support < 0.5)
            or (claim >= 0.5 and number_not_in_facts))   # numbers are checked in code, not by Jev

And the way back, around every call:

mode = global_switch("jev_intake")   # "act", "shadow" or "off": one database row, read every 60 s
try:
    answers = ask_jev(state, questions)
except JevUnavailable:
    answers = None                   # the job takes the old path

The same request also works against Laya’s open server, which you run yourself. The thresholds do not carry over from one model to the other, so calibrate them again.

Sources: Nate B. Jones, Jev and the Missing Primitive and Find the Jev-shaped problems in your software; TypeSafe documentation; Nandakishor Mukkunnoth, Laya.