“Men learn while they teach.”

Seneca, Letters to Lucilius, 7.8

The last piece ended with an update. Three days after our first real Jev calls, I learned that an open model of the same kind already existed. Laya, from Nandakishor Mukkunnoth and ConvAI Innovations, reads text, writes none and returns typed decisions with probabilities, like Jev. Its weights are open under Apache 2.0, and its server accepts the same requests that we send to Jev.

So we sent it our requests. On a small staging test, it found the right country for 26 percent of the job ads, and the right work mode for 50 percent, where Jev had 89. On the test set that we built the same day, the untrained Laya answered 44 percent of our questions correctly. Jev answered 94 percent.

That number was not a verdict on Laya. It told us what the work would be. An open model that decides is a starting point, and you teach it your questions. For us, the training is the product. We were not the only ones to see this: AJ Awan’s comparison of Laya and Jev reports a typed-decision score of 0.362 for Laya out of the box, and 0.766 after fine-tuning.

This is the story of the eight days since then: how we taught it, where we trained it, how it now runs next to Jev on every real request, and what still keeps us from the switch.

What Jev carries today

First, the thing we want to replace works, and it works well.

Since 23 September, Jev has answered 35,812 logged calls for us, with 76 million input tokens, for USD 3.68 in total. In the last seven days it made 30,738 calls. Almost all of them sit in the job pipeline and in Be Found: 14,920 checks at the door before any language model reads an ad, 11,820 role checks in Be Found, and 3,968 match judgments after the extraction.

The effect is in the language-model calls. In the two weeks before the first Jev gate (9 to 22 September), all our agents together made about 4,400 language-model calls a day. In the nine days after it (24 September to 2 October), they made about 2,000 a day, while the number of new jobs per day almost doubled. Per 100 new jobs, that is about 300 calls before and about 100 after (73 if we count one very large crawl on 1 October). Language-model tokens fell from about 11 million a day to about 5 million. This is a broader count than the one in the last piece, across all agents and mixed traffic, so read it as a direction, not as a controlled test. The direction is clear.

So money is not the reason for Laya. At our volume, Jev costs a few dollars a week. The reasons are in Kitsuno’s principles. The eighth says “open-source models for AI”. The ninth says “data stays in European data centers”. Jev is closed and runs in the United States. I made the exception with open eyes, and I promised a way back. An open model on our own server in Europe is that way back. It also gives us something that Jev cannot: a model that we can train for our own questions.

Rule zero: never learn from the model you want to replace

The fastest way to train Laya would have been to copy Jev. We had thousands of Jev answers on real job ads, and each one is a ready label.

We did not use one of them. TypeSafe’s customer agreement (section 2.3(b), as updated on 23 September) prohibits using Jev’s output “to perform model distillation, train a model to imitate the output of the Services, or develop … a similar or competing product”. That is our reading, not legal advice, but the text is clear enough. Jev’s answers are a comparison for us, never a label and never a calibration target. We built this into the system: the database role of the Laya worker cannot read Jev’s answers at all.

That rule made the work slower, and it made the result more honest. When Laya agrees with Jev, it did not learn the answer from Jev. When it disagrees, the disagreement can mean something.

The teachers

Without Jev, the labels had to come from somewhere else. A model that writes the training labels is called the teacher, and we had two.

The first was a local model with 9 billion parameters, on the external GPU of my old MacBook. It cost nothing, and it was not good enough. Its best version answered 91 percent of our test questions correctly, and a student does not pass its teacher. Laya trained on it reached 86 percent. We tried notes in the prompt for the hard cases, and they raised the score on the 100 ads that we wrote them for, but not on the test set. We also measured two large open models as teachers. They came closer (93 and 92 percent), but not close enough.

So we used a frontier model as the teacher, and human review for the questions that were unclear or borderline. Uncertain answers were left out of the training data. In total, about 4,400 job ads with about 44,000 questions went into the training. The reference answers in our test sets were made the same way: the frontier model first, and I reviewed the open and borderline cases myself.

As we read the frontier model provider’s terms, they allow outputs to train specialised classifiers and extraction tools that do not compete with the provider’s own models. Laya, as we train it, is such a classifier: it reads a job ad and picks from answers that we defined. The wording of the terms is not fully clear on every point, so we asked their legal team for a written confirmation. We have no answer yet. If the answer is no, we train Laya again with labels from another teacher, and build a new test set to go with them.

From 42 hours to 37 minutes

The first training runs ran on the CPU of the same old MacBook, a 2017 MacBook Pro that now runs Omarchy Linux. A run took 39 to 42 hours, with the processor at 93 to 100 °C. The first run reached 86 percent on our test set. The second, with the frontier teacher, reached 91 percent. Jev was at 94.

On 2 October I changed the goal. Until then, Laya was to become as good as Jev. From that day, it was to replace Jev, and the result had to be better. That needed more runs than a laptop can do in a week.

We moved the training to a free Kaggle GPU. The same run took 37 minutes instead of 42 hours, with the same result. I did not want to rent a GPU: for that money, I can run Jev for a long time. Kaggle belongs to Google and runs in the United States, so we drew a line there too. Kaggle gets public job ads, teacher labels and test questions. It never gets the reference answers, and it never gets data about the people who use Kitsuno.

The test set that chooses is no longer a test

With 37-minute runs, we trained a series of models in one day. The second Kaggle model, k2, answered 96.2 percent of our test questions correctly. Jev had 94.1. On paper, Laya had passed Jev.

But we had used that test set to choose between models. A test set that chooses becomes part of the training, just more slowly. So we built a second test set from 150 new job ads, and decided that it would never choose a model. On the clean set, k2 and Jev were equal: 93.4 percent each.

That second set is now the one I trust. On both sets together (574 decisions), the current model k6 is at 95.8 percent, and Jev at 93.7.

ModelFirst setClean setBoth
Laya, untrained44.9 %43.9 %
k296.2 %93.4 %94.8 %
k4 (more training ads)95.8 %95.5 %95.6 %
k6 (current)95.8 %95.8 %95.8 %
Jev94.1 %93.4 %93.7 %

Better is not the same as proven. Over the 574 decisions, 31 are right only for k6 and 19 only for Jev. That difference is not significant yet (p = 0.12). On one question type, it is: k6 reads the geographic scope of a job correctly 15 times where Jev does not, and never the other way round. Jev reads a country in the location line as “open only to people in that country”, where our question calls that case unclear. On seniority, Jev is better (9 to 5), and seniority is Laya’s weak point.

Swapnil Talekar argues that Jev’s real innovation is calibration: when it says 90 percent, it is right about 90 percent of the time. We hold Laya to the same test. When k6 is at least 90 percent sure, which is the case for 86 percent of the decisions, it is right 97.8 percent of the time. On the same decisions, Jev is right 95.7 percent of the time. Seniority is again the exception: there Laya is 96 percent sure and 82 percent right.

There is one more caveat, and it matters. The same frontier model wrote Laya’s training labels and the first draft of the reference answers. That favours Laya. So I looked for checks that do not depend on that teacher. One is the 53 decisions that I made myself, in the open and borderline cases and in blind spot checks: k6 got 48 of them right, Jev 36. The other is production.

Shadow on prod

A shadow run means that the new model gets a copy of every real request and answers it, but the old model’s answer is the one that counts. Nobody depends on the new answer. It is recorded and compared.

Since 2 October, 22:03 UTC, every job-side decision that Jev makes in Kitsuno is also made by Laya: the checks at the door and the role checks in the job pipeline and in Be Found. These calls carry job ads and target roles, and no identity data. The chat and the writer are not part of the shadow run. Jev still makes every production decision.

In its first 18 hours, the shadow run collected 1,158 real items. Every Laya model answers every item, so we can compare models with each other and with Jev on exactly the same requests.

The first day showed something that our test sets could not show. Laya read Portuguese job ads as jobs in Brazil, or as “country not stated”. Our test sets had few Portuguese ads, so they never saw the problem. One day of production showed it. We trained the next model, k6, with 265 production ads, each with the country that the source itself states (never Jev’s answer). On 276 other production ads that k6 had never seen, it found the right country 96.4 percent of the time, where the model before it had 91.3 percent. On the Portuguese ads, it went from 60 of 69 to 69 of 69. Production found the problem, and one more training run fixed it.

On production items, k6 now gives the same answer as Jev for 97 percent of the language questions and of the role questions, 95 percent of the employment types, 93 percent of the work modes and 92 percent of the geographic scopes. Seniority is lower. Where the ad shows a level, they agree on 89 percent of 491 items, and where they differ, neither model is clearly the stricter one. The most common difference: Jev says “director”, Laya says “lead”. Who is right there is the next thing we check.

From too slow to about four seconds

Accuracy was one half of the problem. On 3 October, Laya needed 12.7 seconds for one intake item on our server’s CPU. Jev needs 0.29 seconds per call. My words that day: if we don’t get near Jev speed, a big batch takes too long to digest in real production, and maybe tomorrow we talk 10,000 intakes.

Two changes did most of the work. The first was a number format: our server’s CPU computes 16-bit numbers (bf16) about twice as fast as 32-bit numbers, with the same answers. The second was in how Laya reads. Until then, it read the job ad once for every question, so ten times for ten questions. We trained a model that gets all questions and the ad in one sequence, so it reads the ad once. Same accuracy, half the time. As I said at the time: that is what Jev does. One read. The trick that halved our time was already in the model we want to replace.

StepMedian time per intake item
fp32 on the server CPU12.7 s
bf168.0 s
bf16, read the ad once4.1 s
k6 on the server (2 workers)3.7 s
k6 on the old MacBook’s GPU3.3 s
Jev (API)0.29 s

Laya is still more than ten times slower than Jev per item. But two server workers now do about 32 items a minute, so 10,000 intakes take about six and a half hours a day, and about four with a second lane. No GPU is rented.

The GPU that PyTorch cannot use

The second lane is the old MacBook again. It has an AMD Radeon Pro 580X in an external GPU box, with 8 GB of memory. PyTorch cannot use it, because AMD’s ROCm does not support that generation of chips.

Instead, we compiled Laya with IREE for Vulkan, a graphics interface that the card does support. On 20 test ads, the Radeon gave the same top answer as PyTorch in 189 of 189 questions. On 20 production items, answered again on the server, it was 188 of 188. The same card also runs the local model that reads tenders for Yako at night. Together they use 7.8 of its 8 GB.

The server and the MacBook take items from the same queue. A free lane takes the next items, so the faster lane does more of the work. If a lane stops, the other lane gets its open items after ten minutes. If the GPU fails, the MacBook stores no error answers and leaves the items to the server. When the MacBook is offline, nothing breaks. When it is online, it helps.

We also tried things that did not work. 8-bit numbers lost too much accuracy. llama.cpp on the Radeon was fast, but it cannot do the “read once” format. A small standby server in Helsinki had too little memory.

What went wrong

  • The untrained base model answered production for 16 hours, and nobody saw it. Our notes said “staging only”, and the data said otherwise. Now the server accepts only trained Kitsuno models.
  • A test run ran our server out of memory, twice. Only the test processes were killed, but since then every test runs with its own memory limit. Later, the workers were killed three times at their memory limit, until we capped the length of an input.
  • An installer passed its own check (189 of 189) and then stopped, because the script read “5 of 189” from the model name “k5” in the result line.
  • A calibration fix for seniority missed seniority. The calibration set had only 47 ads, too few seniority questions for the rule to work.

None of these broke anything for a user, because Jev made every decision while they happened. That is the point of a shadow run.

Ten rules we build by now

  1. Shadow first. The new model answers every real request next to the old one, and nobody depends on it.
  2. Give every model the same items. Compare only on the same items.
  3. Never learn from the model you want to replace.
  4. Labels come from what the source states, or from an independent teacher.
  5. Keep one test set that has never chosen a model.
  6. Train on today’s production data. It shows what the test set cannot.
  7. Read the input once, and ask all questions about it in one pass.
  8. Every lane fails closed. A lane that cannot answer leaves the item to another lane.
  9. Measure speed on real items, with the real limits of cores and memory.
  10. Write the switch criteria before you see the results.

Where it stands

Leonardo Gonzalez writes that the value of these models depends on how many cases they cover and how often the accepted cases are wrong, measured against the cost of the fallback. Our switch criteria ask the same question. We wrote them on 25 September, before any trained model existed. This is where k6 stands:

CriterionStatus
Per question type, at least as accurate as Jev4 of 8 types
No type with 20 or more items more than 10 points below Jevmet (largest gap: seniority, 7.2 points)
Calibration error below 0.05met (0.027)
Gate questions: the same production decision as Jev on 95 % of the itemsnot met (role keep-or-prune: 83 %)

So Laya does not decide yet. For the next days, k6 stays unchanged on both lanes and collects shadow data. I look at it every day. Then there are two ways forward. Either the data supports a switch, question type by question type, or it tells us what the next model, k7, must learn: a larger calibration set for seniority, and a careful look at the cases where Jev and Laya disagree on seniority, to see who is right.

When the switch comes, the three paths in the shadow run will cost nothing per call. Last week they carried 87 percent of Jev’s logged calls. TypeSafe named Jev after William Stanley Jevons, who saw that cheaper coal leads to more coal being burned. A decision that costs nothing per call is the far end of that curve. Jev stays, for now, where Laya is not tested yet: the match judgment after the extraction, the chat and the writer.

I also want to say what has not changed. We wrote to TypeSafe about data retention in September, and we have no answer yet. Jev did its work well, and it still does. It showed us a category of model, and it made us walk every flow of Kitsuno again. Laya exists because Nandakishor Mukkunnoth decided that this kind of model should be open to everyone, with the code, the weights and a training notebook. The credit for the way back is his.

Seneca wrote to Lucilius that people learn while they teach. We spent eight days teaching a model our questions. The model learned. We learned more: about our test sets, about our own data, and about a GPU in a box that everyone had given up on. In about a week, I plan to write the next part: which decisions Jev hands over, and how that brings the job side of Kitsuno back in line with our eighth and ninth principles.

For builders: the numbers and the code

All models on both test sets

Accuracy on the decisions of the two test sets. The first set (287 decisions) was used to choose between models; the second (287 decisions, 150 new job ads) never chose a model.

ModelTrained onFirst setClean setBoth
Laya, untrained0.4490.439
r1MacBook CPU, local 9B teacher, 39 h0.864
r2MacBook CPU, frontier teacher, 42 h0.913
k1Kaggle T4, 1 pass, 37 min0.916
k23 passes0.9620.9340.948
k35 passes, calibrated0.9580.9410.949
k4more training ads0.9580.9550.956
k5read the ad once0.9620.9550.958
k6production country examples0.9580.9580.958
Jev (jev-1.13.0)0.9410.9340.937

k6 and Jev per question type, on all 574 decisions:

Question typek6Jev
Work mode0.9590.919
Employment0.9740.987
Seniority signal0.9860.973
Seniority0.8210.893
Geo scope0.9730.773
Language0.9730.986
Role0.9730.986
Country0.9720.972

Same answer as Jev on production items (k6, shadow run):

Question typeSame answer as Jev
Language97 %
Role (per question)97 %
Employment95 %
Work mode93 %
Geo scope92 %
Seniority signal90 %
Country89 %
Seniority67 %

Production uses the seniority answer only when the ad shows a level. On those items, k6 and Jev agree on 89 percent of 491. Without a level they agree on only 47 percent, but production does not use those answers.

The training recipe

  • Base: Laya’s multilingual checkpoint (a ModernBERT-type encoder), laya 0.3.20.
  • Method: the fine-tuning method of the Laya notebook. A policy-gradient step with a proper scoring rule as the reward, plus a soft cross-entropy on the teacher’s probabilities. The encoder and the decision head train; the large embedding table stays frozen.
  • Labels: about 44,000 questions on about 4,400 job ads. Answers that the teacher was not sure about stay out.
  • Calibration: one temperature per question type, fitted after training. Soft labels first made the model under-confident (it was right more often than it said). Fitting the temperature against the teacher’s top answer took the calibration error from 0.107 to 0.015.
  • Hardware: a free Kaggle T4. Kaggle gets public job ads, teacher labels and the test questions, never the reference answers and never seeker data.

The tap

Every job-side Jev call goes through one function after Jev has answered. It never raises, it waits at most two seconds, and it writes a copy only when the path is switched on. The switch is a database row, read every 60 seconds:

# Job ads and target roles only, no identity data.
ALLOWED_PATHS = ("jev_intake", "jev_role_gate", "jev_handshake_gate")

async def tap(db_pool, path, request, answers,
              model=None, ms=None, ref=None, labels=None, extra=None):
    try:
        # '__global__' row: {"rate": 1.0, "paths": [...]}
        if db_pool is None or not await sampled(db_pool, path):
            return False
        req = {"state": request.get("state"),
               "questions": request.get("questions")}
        if extra:
            # questions only Laya answers, e.g. the country
            # when the source itself states it
            req["laya_extra"] = extra
        await asyncio.wait_for(db_pool.execute(
            "INSERT INTO laya_shadow_items (path, ref, labels, jev_model,"
            " jev_ms, request, jev_answers) VALUES ($1, $2::jsonb,"
            " $3::jsonb, $4, $5, $6::json, $7::jsonb)", ...), timeout=2.0)
        return True
    except Exception:
        return False   # never break the Jev path

The request column is json, not jsonb. jsonb sorts the keys, and that changed the order of the options in a Choice.

Never learn from the incumbent, in SQL

The worker’s database role can read the inputs and move the lease. It cannot read Jev’s answers, the labels or the references:

REVOKE ALL ON laya_shadow_items, laya_shadow_answers FROM laya_shadow;
GRANT SELECT (id, created_at, path, request,
              leased_at, leased_model, lease_count)
      ON laya_shadow_items TO laya_shadow;
GRANT UPDATE (leased_at, leased_model, lease_count)
      ON laya_shadow_items TO laya_shadow;
GRANT SELECT, INSERT, UPDATE ON laya_shadow_answers TO laya_shadow;

Two lanes, one queue

Both lanes pull with the same model id. A lease of ten minutes gives each item to one lane; when a lane stops, its items come free again:

WITH picked AS (
    SELECT i.id FROM laya_shadow_items i
    WHERE i.created_at > now() - make_interval(days => 14)
      AND i.request IS NOT NULL
      AND i.lease_count < 10
      AND NOT EXISTS (SELECT 1 FROM laya_shadow_answers a
                      WHERE a.item_id = i.id AND a.laya_model = $model)
      AND (i.leased_at IS NULL
           OR i.leased_at < now() - make_interval(mins => 10)
           OR i.leased_model IS DISTINCT FROM $model)
    ORDER BY i.created_at DESC
    LIMIT $n
    FOR UPDATE SKIP LOCKED
)
UPDATE laya_shadow_items i
   SET leased_at = now(), leased_model = $model, lease_count = i.lease_count + 1
  FROM picked WHERE i.id = picked.id
RETURNING i.id, i.path, i.request;

Every model has its own answers, keyed by the model id, so k4, k5, k6 and Jev can be compared on exactly the same items. The worker also stores Jev’s confidence formula next to Laya’s own fields, (n * p_max - 1) / (n - 1) for n options, so the report compares like with like.

The GPU lane fails closed. When the device or the runtime fails, it pushes what it has answered, pushes no error records (an error record counts as an answer and would keep the item from the other lane), waits 15 minutes (longer than the lease) and exits. systemd starts it again. While the Yako night reader runs on the same GPU, the lane idles without its model in memory.

Read the ad once

Laya’s own format builds one sequence per question: [CLS] question [SEP] [MASK] option … [SEP] ad [SEP]. Ten questions read the ad ten times. The packed format (k5 and later) puts every question block and the ad into one sequence (shown here on several lines):

[CLS] q1 [SEP] [MASK] o … [SEP]
[CLS] q2 [SEP] [MASK] o … [SEP]
…
ad tokens [SEP]

Three rules keep the answers the same as in a row of their own:

  • Position ids: each question block ends just before the ad, and the ad starts at the same position for every block. The encoder uses rotary positions, which are relative, so every block sees the ad at the same distances as in its own row.
  • Attention: a block attends to itself and to the part of the ad that its own row would hold. The ad attends only to the ad, so it is read once, independent of the questions.
  • Sliding-window layers keep their rule on positions (64 tokens), and pad tokens attend only to themselves.

The weights and their layout do not change; only the input and the masks do. A model must still be trained in this format: k5 kept the accuracy of k4 and halved the time. On a CPU the attention scores of a pack are materialized, so serving splits the questions of an item into packs of at most 1,536 tokens. That split is exact, because a block depends only on itself and the ad.

The Radeon lane

PyTorch has no backend for a Polaris GPU (AMD’s ROCm does not support it), so we compiled the model with IREE for vulkan-spirv:

  • The compiled core is the encoder (with plain matmul attention instead of sdpa), the type embedding and the two head layers. It returns the hidden states.
  • The packs, the masks (as additive float masks, 0 or −1e9), the scorer at the answer markers and the decoding stay on the CPU in PyTorch.
  • The core is compiled for fixed pack sizes and the weights load from the model folder at start, so a later model with the same structure (k6) runs on the same compiled files.
  • The check before a lane goes live: the same items through the compiled core and through PyTorch, same top answer on every question (189 of 189 on 20 test ads; 188 of 188 on 20 production items answered again on the server).

What we did not use

  • int8 weights: too much accuracy lost.
  • llama.cpp on the Radeon: fast, but it cannot do the “read once” masks.
  • Rented GPUs, for training or serving.

I build Kitsuno with Claude as my engineering partner.

Sources: Nandakishor Mukkunnoth, Laya and its server; TypeSafe documentation and Master Customer Agreement; Nate B. Jones, Jev and the Missing Primitive; Leonardo Gonzalez, Jev’s Open Rivals Test the Decision Model Idea; Swapnil Talekar, Jev and the Return of the Classifiers; AJ Awan, Laya: The Open-Source Jev Alternative, Benchmarked Honestly; IREE; the previous piece, Jev in production.