Interlude: Jev and the economics of no-training
I re-ran an old AWS Comprehend project against TypeSafe's Jev. With no training data it beat the trained classifier at roughly a two-hundredth of the cost. It also showed where the risk goes when you stop training models and start writing questions.
Get an email when the next article lands
A short detour from the series. Jev is the most talked-about model release this month, I had a dataset sitting in a drawer that was a near-perfect test for it, and the result says something about where the work in AI engineering is moving. The numbered articles pick up again next week.
I dug out the training data from a custom entity recognition model I’d built on AWS Comprehend for a past personal project and used it to test TypeSafe’s Jev, the judgment model that’s been getting a lot of press over the last week. Same data, split the same way into training data and test data, but a very different idea of how the underlying AI should be built.
The task: classify sections of documentation by what they contain: internal references, pointing to another part of the same document; external references, pointing out to other documentation; both; or neither.
What each approach needed
Comprehend needed around 600 labelled sections. Getting them was its own small project: an LLM drafted a first pass of labels, a human (me) reviewed every one, and then the model was refined through a human-in-the-loop retraining loop built on Comprehend flywheels and SageMaker Augmented AI. So, roughly five hours of human review, to create a gold standard, before there was a model at all. That gold standard was then divided in two: training data, which the model learns from, and test data, kept aside so it can be scored on sections it has never seen. Then a training run, and a hosted endpoint to make the model available for use.
Jev needed zero labelled sections, no training step, and about an hour of massaging its prompt. You describe what you’re looking for in plain language and ask yes/no questions, here one for internal references and one for external, and for each question you get back the probability that the answer is yes. All of the test data scores in under ten seconds of API calls.
Jev is essentially a zero-shot classifier, so the fair test was to train a fresh Comprehend classifier on the same training data, lock both configurations before looking at the test data, and put the two side by side on the same ~250 test sections that neither had seen. Usually when we’re handling blocks of text in a classifier, they are referred to as documents, so for the rest of the article we’ll call these sections documents.
Jev beat my custom Comprehend model for both internal and external references: 99.5 F1 against 93.9 for internal, 84.4 against 72.4 for external (F1. is a standard metric for machine learning models) Jev won on MCC (Matthews Correlation Coefficient) too, which I call out because MCC can’t be flattered by lopsided test data the way F1 can.
The costs were even more interesting
I was pleasantly surprised with the performance, but the costs were even more interesting. Training the Comprehend classifier cost $1.48 and 34 minutes. Classifying the documents with it cost another $0.38. Jev did the same documents for under a cent, in nine seconds, with no training cost/time. That’s roughly 200x on the total cost, and it still flatters Comprehend, because it leaves out the five hours of human review behind the training data, which cost far more than the rest of this experiment put together. Where Comprehend does potentially win is live, one-at-a-time requests: it offers a real-time endpoint built for low-latency answers, but that comes with a significant cost. (Jev’s nine seconds was for all 251 test documents, eight at a time, so roughly a third of a second per request.)
The running costs diverge far more than the build costs. A trained Comprehend model bills $0.50 a month just to exist, and a real-time endpoint to answer live requests runs about $1,296 a month. Classification can run asynchronously for far less if you don’t need live answers, which is what I did for this test but that takes multiple minutes to return results. Jev’s pricing is still provisional, but there is no charge for output tokens and input tokens are extremely cheap compared to what we’re used to, and a difference in cost that large makes this solution well worth considering.
The scores are more useful, not just better
Both systems give a probability score for every label, and both were judged at a fixed cut-off of 0.5. The interesting part is what happens when you move that cut-off, and the chart below lets you explore exactly that.
Almost all of Comprehend’s scores, 94.8%, sit below 0.05 or above 0.95. That includes its mistakes: of its 29 false external-reference calls, 26 score above 0.9. It isn’t unsure about those; it’s confidently wrong, so no cut-off can separate them from its correct answers, and the only way to improve it is to retrain.
Jev spreads its uncertainty. Only 2 of its 14 false calls clear 0.9, so raising the cut-off to 0.8 removes twelve of them for the loss of one correct answer. That’s genuinely useful in production: you can send the uncertain middle to a person, let the confident ends through, and tune that trade-off without touching the model. How do you work out the right threshold for Jev, to get the best results for your query? You can’t without human-labled data: without those five hours of review there’d be no ground truth to score either system against. Zero-shot doesn’t mean you get to skip the labelling, it just means the labelling moves from the training data to the test data, and you need far less of it.
The improvement loop, and a new failure pattern
My data set was small and doesn’t tell you how either approach scales. But the biggest takeaway for me is the improvement loop. To make the trained model better: label more documents, run them through human review, retrain, redeploy. To make Jev better: edit the question, rerun. While I was building this, one sentence added to the question I was asking Jev was worth about 10 F1 points on its own.
TypeSafe’s launch post says Jev “can’t hallucinate” and “never makes type errors”, and in all my testing, that was true. But Jev has a more subtle failure pattern, it can only choose from the answers you offer it. A multiple-choice question in Jev returns a probability for each option, and those probabilities are relative to each other: across every response I recorded, they add up to 1. So if the right answer isn’t on the list, Jev can’t say so. It picks the closest thing that is. So hallucination may no longer be the risk but failing to provide the right options can give you just as dangerous a result.
I tested that directly. Twenty of the test documents contain no reference of either kind. I asked each document one question with three options, internal only, external only, or both, and deliberately left out “neither”.
It got all twenty wrong, fourteen of them at 0.8 confidence or higher. Every answer was a perfectly valid option; nothing was hallucinated. Add “neither” back to the list and it got eighteen right. Nothing else changed.
My test case was simple, but if yours is more complex, how do you ensure you haven’t missed an option? The mistakes were hidden well: the missing option didn’t affect the other ~230 documents at all, so the overall accuracy still read a respectable 87%, and you’d only find the problem by looking at that one class, where the score was exactly zero. Confidence didn’t give it away, the wrong answers were a little less confident than the right ones, but not by enough to catch them: sending every result under 0.8 to a person would have caught six of the twenty but adding one option to the list fixed what no threshold could.
I’ve seen the same shape somewhere else too. I’m trying Jev as the judge in a hook that decides whether my coding agent/harness is allowed to stop work, and rewording one yes/no question, from asking what the agent’s message claimed to asking what the rules actually permitted, moved the same case from 0.78 to 0.19. Same model, same input, different question…it feels almost like we’re back to the early days of prompt engineering. There will be more on this in a future article in this series.
So the practical rules I’ve taken away are short:
- Always give it a way to say no. “None”, “neither”, “other”. If the answer can be absent, the option has to be there.
- Ask about the case you care about directly. A yes/no question is scored on its own, so “no” is always available. The yes/no version of my original questions got 17 of those same 20 right without any “neither” option. The Jev Choice and Score both have this failure mode, while Noul (ie. true/false questions) does not.
- Test with the case you didn’t plan for, and read the per-class numbers, not just the headline numbers.
None of that takes anything away from the result. Jev beat a purpose-trained model on its own data, at a fraction of the cost, with more useful confidence scores, and I’d use it again for this kind of data processing/decision.
Now, if I’m talking purely about economics, then I should probably call out that I could use a BERT/deBERTa based model, for classification and run it locally on my laptop. I don’t know how that would perform against this test but maybe that will be a future article.
So, for cloud inference the cost structure of this kind of NLP task has inverted. You no longer pay to label, train and host a model before you can ask it anything. What you pay for instead is getting the question right, because that’s now where the improvement comes from, and where the mistakes come from too. It’s a common theme in this series: the problems haven’t changed, the bottlenecks have just moved again.


