Where the decision threshold actually bites

Both systems emit a per-label score, and both were scored at a frozen 0.5 cut. Moving that cut shows how much headroom each one has. Jev’s curve responds; Comprehend’s barely moves, because almost all of its scores sit pinned at 0 or 1, including the ones it gets wrong. Measured on the same 251 test documents.

F1 as the threshold moves

Higher is better. A flat line means the threshold is not a usable lever: the errors sit at the same extreme scores as the correct answers, so no cut separates them.

Comprehend Jev dotted rule and dots mark the 0.5 cut we report

Internal references, log scale

On the linear panel above both lines crowd the ceiling, so the gap between them is hard to read. Putting F1 itself on a log axis does not help: every value is near 100, and log compresses exactly there (the interesting band would shrink from 6.3% of the axis to 4.7%). What separates them is a log scale on the distance from a perfect 100. Jev sits 0.50 off perfect at the 0.5 cut and Comprehend 6.05, which is a factor of 12, just over one decade. The axis is labelled in F1, so read the numbers normally; only the spacing is logarithmic, and better is still up.

Comprehend Jev gridlines are evenly spaced in “distance from 100”, not in F1

Why: where the scores sit

Every score each model produced, both labels pooled. Comprehend is effectively a hard label wearing a decimal point: 94.8% of its scores fall below 0.05 or above 0.95. Jev spreads its uncertainty across the range (22.9% at the extremes), which is what makes its threshold curve move.

Table view