Both systems emit a per-label score, and both were scored at a frozen 0.5 cut. Moving that cut shows how much headroom each one has. Jev’s curve responds; Comprehend’s barely moves, because almost all of its scores sit pinned at 0 or 1, including the ones it gets wrong. Measured on the same 251 test documents.
Higher is better. A flat line means the threshold is not a usable lever: the errors sit at the same extreme scores as the correct answers, so no cut separates them.
On the linear panel above both lines crowd the ceiling, so the gap between them is hard to read. Putting F1 itself on a log axis does not help: every value is near 100, and log compresses exactly there (the interesting band would shrink from 6.3% of the axis to 4.7%). What separates them is a log scale on the distance from a perfect 100. Jev sits 0.50 off perfect at the 0.5 cut and Comprehend 6.05, which is a factor of 12, just over one decade. The axis is labelled in F1, so read the numbers normally; only the spacing is logarithmic, and better is still up.
Every score each model produced, both labels pooled. Comprehend is effectively a hard label wearing a decimal point: 94.8% of its scores fall below 0.05 or above 0.95. Jev spreads its uncertainty across the range (22.9% at the extremes), which is what makes its threshold curve move.