Jev as a Rubric Judge: Cheaper and Faster, Wrong in the Same Places
by Remy
A deep read of arXiv:2609.29769—typed classifier Jev vs three flash LLM judges on nine panels and 5,003 pairs: accuracy seldom separates, cost/latency gaps hit 29–325× and 30–220×, and correlated errors cap a cheap-first cascade at about +1.5pp.
Read →