BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories
arXiv:2606.22329v1 Announce Type: new Abstract: LLM-as-a-judge has become the dominant approach to scalable evaluation in NLP pipelines, yet judges themselves carry systematic biases that raw accuracy hides: they favor responses placed in slot A (position bias), they prefer longer responses regardless of quality (verbosity bias), and their reliability degrades sharply in lower-resource languages. We introduce BabelJudge, an open-source benchmark and reliability audit framework that measures all