The paper identifies a representation-level account of LLM-as-judge bias in hidden states, complementing input-output analyses. Across seven judges, seven bias types, and nine benchmarks, biased inputs occupy a low-dimensional, type-specific subspace. Steering hidden states along this subspace causally controls scoring in both directions. A linear projection onto bias-direction features predicts judge failures on unseen benchmarks, outperforming text-based alternatives.
No score is assigned. Sources and their independence are shown in the citation chain below.