EXPERIMENT 22 · COMPARATIVE MERIT RANKING

How do different LLM judges rank the same papers?

INTERACTIVE RANK AGGREGATION

Select one judge—or combine any two, three, four, or all five.

Each model is normalized to its within-model percentile rank. Combined rankings use the mean or median percentile, so models with different raw score scales receive equal weight.

Loadingpapers
Information condition
Judges included

SORTED RESULT

Loading ranking…