256-CANDIDATE MULTIMODEL AUDIT · SOL-HIGH JUDGE

Which models anticipate human reviewer criticism?

Covered at ≥ 0.70 Below threshold GPT-5.6 Sol High dense matching

Loading the audit bundle…