The Wiggle Framework for LLM as a Judge
The Wiggle Framework is an evaluation method introduced in Meta’s research paper, “Jagged Judges: Epistemic Stability Under Perturbation.” It is designed to test how easily large language model (LLM) judges abandon their original verdicts when challenged. Why this matters Traditional LLM evaluation often measures a judge’s accuracy once against a fixed gold-standard dataset. That tells us how […]
Read More →