July 17, 2026
Your LLM judge is wrong on the median session, not the tail
Most agent-eval work targets the scary tail. But a flattened LLM-as-a-judge is wrong on the median trajectory. Meet the agent judge that investigates instead of scoring.
Read more →