Free tool · no signup · runs in your browser
LLM-as-a-Judge eval builder
Pick a failure mode and get a complete evaluator prompt — single criterion, pre-specified procedure steps, bias defenses, and the research behind each choice. Built on the same frame as the generate-eval skill; nothing you enter leaves this page.
1 · What failure mode should the judge catch?
2 · Is there a ground truth to check against?
A reference is a ground truth the judge can check against — retrieved chunks for RAG, gold SQL, a reference summary.
3 · One output, or comparing two?
4 · Bias defenses
Evaluator prompt
Run settings
Why this config
When you run it
Research
Next