observability-oss skills & plugins for progress-observability
Free tool · no signup · runs in your browser

LLM-as-a-Judge eval builder

Pick a failure mode and get a complete evaluator prompt — single criterion, pre-specified procedure steps, bias defenses, and the research behind each choice. Built on the same frame as the generate-eval skill; nothing you enter leaves this page.

1 · What failure mode should the judge catch?

2 · Is there a ground truth to check against?

A reference is a ground truth the judge can check against — retrieved chunks for RAG, gold SQL, a reference summary.

3 · One output, or comparing two?

4 · Bias defenses

Evaluator prompt

Run settings

Why this config

When you run it

    Research