Classification
2- Complexity
- Medium
- Impact area
- Technical
A model evaluates outputs against criteria, rubrics, or comparative examples — scoring one output at a time (pointwise) or picking the better of two (pairwise, more robust to score drift).
360° overview
Six perspectives place the building block in context. The numbers show where each perspective continues in the reading path.
The building block at a glance
LLM-as-Judge
360°
Integrations
LLM-as-judge evaluators over datasets.
+2
A model evaluates outputs against criteria, rubrics, or comparative examples — scoring one output at a time (pointwise) or picking the better of two (pairwise, more robust to score drift). Structuring the judgment as chain-of-thought reasoning before a form-filled score (the G-Eval paradigm) improves alignment with human ratings.
Value stream stage: Build