What is LLM-as-a-judge?
LLM-as-a-judge is an evaluation setup in which a large language model (LLM) grades output against written criteria. The grading model is also called an LLM judge or an LLM evaluator. Teams use a judge when no fixed expected value can capture what counts as good, e.g. whether a summary is accurate.
Judges are often part of an eval. An eval, short for evaluation, is a repeatable test of an AI system that runs a fixed set of tasks and scores the results, so versions can be compared. A judge can score those results in place of checks written in code, or alongside them.
In a 2023 paper, Zheng and colleagues present LLM judges as a scalable way to approximate human preferences, which are expensive to collect. In their tests on answers from chat assistants, the strongest judges agreed with human preferences about as often as people agreed with each other.
When a team checks AI-generated code, a judge can grade a coding agent's diff or summary. The judge reads that text and runs nothing.
How does an LLM judge work?
A judge fits into an eval harness, the program that runs the tasks, in four steps:
- A person writes a rubric, the criteria and score scale that the judge applies.
- The harness builds a prompt from the rubric, the task, and the output to grade.
- The judge returns a score or a choice, usually with a written reason.
- The harness reads each score and combines the scores across tasks into one result per version.
Zheng and colleagues describe three setups, which teams use alone or together:
- Pairwise comparison. The judge reads one question and two answers, then picks the better answer or declares a tie.
- Single answer grading. The judge gives one answer a score on its own.
- Reference-guided grading. The prompt also holds a reference solution, which helps when a question has one right answer, e.g. a math problem.
When an agent gets several attempts at each task, the harness can combine the judge's grades as pass@k, the chance that at least one of k attempts succeeds. Anthropic's eval guide pairs it with pass^k, the chance that every one of k attempts succeeds, which rewards consistency. A judge that grades an attempt wrongly changes both measures.
What is an example of an LLM judge?
Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme builds an eval to compare two setups of a coding agent. One task asks the agent to "Let customers edit their delivery address during checkout." The eval harness sends the change from each setup to a judge with this rubric:
Grade the change against the request.
5: implements the request and changes nothing else.
3: implements part of the request.
1: does not implement the request, or breaks other behavior.
Reply in JSON with "reason" and "score".
For one setup, the judge receives the request, the diff, and the agent's summary, which says "All tests pass." The diff adds an address form, saves the address, and reloads the cart after the address changes. The judge returns this:
{
"reason": "Adds the address form, saves the address, and reloads the cart.",
"score": 5
}
Next, the developer starts the store with that change, adds 2 items to the cart, starts checkout, and changes the delivery address to 12 Elm Street. The address saved, but the cart emptied. The judge scored the text of the change, and the text looked right. The failure appeared only when the code ran with items in the cart.
An end-to-end test that adds 2 items, changes the address, and expects 2 items in the cart fails on the same code. This example is simplified. A real eval would grade many tasks and compare the judge's scores with a person's grades on a sample.
What biases do LLM judges have?
Studies of LLM judges report these weaknesses:
- Position bias. In a pairwise comparison, the judge picks an answer partly for its place in the prompt, so swapping the two answers can change the winner.
- Verbosity bias. The judge tends to score a longer answer higher, even when it is no clearer or more accurate than a shorter one.
- Self-preference bias. Research shows that models acting as judges tend to give their own output higher ratings. Zheng and colleagues call it "self-enhancement bias." The same effect can apply when an agent is asked to check its own code.
- Limited reasoning. A judge can accept a wrong answer to a math problem that the same model solves correctly when asked on its own.
- Variation between runs. The same judge can give the same output different scores on two runs.
Careful setup reduces these biases but does not remove them. Zheng and colleagues ran each pairwise comparison twice with the answers swapped, and counted a win only when both orders agreed. Anthropic's guide suggests grading each part of a task with a separate judge, and letting the judge answer "Unknown" when it lacks information. It also says to calibrate the judge against grades from human experts.
Asking for the reason before the score, and using a judge from a different model than the one that wrote the output, can also help. Rerunning the judge on the same outputs shows how much its scores vary.
Can an LLM judge whether software works?
A plain LLM judge cannot observe whether software works, because it never starts the software. Its grade of code, a diff, or a transcript is a prediction from text. AI code review that only reads the diff has the same limit, the gap between reading code and running code.
Anthropic's guide separates an agent's transcript, the record of what the agent did, from the outcome, "the final state in the environment at the end of the trial." A judge that reads an agent's summary grades the transcript. For coding agents, the guide recommends grading the outcome with code, e.g. by running the tests. It suggests judges with clear rubrics for what tests cannot score, e.g. overall code quality.
A judge can take part in a run when an agent harness gives it tools, e.g. a shell. This setup is sometimes called "agent as a judge." Its grade then rests on evidence from the run, e.g. a nonzero exit code. When an agent does exploratory testing, a judge can grade the recorded steps and results instead of the agent's summary. An agent hook that runs the tests when a coding agent stops gives a result that no model wrote.
How is an LLM judge different from a deterministic check?
A deterministic check is code that gives the same answer for the same input, e.g. an assertion that the cart holds 2 items. A judge and a deterministic check can each serve as a test oracle, the source of truth that a result is checked against. They differ on these points:
| Point | LLM judge | Deterministic check |
|---|---|---|
| Criteria | Prose in a rubric | Code, e.g. an assertion |
| Same input, same answer | Not always | Yes, but a flaky test breaks this |
| Cost per check | One or more model calls | Low once written |
| Grades well | Qualities with no single right answer | Exact values, states, and exit codes |
| Main weakness | Biases and variation between runs | Rejects valid output in an unexpected form |
When no exact expected output exists, code can often check a relation between runs, e.g. with metamorphic testing. Many eval setups use code for pass or fail and a judge for what code cannot score.
Who should decide whether a change works?
A person on the team should decide, using evidence from a run of the software. A judge's score can point a reviewer to a weak change.
A plain LLM judge grades text, and a run shows what the software does. RunStory independently runs the software against your change and returns evidence to the coding agent. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.
FAQs
What is an eval in AI?
An eval in AI, short for evaluation, is a repeatable test of an AI system that runs a fixed set of tasks and scores the results. Running the same eval on two versions of a system shows which one does better. An LLM judge is one way to score results that code cannot check.
What is pass@k?
Pass@k is an eval measure of the chance that an agent succeeds at least once in k attempts at a task. The related measure pass^k is the chance that every one of k attempts succeeds. When a judge grades the attempts, its errors change both measures.
What is a rubric for an LLM judge?
A rubric for an LLM judge is the written set of criteria and score levels that the judge applies to each output. A clear rubric says what earns each score, so a person can grade the same outputs and compare the two sets of grades.
How do you check that a judge is reliable?
Checking that a judge is reliable starts with outputs that people have graded by hand. Compare the judge's scores with those grades, rerun it on the same outputs to measure variation, and swap the answer order in pairwise setups.