
Agent-as-a-Judge: Evaluate Agents With Agents
Scoring agent behavior with another agent works, but only if you treat the judge as a system component with its own failure modes.
Blake Aber · Predicate Ventures
Why judging agents is harder than judging outputs
Evaluating a single model response is a bounded problem. You have an input, an output, and a rubric. Compare, score, move on.
Agents break that framing. An agent takes multiple steps, calls tools, revises its plan, and produces a trajectory rather than a single answer. The final output can be correct while the path to it was wasteful, unsafe, or accidentally right.
That is the gap agent-as-a-judge fills. Instead of scoring only the last message, you use an evaluator agent that inspects the full trajectory: the reasoning, the tool calls, the intermediate states, and the final result.
The premise is simple. The agent doing the work is not the best-positioned observer of its own behavior. A separate agent, given the trace and a rubric, can catch errors the working agent cannot see.
What the judge actually reads
A useful judge does not receive a summary. It receives the artifact.
That means the full step log: each planning decision, each tool invocation with its arguments, each tool return value, and the state transitions in between. If you only hand the judge the final answer, you have rebuilt output evaluation and lost the reason for using an agent at all.
Concrete inputs to a trajectory judge:
- The original task specification and any constraints.
- The ordered list of steps the agent took.
- Tool call arguments and raw returns, not paraphrases.
- The final deliverable.
- A rubric that maps to the dimensions you care about.
The rubric matters more than the model. A judge with a vague prompt returns vague scores. A judge with explicit criteria and examples of pass and fail returns something you can act on.
Patterns that hold up in production
Per-step scoring vs. trajectory scoring
Two modes, different uses.
Per-step scoring evaluates each action in isolation: was this tool call valid, was the argument correct, did the step advance the task. This surfaces where a run went wrong.
Trajectory scoring evaluates the run as a whole: did the agent solve the task, was the path efficient, did it respect constraints. This tells you whether the run was good.
Run both. Per-step scoring drives debugging. Trajectory scoring drives the aggregate metrics you report and gate releases on.
Rubric decomposition
A single overall quality score hides too much. Break the rubric into named dimensions and score each separately.
Common dimensions: task completion, tool-use correctness, efficiency, constraint adherence, and safety. When completion is high but efficiency is low, you know the agent is solving the problem the expensive way. A single blended number would have masked that.
Reference-based and reference-free judging
When you have a known-good trajectory or a gold answer, give it to the judge. Reference-based judging is more reliable because the model compares against a fixed target instead of reasoning from scratch.
When you have no reference, which is most of the time in open-ended agent work, the judge scores against the rubric alone. Expect more variance here and account for it.
Pairwise comparison
Asking a judge to score one trajectory on an absolute scale is noisy. Asking it to choose between two trajectories is more stable.
Use pairwise comparison when you are ranking model versions or prompt variants. Use absolute scoring when you need a threshold to gate against.
Failure modes of the judge itself
The judge is an agent. It fails the way agents fail.
Position bias. In pairwise comparison, judges favor whichever trajectory appears first. Swap the order across runs and average. If the winner flips with position, the comparison is not real.
Length bias. Judges reward longer, more detailed trajectories even when brevity was correct. Watch for a correlation between token count and score that your rubric did not intend.
Self-preference. A judge built on the same model family as the worker tends to rate that family's output higher. Cross-family judging reduces this.
Rubric drift. Over many runs, the judge's interpretation of the rubric shifts subtly. Anchor it with fixed few-shot examples and periodically re-validate against human labels.
The last point is the one people skip. A judge that has never been checked against human judgment is an opinion, not a measurement. Calibrate against a labeled set, measure agreement, and re-check when you change the judge prompt or model.
Guardrails are a different job
A judge scores after the fact. A guardrail acts during the run.
Do not conflate them. Judging is evaluation for offline analysis and release gating. Guardrails are runtime checks that block unsafe tool calls, malformed arguments, or actions outside policy before they execute.
A guardrail can be an agent too, running as a fast inline check. But its job is a binary allow or deny on a specific action, not a graded assessment of the whole trajectory. Keep the two separate so a slow, thorough judge never sits in your latency path.
The cost problem
Judging every production run with a full trajectory judge is expensive. The judge reads more tokens than the worker produced, and it runs on top of the work you already paid for.
Control this deliberately.
- Sample. Judge a fraction of production traffic, not all of it, once you trust the metric.
- Tier the judge. Use a cheap judge for a first pass and escalate only ambiguous cases to a stronger, costlier one.
- Judge offline. Move full evaluation out of the request path into batch jobs against logged traces.
- Cache references. Reuse gold trajectories across runs instead of regenerating them.
The goal is a signal you can afford to collect continuously, not a one-time benchmark you run before launch and never again.
Where to start
Begin with a rubric written by the people who own the task, not by the model. Decompose it into named dimensions. Feed the judge full trajectories, not summaries.
Then calibrate against human labels before you trust a single number. Measure judge-human agreement, correct for position and length bias, and only then wire the metric into your release process.
Agents that evaluate agents give you a scalable way to watch behavior you cannot review by hand. They earn that trust only after you have treated the judge as carefully as the system it grades.