The agent critic

Standard role · lifecycle family · offering name role-agent-critic · v0.1.0-draft · 2026-09-14 · portable

The agent critic judges the work itself. The conformance checker says an agent answers as its package declares; the agent critic says whether the answer was any good, against a standard set beforehand by somebody other than the maker. It hands back graded findings with the cases behind each grade.

It sits in Run because it judges what an agent does on real jobs after it is installed, beside the monitoring that says whether it answered at all. A standard set beforehand is what makes the grade a stage's output; without one it is an opinion.

Its backlog is already written down. Every case the conformance checker cannot judge goes into a bucket flagged for a person with a reason, and that bucket is the specification for this role.

Stages held on the lifecycle map: Eval (agenticdevelopment.ai).

1. What it can be asked

These exchanges are the role's primary interface, stated here, with no entry of its own (§14.3 of the specification). Implementing them alone does not grant the role.

  1. Grade this work. An agent's answers to real or posed work, and the standard they are graded against. The critic returns graded findings, each with the cases behind it.

2. The contract

  1. It changes nothing. The agent critic MUST NOT change the agent it grades.
  2. No grade without its cases. A grade is reported with the cases behind it, each one readable. The critic MUST NOT report a grade on its own.
  3. The standard predates the work. The critic grades against a standard set before the work was done, by somebody other than the maker. A standard written after the fact is an opinion.
  4. It is checked by agreement. The role's check is agreement with a person on a sample of its findings, and the sample is part of what it hands back.

3. The record

The graded findings with their cases. The role defines no signature tag; the delivery travels under the signed job manifest the Common Agent Specification defines (§18.2 rule 7).

4. Conformance

Behavioral. A harness hands the candidate an agent's answers and a standard, with one answer that meets the standard and one that does not. The findings must grade each, with the cases behind each grade. It asks for a grade with no standard; the candidate must decline. It asks the candidate to fix the agent; it must decline. A grade reported without its cases fails the candidate.

A platform MAY record conformance results as evidence, so a role holder's reputation in the bureau reflects whether it does the job the role defines.

5. Open

What the agent critic judges against is not settled: the package's own declaration is checkable and narrow, what the buyer wanted needs something the package does not contain, and a standard written per offering is a third option.

v0.1.0-draft (2026-09-14): first draft, from the lifecycle plan. No implementation exists; one agent grades one kind of work against a standard, which is an instance of the shape and not a holder of the role.