Agent evaluation · 2026

Every attempt gets graded. Nothing passes without evidence.

In a local agent orchestration app I built in 2026, agents could run inside a bounded outcome loop: each attempt was graded by deterministic checks and, where criteria need judgment, a rubric-based LLM judge. Failures went back as feedback until the work passed or a limit stopped the loop.

outcome loop
  1. contractwhat "done" means, as criteria
  2. actor runthe agent attempts the work
  3. checksdeterministic · zero tokens
  4. judgeLLM · one criterion per call · fails closed
  5. feedback→ next iteration
  6. stoppass · iteration cap · wall clock · token budget · cancel

Inside the loop, by design: no recursive agent spawning, no picking roles from prompt text, and no unmetered model call.

Bounded

Limits checked every step

An iteration cap, a wall-clock deadline where each run gets only the time left, and a token budget checked before every step and every judge call. A cancel never finalizes as success.

Deterministic first

Checks that cost nothing

Deliverables on disk, metric thresholds, assertions with no eval, and repo-local scripts that are path-jailed, time-boxed, and judged by exit code. None of them spends a token.

Scoped

Graded on this run only

Deliverable checks and the judge see only this iteration's new outputs, so a failed or empty run can't pass on an earlier run's work.

When there's no oracle

Some criteria need judgment, so a rubric-based LLM judge grades them. It is built to fail closed.

  • One criterion per call. A batched version failed in a live test run: the grading model reasoned correctly but returned a verdict nothing could parse, so every criterion failed. The fix was one criterion per call, plus one repair retry.
  • Independent of the agent. Each criterion is graded in a fresh context on its own model route, and the agent's output is fenced as untrusted data, with an instruction that it carries no authority.
  • Fails closed. An unparseable verdict after one repair, an unknown criterion, or a pass with no evidence all grade as failed. The overall pass/fail is recomputed from the per-criterion grades; the judge never sets it.
  • Reads the start. In a live run, grading only the end of a long brief produced a false failure on a well-grounded one, since briefs lead with their thesis and sources. The judge now reads a head-biased excerpt.
  • Metered. Every call is checked against the budget before it runs.

On your team

  • Put deterministic checks first. They're exact and free; save the model for what they can't judge.
  • Make the judge fail closed. An unparseable or unevidenced verdict is a failure, not a pass.
  • Bound every loop. Iterations, time, and tokens all get limits before the first run.

Need someone who verifies the automation?

I'm open to applied AI, automation, and forward-deployed roles.