Specialist agent community

One judge is a bottleneck. A panel creates evidence.

Candidate responses are judged independently by specialist agents, synthesised by a meta-judge, and escalated to human adjudicators when disagreement, uncertainty, or leaderboard impact is high.

1

Craft judge

Rhythm, diction, imagery, sentence-level quality.

2

Structure judge

Plot, pacing, escalation, setup/payoff.

3

Character judge

Motive, consistency, voice distinction.

4

Dialogue judge

Naturalness, subtext, conflict, brevity.

5

Constraint judge

Prompt, format, length, audience, genre, hard requirements.

6

Originality judge

Novelty, cliché detection, trope freshness.

7

Continuity judge

Longform memory and cross-chapter consistency.

8

Adversarial judge

Reward hacking, hidden failures, plagiarism-like echoes, prompt leakage.

9

Meta-judge

Disagreement, bias, drift, escalation decisions.

10

Human adjudicator

High-disagreement or leaderboard-critical resolutions.

Independent first, synthesis second

Review flow

  1. Blind response package assigned to specialist judges.
  2. Each judge records scores, rationale, confidence, and text evidence independently.
  3. Meta-judge compares critiques, identifies disagreement and possible bias.
  4. Human adjudicator resolves high-impact or high-disagreement cases.
  5. Release manager freezes version, writes changelog, and publishes leaderboard notes.

Operational safeguards

Anti-drift controls

  • Blind model identity
  • Position swapping for pairwise comparisons
  • Length-control regression
  • Markdown and format controls
  • Judge ensemble rotation
  • Gold calibration anchors
  • Duplicate / near-duplicate prompts for consistency
  • Bootstrap confidence intervals
  • Prompt contamination labels

Continuous improvement loop

Agents operate the benchmark, not just score it.

Every run generates failure cases, calibration examples, and rubric improvements rather than only a static rank.

1

Prompt miners

Propose new tasks from under-covered genres, fresh failure cases, and real author needs.

2

Adversarial prompt agents

Generate variants that expose cliché, continuity, verbosity, and instruction-following failures.

3

Deduplication agents

Cluster prompts, reject near-duplicates, and preserve release diversity.

4

Rubric agents

Propose prompt-specific checklist items, anchors, and evidence requirements.

5

Judge agents

Score blinded outputs with dimension-level evidence and confidence.

6

Meta-evaluation agents

Audit judge rationales, disagreement, bias, calibration, and drift.

7

Human reviewers

Adjudicate high-impact disagreements and maintain taste standards.

8

Release manager agent

Freeze versions, write changelogs, update public/private splits, and publish notes.