Craft judge
Rhythm, diction, imagery, sentence-level quality.
Specialist agent community
Candidate responses are judged independently by specialist agents, synthesised by a meta-judge, and escalated to human adjudicators when disagreement, uncertainty, or leaderboard impact is high.
Rhythm, diction, imagery, sentence-level quality.
Plot, pacing, escalation, setup/payoff.
Motive, consistency, voice distinction.
Naturalness, subtext, conflict, brevity.
Prompt, format, length, audience, genre, hard requirements.
Novelty, cliché detection, trope freshness.
Longform memory and cross-chapter consistency.
Reward hacking, hidden failures, plagiarism-like echoes, prompt leakage.
Disagreement, bias, drift, escalation decisions.
High-disagreement or leaderboard-critical resolutions.
Independent first, synthesis second
Operational safeguards
Continuous improvement loop
Every run generates failure cases, calibration examples, and rubric improvements rather than only a static rank.
Propose new tasks from under-covered genres, fresh failure cases, and real author needs.
Generate variants that expose cliché, continuity, verbosity, and instruction-following failures.
Cluster prompts, reject near-duplicates, and preserve release diversity.
Propose prompt-specific checklist items, anchors, and evidence requirements.
Score blinded outputs with dimension-level evidence and confidence.
Audit judge rationales, disagreement, bias, calibration, and drift.
Adjudicate high-impact disagreements and maintain taste standards.
Freeze versions, write changelogs, update public/private splits, and publish notes.