Benchmark suite

The full creative-writing skill surface.

CreativeBench v0.1 translates the design doc into public, private, sealed, and rolling tracks covering short fiction, longform, dialogue, revision, style control, constraints, humour, brainstorming, and roleplay.

01

Public development

Public prompts, rubrics, outputs, and judge rationales for reproducible debugging.

Release lane

creativebench_dev

Open design
02

Public leaderboard

Frozen public releases such as creativebench_v1.0 for transparent rankings and regression tracking.

Release lane

creativebench_v1.0

Versioned
03

Private sealed

Hidden prompts and final judging data for contamination resistance and anti-gaming.

Release lane

sealed_v1

Controlled
04

Rolling live

Fresh prompts generated after major model cutoffs to test current capability.

Release lane

live_monthly

Continuous

Prompt coverage

Ten creative categories.

These categories prevent a model from winning by overfitting to one convenient demo style.

Short fiction scenes and flash fictionLongform narrative planning and executionDialogue with subtext, conflict, and voice distinctionCharacterisation, motives, and behavioural consistencyStyle control without living-author imitationRevision and editorial improvementConstraint writing: format, length, POV, audience, forbidden tropesPoetry, lyric, humour, and prose-adjacent formsBrainstorming: premises, twists, titles, and taglinesRoleplay and interactive fiction with persistent world state

Longform staging

Do not ask for one giant answer.

Longform capability is measured as a staged process with planning, generation, continuity checking, revision, and whole-story judgment.

  1. Premise expansion
  2. Cast and motive map
  3. Plot / turning-point plan
  4. Chapter-level plan
  5. Chapter generation
  6. Continuity check
  7. Final revision
  8. Whole-story judging

MVP build pack

Concrete first release scope.

Enough to improve on current writing leaderboards while staying practical to run and refine.

v0.1

80 public prompts across 8 creative categories

Represented in the app information architecture and ready to become executable backend work.

v0.1

40 private sealed prompts

Represented in the app information architecture and ready to become executable backend work.

v0.1

10 longform staged prompts

Represented in the app information architecture and ready to become executable backend work.

v0.1

5 specialist judge agents plus one meta-judge

Represented in the app information architecture and ready to become executable backend work.

v0.1

Pairwise position-swapped comparisons against fixed baselines

Represented in the app information architecture and ready to become executable backend work.

v0.1

Visible subdimension rubric scoring

Represented in the app information architecture and ready to become executable backend work.

v0.1

Slop, repetition, and length diagnostics

Represented in the app information architecture and ready to become executable backend work.

v0.1

Bootstrap confidence intervals

Represented in the app information architecture and ready to become executable backend work.

v0.1

Human adjudication for high-disagreement and leaderboard-critical cases

Represented in the app information architecture and ready to become executable backend work.