Public development
Public prompts, rubrics, outputs, and judge rationales for reproducible debugging.
creativebench_dev
Benchmark suite
CreativeBench v0.1 translates the design doc into public, private, sealed, and rolling tracks covering short fiction, longform, dialogue, revision, style control, constraints, humour, brainstorming, and roleplay.
Public prompts, rubrics, outputs, and judge rationales for reproducible debugging.
creativebench_dev
Frozen public releases such as creativebench_v1.0 for transparent rankings and regression tracking.
creativebench_v1.0
Hidden prompts and final judging data for contamination resistance and anti-gaming.
sealed_v1
Fresh prompts generated after major model cutoffs to test current capability.
live_monthly
Prompt coverage
These categories prevent a model from winning by overfitting to one convenient demo style.
Longform staging
Longform capability is measured as a staged process with planning, generation, continuity checking, revision, and whole-story judgment.
MVP build pack
Enough to improve on current writing leaderboards while staying practical to run and refine.
Represented in the app information architecture and ready to become executable backend work.
Represented in the app information architecture and ready to become executable backend work.
Represented in the app information architecture and ready to become executable backend work.
Represented in the app information architecture and ready to become executable backend work.
Represented in the app information architecture and ready to become executable backend work.
Represented in the app information architecture and ready to become executable backend work.
Represented in the app information architecture and ready to become executable backend work.
Represented in the app information architecture and ready to become executable backend work.
Represented in the app information architecture and ready to become executable backend work.