Rubric score
Calibrated 0–100 dimension scores with concrete text evidence.
Method v0.1
Primary public rank is a blinded pairwise preference score with confidence intervals. Rubrics remain visible as the explanation layer, while diagnostics show verbosity, repetition, slop, uncertainty, and drift.
Calibrated 0–100 dimension scores with concrete text evidence.
Bradley-Terry / Elo / Glicko-style ranking from blinded comparisons with position swaps.
Non-quality metrics: slop, repetition, vocabulary control, length, uncertainty, judge disagreement, and drift.
Rubric dimensions
Every score is explainable and tied to a creative-writing failure or success mode.
Negative criteria
Bias controls
CreativeBench records initial draft → panel critique → standard editorial brief or own-critique revision → original-vs-revision judgment → revision lift and regression risk. It tracks whether the rewrite improves craft without breaking constraints, flattening voice, adding slop, or losing emotional core.