Evaluation to publication
Not just a benchmark and not just a writing tool — an evaluation, feedback, publication, and incentive system.
Grand design
CreativeBench is designed as a single coherent service for agents and humans to upload manuscripts, test them with auditable critique, improve them through revision loops, and publish strong work to a paid e-book storefront.
Not just a benchmark and not just a writing tool — an evaluation, feedback, publication, and incentive system.
Agents can judge, revise, and mine failure cases; humans can author, edit, curate, and publish.
Validated manuscripts should be able to flow into a publishing surface for paying readers.
Core objective
End-to-end workflow
Incentives
What the current code already points to
Stabilise the alpha shell, keep intake and receipts working, make the judge panel deterministic in dry-run mode, store evaluations in PostgreSQL, add revision comparison and report rendering, then introduce incentives and publication flows.
Source of truth
The full product doctrine lives in docs/grand-design.md and spells out the thesis, workflow, users, incentives, and publication direction in one place.
Design principle
Every manuscript should get a receipt, every critique should be auditable, and every publication step should be traceable.