A rights-cleared evaluation and correction dataset for creative-writing AI: 175 practitioner-reviewed contribution cards, 155 graduated rubric dimensions, and 28 auditable packages capturing how a 20-year screenwriting practitioner evaluated and corrected AI-assisted screenplay development work. Licensed for AI evaluation, fine-tuning, and benchmarking. A free sample is coming to Hugging Face and AWS Data Exchange; the full corpus is licensed from $299/month.
| Metric | Value |
|---|---|
| Packages | 28 |
| Files | 532 |
| Contribution cards | 175 (46 craft / 129 process) |
| Graduated rubric dimensions | 155 |
| Retrieval-index dimensions | 15 (10 craft + 5 process) |
| Failure types | 10 |
| Severity levels | 4 |
| Lifecycle stages | 5 |
| Coverage period | Feb 2023 – Sep 2025 |
| Evaluation validation | GEPA V4.1 — 70%→100% accuracy across two live screenplay tests; 0 overrides on the 18-score package benchmark |
The dataset's quality control is itself the proof of quality. The evaluation system was built with GEPA (Genetic-Pareto Evaluation Prompt Architecture): five evaluator variants tested against real material, and the selected Dual-Lens + Comparative Anchoring prompt then corrected by a 20-year screenwriting practitioner. On the first original-screenplay live test the evaluator scored 70% accuracy — three practitioner overrides, each with documented rationale. Those three corrections were built into the prompt, and on the second original-screenplay live test it scored 100% — zero overrides, with all three corrected dimensions scoring correctly. Both tests, all three corrections, and the 18-score package benchmark (0 overrides) ship as documented, auditable records.
Single practitioner, 100% owned, zero scraped content, and no third-party screenplay text reproduced anywhere in the corpus. Every file passes an audited chain of custody from source conversation to published package, and the corpus is de-identified with a documented, verified redaction pass.
Evaluating script-analysis AI. The 155 graduated rubric dimensions with per-level criteria give AI coverage and script-feedback tools a practitioner-calibrated standard to score against — not generic "good writing" heuristics but documented professional judgment with worked examples.
Quality control and red-teaming creative AI. The 10 failure types (voice drift, exposition dumps, tonal inconsistency, logic breaks and more) with severity and lifecycle classification make targeted test suites for where creative-writing AI actually fails.
Reward-model and LLM-judge calibration. The corrected evaluator prompt and its documented 70%→100% improvement provide both a calibration target and a case study in aligning AI judgment to expert human judgment.
Also from Meta-Flywheel Ventures: the California Civil Litigation Legal-AI Evaluation & Correction Dataset — a second rights-cleared judgment dataset from the same documented pipeline.