Professional Screenwriting AI Evaluation & Correction Dataset

A rights-cleared evaluation and correction dataset for creative-writing AI: 175 practitioner-reviewed contribution cards, 155 graduated rubric dimensions, and 28 auditable packages capturing how a 20-year screenwriting practitioner evaluated and corrected AI-assisted screenplay development work. Licensed for AI evaluation, fine-tuning, and benchmarking. A free sample is coming to Hugging Face and AWS Data Exchange; the full corpus is licensed from $299/month.

Dataset statistics

MetricValue
Packages28
Files532
Contribution cards175 (46 craft / 129 process)
Graduated rubric dimensions155
Retrieval-index dimensions15 (10 craft + 5 process)
Failure types10
Severity levels4
Lifecycle stages5
Coverage periodFeb 2023 – Sep 2025
Evaluation validationGEPA V4.1 — 70%→100% accuracy across two live screenplay tests; 0 overrides on the 18-score package benchmark

Provenance: the documented judgment story

The dataset's quality control is itself the proof of quality. The evaluation system was built with GEPA (Genetic-Pareto Evaluation Prompt Architecture): five evaluator variants tested against real material, and the selected Dual-Lens + Comparative Anchoring prompt then corrected by a 20-year screenwriting practitioner. On the first original-screenplay live test the evaluator scored 70% accuracy — three practitioner overrides, each with documented rationale. Those three corrections were built into the prompt, and on the second original-screenplay live test it scored 100% — zero overrides, with all three corrected dimensions scoring correctly. Both tests, all three corrections, and the 18-score package benchmark (0 overrides) ship as documented, auditable records.

Single practitioner, 100% owned, zero scraped content, and no third-party screenplay text reproduced anywhere in the corpus. Every file passes an audited chain of custody from source conversation to published package, and the corpus is de-identified with a documented, verified redaction pass.

What it's for

Evaluating script-analysis AI. The 155 graduated rubric dimensions with per-level criteria give AI coverage and script-feedback tools a practitioner-calibrated standard to score against — not generic "good writing" heuristics but documented professional judgment with worked examples.

Quality control and red-teaming creative AI. The 10 failure types (voice drift, exposition dumps, tonal inconsistency, logic breaks and more) with severity and lifecycle classification make targeted test suites for where creative-writing AI actually fails.

Reward-model and LLM-judge calibration. The corrected evaluator prompt and its documented 70%→100% improvement provide both a calibration target and a case study in aligning AI judgment to expert human judgment.

Frequently asked questions

Is this dataset rights-cleared for commercial AI use?
Yes. The corpus derives entirely from the practitioner's own development work and original projects, owned 100% by the provider. It contains no scraped content and no third-party screenplay text, so there are no upstream rights to clear.
Does it contain scripts from movies or shows?
No. Published works appear only as titles in analysis and critique context. The dataset is judgment data — evaluations, corrections, rubrics, and workflow specifications — not a screenplay corpus.
How was the evaluation system validated?
Through GEPA prompt-variant testing followed by practitioner correction: 70% accuracy on the first original-screenplay live test, three documented corrections, then 100% accuracy on the second live test, plus an 18-score package benchmark with zero overrides. The corrected prompt and both test records ship with the dataset.
What formats are the files in?
Markdown, JSON, and CSV: contribution cards, graduated rubrics, machine-readable eval schemas, and a queryable retrieval index classified by dimension, failure type, severity, and lifecycle stage. Delivered as a ZIP through AWS Data Exchange.
Is there a free sample?
Yes — a gated sample (15 contribution cards spanning craft and process types, one complete graduated rubric, and its eval schema) is publishing on Hugging Face, with a free sample also planned for AWS Data Exchange.
How is it licensed and priced?
Licensed from $299/month through AWS Data Exchange. AWS handles billing, delivery, and entitlement; specific terms and durations are set out in the AWS Data Exchange offer. Model-training rights are licensed separately.

License & access

Also from Meta-Flywheel Ventures: the California Civil Litigation Legal-AI Evaluation & Correction Dataset — a second rights-cleared judgment dataset from the same documented pipeline.