Workplace promptquality benchmark.
A reproducible plan for examining recurring instruction failures across representative workplace tasks. No benchmark result appears here until the corpus, scoring, review, and limitations are complete.
Research question
Where does usable intent disappear before AI receives it?
The benchmark will describe patterns in the sampled prompts. It will not claim that the sample represents every worker, company, model, or industry.
Outcome clarity
Score 0–2 using a published anchor and retain reviewer notes for disagreements.
Audience and context
Score 0–2 using a published anchor and retain reviewer notes for disagreements.
Input sufficiency
Score 0–2 using a published anchor and retain reviewer notes for disagreements.
Constraints and exclusions
Score 0–2 using a published anchor and retain reviewer notes for disagreements.
Output contract
Score 0–2 using a published anchor and retain reviewer notes for disagreements.
Evidence and uncertainty
Score 0–2 using a published anchor and retain reviewer notes for disagreements.
Sensitive-data handling
Score 0–2 using a published anchor and retain reviewer notes for disagreements.
Human review and next action
Score 0–2 using a published anchor and retain reviewer notes for disagreements.
Reproducible protocol
The study can be rerun, challenged, and improved.
1 · Corpus
Collect 50–100 synthetic or permissioned prompts across common work domains. Record source type and selection rule; exclude confidential content.
2 · Blind scoring
Two reviewers score de-identified prompts against the eight published dimensions without seeing author identity or expected product impact.
3 · Agreement
Report raw agreement and dimension-level disagreement. Resolve only after preserving both original scores.
4 · Analysis
Report distributions and recurring omissions. Separate observed patterns from interpretations and product hypotheses.
5 · Holdout
Reserve prompts not used while refining the rubric, then apply the frozen rubric to the holdout set.
6 · Release
Publish the corpus where permission allows, scoring anchors, code or worksheet, limitations, and a dated reproducibility record.
Dataset description
Task mix, source type, inclusion and exclusion rules, deduplication, and the final item count.
Reporting
Dimension distributions, agreement, recurring failures, uncertainty, and a machine-readable score file.
Safeguards
No private customer prompts without permission, no fabricated findings, and no identifying content in released examples.
Limits disclosed in advance
A benchmark is evidence about its sample—not a universal law.
- Domain and language coverage will be incomplete.
- Synthetic prompts may not reproduce every real workplace constraint.
- Rubric scores contain human judgment even with anchors and dual review.
- The study will measure instruction quality, not guarantee downstream model accuracy.
- BeforePrompt’s involvement creates a conflict of interest that must be disclosed with the results.