ExploringMethodology published · Results not yet claimed

Workplace promptquality benchmark.

A reproducible plan for examining recurring instruction failures across representative workplace tasks. No benchmark result appears here until the corpus, scoring, review, and limitations are complete.

Research question

Where does usable intent disappear before AI receives it?

The benchmark will describe patterns in the sampled prompts. It will not claim that the sample represents every worker, company, model, or industry.

01

Outcome clarity

Score 0–2 using a published anchor and retain reviewer notes for disagreements.

02

Audience and context

Score 0–2 using a published anchor and retain reviewer notes for disagreements.

03

Input sufficiency

Score 0–2 using a published anchor and retain reviewer notes for disagreements.

04

Constraints and exclusions

Score 0–2 using a published anchor and retain reviewer notes for disagreements.

05

Output contract

Score 0–2 using a published anchor and retain reviewer notes for disagreements.

06

Evidence and uncertainty

Score 0–2 using a published anchor and retain reviewer notes for disagreements.

07

Sensitive-data handling

Score 0–2 using a published anchor and retain reviewer notes for disagreements.

08

Human review and next action

Score 0–2 using a published anchor and retain reviewer notes for disagreements.

Reproducible protocol

The study can be rerun, challenged, and improved.

  1. 1 · Corpus

    Collect 50–100 synthetic or permissioned prompts across common work domains. Record source type and selection rule; exclude confidential content.

  2. 2 · Blind scoring

    Two reviewers score de-identified prompts against the eight published dimensions without seeing author identity or expected product impact.

  3. 3 · Agreement

    Report raw agreement and dimension-level disagreement. Resolve only after preserving both original scores.

  4. 4 · Analysis

    Report distributions and recurring omissions. Separate observed patterns from interpretations and product hypotheses.

  5. 5 · Holdout

    Reserve prompts not used while refining the rubric, then apply the frozen rubric to the holdout set.

  6. 6 · Release

    Publish the corpus where permission allows, scoring anchors, code or worksheet, limitations, and a dated reproducibility record.

Dataset description

Task mix, source type, inclusion and exclusion rules, deduplication, and the final item count.

Reporting

Dimension distributions, agreement, recurring failures, uncertainty, and a machine-readable score file.

Safeguards

No private customer prompts without permission, no fabricated findings, and no identifying content in released examples.

Limits disclosed in advance

A benchmark is evidence about its sample—not a universal law.

  • Domain and language coverage will be incomplete.
  • Synthetic prompts may not reproduce every real workplace constraint.
  • Rubric scores contain human judgment even with anchors and dual review.
  • The study will measure instruction quality, not guarantee downstream model accuracy.
  • BeforePrompt’s involvement creates a conflict of interest that must be disclosed with the results.