When a study or agency uses AI to score student work or code open-ended responses, different models hand back different scores for the same response. That disagreement is a reliability threat to any finding built on those scores. This lab measures the instability on your own scores and reports a defensible bounded range, with a human reference anchoring the calibration.
Check whether essay scores from an AI scorer would hold up if a different model had been used.
Test how stable AI-assigned codes are for survey and interview responses before the analysis begins.
Compare candidate models for a scoring pipeline and document how sensitive results are to that choice.
One row per item (an essay, a short response, a portfolio). One column per model, two to four. Add human reference scores in the last column if you have blinded human ratings. Any numeric scale works, set the bounds below.
Paste comma-separated or tab-separated rows (a copied spreadsheet selection works), or choose a .csv or .txt file. Include a header row with an item column, your model columns, and an optional column named Human.
Each model scores the selected item. The shaded band is the range across models. The diamond is the human reference, shown when you provide one.
Rank your items by whichever model you trust, then compare each to its rank under the human reference. The same items move up and down depending only on the scorer.
| Rank | Item | Score | Range | vs human | Certainty |
|---|
The instability is the finding, not a bug. These methods turn it into measurement you can defend.
Instead of one AI score, report an interval reflecting what the models can and cannot agree on, so a ranking is asserted only when the data support it.
A blinded sample of human ratings anchors the AI scores and lets you estimate and correct model-specific bias.
Which model you use shifts the answer. Measure that sensitivity directly so findings do not quietly depend on a vendor choice.
Method draws on the Principal Investigator's research on multi-model measurement instability (NBER Working Paper 35110) and platform-selection effects in AI measurement (arXiv 2605.21743).
Package the loaded scores as a branded one-page PDF a reviewer, program officer, or vendor-selection committee can read in a minute. Built entirely from the data above, on this page.