EduWork.AI · Evaluation Quality Assurance

Can you trust an AI score? Measure it.

When a study or agency uses AI to score student work or code open-ended responses, different models hand back different scores for the same response. That disagreement is a reliability threat to any finding built on those scores. This lab measures the instability on your own scores and reports a defensible bounded range, with a human reference anchoring the calibration.

Works in your browser. Enter scores from two or more models in the grid below, or load the example to try it instantly. The page computes the per-item range, an overall stability figure, the ranking under each model, and a bounded range you can report, then packages it as a one-page reliability report. This is the human-oversight mission in practice. Everything runs on this page. Nothing is uploaded.
1
Enter or load scoresType scores from two or more models into the grid, paste a spreadsheet selection, or load the example.
2
See the spread and ranking flipsThe lab shows each item's score range across models and which items change rank depending on the scorer.
3
Report the bounded rangeCite the range, not one number, and keep a human in the loop to anchor and check the calibration.
Where you would use this

AI-scored writing assessments

Check whether essay scores from an AI scorer would hold up if a different model had been used.

AI coding of open-ended responses

Test how stable AI-assigned codes are for survey and interview responses before the analysis begins.

Vendor model selection

Compare candidate models for a scoring pipeline and document how sensitive results are to that choice.

Enter your scores

One row per item (an essay, a short response, a portfolio). One column per model, two to four. Add human reference scores in the last column if you have blinded human ratings. Any numeric scale works, set the bounds below.

to
Paste or upload a table instead

Paste comma-separated or tab-separated rows (a copied spreadsheet selection works), or choose a .csv or .txt file. Include a header row with an item column, your model columns, and an optional column named Human.

Score one item across models

Each model scores the selected item. The shaded band is the range across models. The diamond is the human reference, shown when you provide one.

The ranking flips with the model

Rank your items by whichever model you trust, then compare each to its rank under the human reference. The same items move up and down depending only on the scorer.

RankItemScoreRangevs humanCertainty

What this measures, and what to do about it

The instability is the finding, not a bug. These methods turn it into measurement you can defend.

Report bounds, not a point

Instead of one AI score, report an interval reflecting what the models can and cannot agree on, so a ranking is asserted only when the data support it.

Calibrate to human reference

A blinded sample of human ratings anchors the AI scores and lets you estimate and correct model-specific bias.

Account for model selection

Which model you use shifts the answer. Measure that sensitivity directly so findings do not quietly depend on a vendor choice.

Method draws on the Principal Investigator's research on multi-model measurement instability (NBER Working Paper 35110) and platform-selection effects in AI measurement (arXiv 2605.21743).

Reliability report

Package the loaded scores as a branded one-page PDF a reviewer, program officer, or vendor-selection committee can read in a minute. Built entirely from the data above, on this page.

One Letter page. In the print dialog choose "Save as PDF".
More EduWork.AI tools All tools →