How the score is built
Readable maths, named sources, and the limits of what a reading can tell you.
1. You describe the work, not the job title
Two people with the same title do different work. The reading is built from the 3 to 8 tasks you actually do, in your words. A job title alone tells the model where to look; the tasks are what gets rated.
2. Each task is rated by an AI model
An Azure OpenAI model (gpt-5-mini) reads each task and assigns one of three levels, with a confidence:
- Protected: little current automation pressure. The task leans on judgment, relationships, physical context, accountability or taste.
- Watching: partly automatable today. AI can draft or assist, a person still owns the result.
- Exposed: a current AI capability plausibly does the core of the task today. The model must name that capability; if it cannot, the rating is automatically lowered to Watching.
Task text is treated strictly as data. Text that tries to instruct the rater is ignored, flagged, and rated at low confidence.
3. The score is fixed arithmetic
The model never produces the number. The score (rubric v1) is the average of the task levels, where Protected counts 0, Watching counts 50 and Exposed counts 100, every task weighing the same. That is why a re-run on the same tasks is comparable: only the levels can move. The results page shows how many points each task adds.
If you enter hours per task in the capacity calculator, a second, time-weighted figure is shown alongside, so a task that fills your week counts for more than one you touch monthly. The canonical score stays unweighted so that readings remain comparable over time.
4. Bands
The score is also shown as a band, because a single number implies more precision than exists:
| Score | Band | Reads as |
|---|---|---|
| 0–24 | Low overlap | Most of your tasks are protected today. |
| 25–49 | Moderate overlap | Some tasks are worth watching; AI can assist at the edges. |
| 50–74 | Elevated overlap | A real share of your tasks can be AI-drafted now; you review and own them. |
| 75–100 | High overlap | Most of your tasks can be AI-drafted today; your value is in directing and judging. |
5. Where the capability references come from
Today, from the model's own knowledge, and they are labelled as unverified on the page. They are a pointer to what to look up, not a citation. The next version grounds them in a dated, public evidence corpus: the MIT Work Analytics Laboratory's evidence-grounded exposure scores for the 18,796 task statements in O*NET (MIT licence), refreshed monthly, which also allows a reading to say what changed and why.
6. Exposure is not replacement
An Exposed task is one where AI can produce a first draft now. In practice a person still directs the tool, checks the output and owns the result; measured time savings are routinely smaller than people estimate, and part of the saved time goes back into review. The learning paths on each task describe that working mode. The score says nothing about whether a job will exist, and it is never shared with an employer.
7. Limits
- The model can mis-rate a task it does not understand. Confidence is shown so you can weigh it.
- References are unverified until the evidence corpus is wired in.
- Role buckets in the overview are built from job titles with a simple normaliser, not a taxonomy.
- The reading reflects capabilities as the model knows them; new tools ship monthly, which is why re-running matters.
Sources and attribution
- O*NET OnLine, U.S. Department of Labor, task statements (CC BY 4.0), the taxonomy the next version maps to.
- MIT Work Analytics Laboratory, evidence-grounded AI exposure dataset (MIT licence).
- Eloundou et al. (2024), Science, on task-level exposure; METR (2025) on measured developer productivity with AI.