Two made-up organizations, 60 past decisions each: Larkspur Garden Co.'s accounts payable (pay an invoice or hold it, and why) and the Hollis home-robotics lab (what the robot saw). Their histories are synthetic, and their labels follow cleaner rules than a real team's do; real histories will score lower. Every number below was measured on September 23 and 24, 2026, and every one can be reproduced from the PlatAtlas repository.
An AI decision model (JEV, from TypeSafe) judged the same 60 past decisions four times. Only the words describing each answer changed.
| What each answer said | Accounts payable, wrong of 60 | Robotics lab, wrong of 60 |
|---|---|---|
| Placeholders: "Recorded as X in N past cases" (what an import writes) | 13 (22%) | 16 (27%) |
| Drafted from rules counted in the history, held out: drafted from one half, graded on the other | 8 (13%) | 9 (15%) |
| Drafted from rules counted in the same rows it is graded on (flattering; shown for honesty) | 0 (0%) | 2 (3%) |
| The steward's own words | 1 (2%) | 0 (0%) |
Run to run, the same setup varied by one or two cases (JEV's placeholder run on accounts payable scored 15 and then 13).
What it says: descriptions drafted automatically from a team's own history cut the AI's mistakes by about half before anyone writes a word. A steward's own words take it the rest of the way. That is why the start page drafts them for you, and why the first job we give a steward is to replace them.
We first measured the drafts on the same rows they were drafted from, and they looked perfect. Held out, on cases they had never seen, they did worse than the placeholders on accounts payable, because a small history lets rules fit coincidences (an account number read as a quantity; the everyday answer described by two unrelated conditions). We fixed three things (no thresholds on identifiers, the everyday answer stays the default, rare answers keep plain words); every number above is after the fix. The start page's AI check now drafts from one half of your history and grades on the other.
The same comparison with GLiNER2, an open model running on a Raspberry Pi with no training, is next; it will be added here.
We planted wrong decisions at random in each history (20 repeats each) and counted how many the start page's "worth a second look" list found, and how many flags pointed at decisions that were in fact right. Counting only, no AI.
| Mistakes planted in 60 | Accounts payable: found | Robotics lab: found |
|---|---|---|
| 1 | 90% (18 of 20), 1.5 flags a run | 100% (20 of 20), 3.5 flags a run |
| 3 | 68% (41 of 60), 3.3 flags a run | 95% (57 of 60), 5.2 flags a run |
| 6 | 47% (56 of 120), 4.8 flags a run | 86% (103 of 120), 7.5 flags a run |
| None (flags on a clean history) | 0 flags | 3 flags |
With one mistake planted, 12 of 30 accounts-payable flags and 51 of 71 lab flags were right decisions that simply stand out (the lab has 3 such cases even with no mistakes). That is the trade we chose for a list a person reads in a minute: each flag says why it stands out, and a person decides. An earlier version that flagged only exceptions to the strictest rules had no false alarms but found only 9 to 25% of the mistakes.
| Engine | Wrong of 60 | Cost per 1,000 cases |
|---|---|---|
| JEV (TypeSafe) | 0 | about $0.04 |
| GLiNER2 trained on the lab's own labels (cross-validated: each case answered by a model that never saw it) | 0 | $0, on a Raspberry Pi |
| Claude Sonnet 5, effort low | 0 | $1.79 |
| Claude Haiku 4.5 | 4 | $1.49 |
| GLiNER2, untrained | 24 | $0 |
| The robot's own object detector | 32 | $0 |
In the PlatAtlas repository: scripts/experiments/descriptions.ts (section 1),
scripts/experiments/mistakes.mjs (section 2), and docs/results/2026-09-23-household/ (section 3). The counting
runs in your browser at /demo/start/, on your own spreadsheet, with nothing uploaded.
All model answers are inferred; what each team decided is the reference. Nothing here is a guarantee.