PlatAtlasTry it with your spreadsheet

What we measured

Two made-up organizations, 60 past decisions each: Larkspur Garden Co.'s accounts payable (pay an invoice or hold it, and why) and the Hollis home-robotics lab (what the robot saw). Their histories are synthetic, and their labels follow cleaner rules than a real team's do; real histories will score lower. Every number below was measured on September 23 and 24, 2026, and every one can be reproduced from the PlatAtlas repository.

1. An answer's words matter more than anything else

An AI decision model (JEV, from TypeSafe) judged the same 60 past decisions four times. Only the words describing each answer changed.

What each answer saidAccounts payable, wrong of 60Robotics lab, wrong of 60
Placeholders: "Recorded as X in N past cases" (what an import writes)13 (22%)16 (27%)
Drafted from rules counted in the history, held out: drafted from one half, graded on the other8 (13%)9 (15%)
Drafted from rules counted in the same rows it is graded on (flattering; shown for honesty)0 (0%)2 (3%)
The steward's own words1 (2%)0 (0%)

Run to run, the same setup varied by one or two cases (JEV's placeholder run on accounts payable scored 15 and then 13).

What it says: descriptions drafted automatically from a team's own history cut the AI's mistakes by about half before anyone writes a word. A steward's own words take it the rest of the way. That is why the start page drafts them for you, and why the first job we give a steward is to replace them.

We first measured the drafts on the same rows they were drafted from, and they looked perfect. Held out, on cases they had never seen, they did worse than the placeholders on accounts payable, because a small history lets rules fit coincidences (an account number read as a quantity; the everyday answer described by two unrelated conditions). We fixed three things (no thresholds on identifiers, the everyday answer stays the default, rare answers keep plain words); every number above is after the fix. The start page's AI check now drafts from one half of your history and grades on the other.

The same comparison with GLiNER2, an open model running on a Raspberry Pi with no training, is next; it will be added here.

2. Flags that catch people's mistakes

We planted wrong decisions at random in each history (20 repeats each) and counted how many the start page's "worth a second look" list found, and how many flags pointed at decisions that were in fact right. Counting only, no AI.

Mistakes planted in 60Accounts payable: foundRobotics lab: found
190% (18 of 20), 1.5 flags a run100% (20 of 20), 3.5 flags a run
368% (41 of 60), 3.3 flags a run95% (57 of 60), 5.2 flags a run
647% (56 of 120), 4.8 flags a run86% (103 of 120), 7.5 flags a run
None (flags on a clean history)0 flags3 flags

With one mistake planted, 12 of 30 accounts-payable flags and 51 of 71 lab flags were right decisions that simply stand out (the lab has 3 such cases even with no mistakes). That is the trade we chose for a list a person reads in a minute: each flag says why it stands out, and a person decides. An earlier version that flagged only exceptions to the strictest rules had no false alarms but found only 9 to 25% of the mistakes.

3. Decision models on the robotics lab

EngineWrong of 60Cost per 1,000 cases
JEV (TypeSafe)0about $0.04
GLiNER2 trained on the lab's own labels (cross-validated: each case answered by a model that never saw it)0$0, on a Raspberry Pi
Claude Sonnet 5, effort low0$1.79
Claude Haiku 4.54$1.49
GLiNER2, untrained24$0
The robot's own object detector32$0

How to reproduce

In the PlatAtlas repository: scripts/experiments/descriptions.ts (section 1), scripts/experiments/mistakes.mjs (section 2), and docs/results/2026-09-23-household/ (section 3). The counting runs in your browser at /demo/start/, on your own spreadsheet, with nothing uploaded.

All model answers are inferred; what each team decided is the reference. Nothing here is a guarantee.