PlatAtlas

How it works, technically: a home-robotics lab

About two minutes, no sound needed. The lab is made up and its detector is simulated; every number on screen was measured.

Read it instead

A home robot sees a small container by the spice rack. Its detector is 95% sure it is a spice jar. It is medicine: a child-resistant cap and a pharmacy label. Being sure is not being right.

The lab keeps a record: plain files in its own repository. The question, each answer in the lab lead's words, who may decide what, the settings, 60 past labels, and everything waiting and signed.

Every model answers the same 60 past cases. JEV, Claude Sonnet 5 and GLiNER2 trained on the lab's own labels got none wrong; Claude Haiku got 4; untrained GLiNER2 got 24; the robot's own detector got 32.

Models can be combined: cheap ones first, stronger ones when unsure, people where the lab says so. The measure chooses which combination, not taste.

Rae, an annotator, labels blind: she never sees a machine's guess, so her labels can retrain the detector. Her inbox layout is hers, built by asking her assistant.

Settings are layered: the organization, the role, the person, and the person's agent, which can only be narrower. A locked setting becomes a request the lead approves or declines.

People talk about a case with cc and bcc, tagging each other, the assistant and other cases. Nothing written there signs anything.

Each signature keeps what the person was shown, which models ran and which settings applied. Mistakes and bad data are flagged, never erased, and the lab's labels become a training set.

Everything regenerates except the record: models, combinations, inboxes and layouts are measured, swapped and rebuilt; the questions, rules and past decisions stay the organization's own.

Measured on a synthetic history whose labels follow clean rules; a real lab's cases will be harder. Details: the results in the PlatAtlas repository (docs/results/2026-09-23-household).