The machinery works. The questions didn't deserve you.
Placement ran on items derived from public LLM benchmarks. Audit that bank and its provenance falls apart: scraped test-prep, pipelined into a benchmark, deployed because it was convenient. Nobody along that chain ever asked whether an item was fit to measure a person. Roughly one in twelve to one in twenty knowledge items is outright broken — and the breakage concentrates at the top of the scale, because an ambiguous item and a hard item look identical to a difficulty estimator. The hardest rungs were the least trustworthy ones. That is the opposite of what a ladder is for.
So the whole bank is pulled from human measurement — not the flagged subset, all of it. It keeps exactly one job: calibrating models against each other, below.
Professionally authored, public-domain ability items (Condon & Revelle 2014, SAPA project). We refit them from scratch: joint maximum-likelihood Rasch on 23,257 responses from 1,509 people, sum-to-zero identification, extreme respondents held out as censored rather than pinned to a bound.
Origin unset — read this ladder as relative. What survives independent of who sat the items is the spacing: rotate.8 stands 29.6 W above reason.17, and that difference is sample-free (Rasch specific objectivity). Where the ladder sits on the absolute axis is not. We anchored it by declaring this sample's mean to be W 520, the published adult mean — and that assumption is unverified, because the dataset ships with no demographics at all: 1,525 web volunteers recruited through the SAPA project in August 2012, no age, no education, no country. A self-selected sample of people who opt into online cognitive testing is not an adult norm sample, and if its true mean differs, every absolute number here moves together.
The 10.79 W spread is likewise a property of this sample, not of adults. It happens to sit near the published adult SD of about 10; that is suggestive and it is not a replication, and we are not going to dress it up as one. Fixing this needs a linking design — common items sat by both this bank and an instrument with published age norms — not a better guess about who these volunteers were.
One thing does not depend on the anchor. The censoring doctrine shows up in live data on the first try: 46 of 1,509 respondents (3.0%) topped out — every item correct. Their ability is not the top of the bank; it is at least the top of the bank. A sixteen-item bank cannot see them, and any instrument that reported those 46 people as equal would be reporting its own edge as if it were theirs.
Placement status: the engine is wired to these difficulties, but the item stems are distributed under a registration agreement with the ICAR team and are not in hand. No placement goes live on inferred or reconstructed items — that is the mistake we just finished undoing. One email unblocks it.
Each source gets audited and signed off before a human sees it. An item that cannot survive scrutiny cannot be allowed to score a person — a bad item at the top of a scale doesn't merely add noise, it manufactures a false ceiling.
Maximum-likelihood Rasch fits on the benchmark bank (MMLU-Pro / GPQA / MATH-hard, 20 open-weight models, 2024-era), computed from per-item outcomes rather than reported headline scores. Valid for comparing these models to each other on this bank, and for nothing else. Not a human scale.
Do not read this table against the human band above. The two banks were gauged separately (each origin declared, neither linked), and no common-item design connects them. Comparing them would be reading two thermometers with unrelated zero points. Linking them honestly needs items that both humans and models actually sit — that design is the next real piece of work, and it is what the dyad ladder ultimately rests on.
W is the Rasch scale used by the Woodcock–Johnson battery: 10 W = a factor of 3 in the
odds of solving an item (Woodcock & Dahl 1971; W = 500 + 9.1024·logit). It is
sample-free — it measures against item difficulty, not against a population — which is why a
child, an adult and a model can occupy the same column without anyone being converted into a
percentile of anyone else.
brainelo (BE) = 1000 + 19.0849·(W − 500), gauged so 400 BE is
one decade of odds, the way Elo is. Deviation IQ is a rank laundered through an assumed normal
curve; this is a quantity. The difference matters most exactly where the tails are — which is
where every existing instrument stops resolving and reports its own edge as though it were yours.