brainelo

an absolute ladder · humans, machines, and pairs on one scale
placement suspended

Item bank under audit.

The machinery works. The questions didn't deserve you.

Placement ran on items derived from public LLM benchmarks. Audit that bank and its provenance falls apart: scraped test-prep, pipelined into a benchmark, deployed because it was convenient. Nobody along that chain ever asked whether an item was fit to measure a person. Roughly one in twelve to one in twenty knowledge items is outright broken — and the breakage concentrates at the top of the scale, because an ambiguous item and a hard item look identical to a difficulty estimator. The hardest rungs were the least trustworthy ones. That is the opposite of what a ladder is for.

So the whole bank is pulled from human measurement — not the flagged subset, all of it. It keeps exactly one job: calibrating models against each other, below.

first worthy bank — ICAR-16, calibrated on 1,525 human respondents

Professionally authored, public-domain ability items (Condon & Revelle 2014, SAPA project). We refit them from scratch: joint maximum-likelihood Rasch on 23,257 responses from 1,509 people, sum-to-zero identification, extreme respondents held out as censored rather than pinned to a bound.

ladder span (relative)
29.6 W from easiest to hardest item
sample spread
10.79 W

Origin unset — read this ladder as relative. What survives independent of who sat the items is the spacing: rotate.8 stands 29.6 W above reason.17, and that difference is sample-free (Rasch specific objectivity). Where the ladder sits on the absolute axis is not. We anchored it by declaring this sample's mean to be W 520, the published adult mean — and that assumption is unverified, because the dataset ships with no demographics at all: 1,525 web volunteers recruited through the SAPA project in August 2012, no age, no education, no country. A self-selected sample of people who opt into online cognitive testing is not an adult norm sample, and if its true mean differs, every absolute number here moves together.

The 10.79 W spread is likewise a property of this sample, not of adults. It happens to sit near the published adult SD of about 10; that is suggestive and it is not a replication, and we are not going to dress it up as one. Fixing this needs a linking design — common items sat by both this bank and an instrument with published age norms — not a better guess about who these volunteers were.

One thing does not depend on the anchor. The censoring doctrine shows up in live data on the first try: 46 of 1,509 respondents (3.0%) topped out — every item correct. Their ability is not the top of the bank; it is at least the top of the bank. A sixteen-item bank cannot see them, and any instrument that reported those 46 people as equal would be reporting its own edge as if it were theirs.

Placement status: the engine is wired to these difficulties, but the item stems are distributed under a registration agreement with the ICAR team and are not in hand. No placement goes live on inferred or reconstructed items — that is the mistake we just finished undoing. One email unblocks it.

further sources, none deployed yet

Each source gets audited and signed off before a human sees it. An item that cannot survive scrutiny cannot be allowed to score a person — a bad item at the top of a scale doesn't merely add noise, it manufactures a false ceiling.

model-vs-model calibration — the one job the benchmark bank keeps

Maximum-likelihood Rasch fits on the benchmark bank (MMLU-Pro / GPQA / MATH-hard, 20 open-weight models, 2024-era), computed from per-item outcomes rather than reported headline scores. Valid for comparing these models to each other on this bank, and for nothing else. Not a human scale.

Do not read this table against the human band above. The two banks were gauged separately (each origin declared, neither linked), and no common-item design connects them. Comparing them would be reading two thermometers with unrelated zero points. Linking them honestly needs items that both humans and models actually sit — that design is the next real piece of work, and it is what the dyad ladder ultimately rests on.

the scale

W is the Rasch scale used by the Woodcock–Johnson battery: 10 W = a factor of 3 in the odds of solving an item (Woodcock & Dahl 1971; W = 500 + 9.1024·logit). It is sample-free — it measures against item difficulty, not against a population — which is why a child, an adult and a model can occupy the same column without anyone being converted into a percentile of anyone else.

brainelo (BE) = 1000 + 19.0849·(W − 500), gauged so 400 BE is one decade of odds, the way Elo is. Deviation IQ is a rank laundered through an assumed normal curve; this is a quantity. The difference matters most exactly where the tails are — which is where every existing instrument stops resolving and reports its own edge as though it were yours.

Standing commitments. Ranks are floors: a ceiling score means the bank ran out, not that you did. Items are audited before they may rate a person, and retired when they leak. Age is a covariate of a growth model, never part of the measure.
What this is not. Not an IQ. It does not predict school, work or life. Practising the tasks trains the tasks — near transfer is real, far transfer is not, and we will report what we measure rather than promise what we would like.