Blog

Why we wait for n≥500 before we say a cohort is calibrated

Lux is a classical test theory, fixed-form assessment in its current phase. That is a deliberate choice, not a placeholder. It means every person who takes Lux today answers the same well-tested item set, scored against evidence that already exists, rather than an adaptive form we couldn’t yet back with enough data. The more adaptive, item-response-theory-based version of Lux (the kind that can tailor which questions a person sees next) is real Phase 2 work, and it is explicitly gated on reaching at least 500 eligible responses in a given calibration population before we’ll treat that population’s data as ready to calibrate against.

What the gate actually protects against

Five hundred is not an arbitrary round number. Calibration work, building the statistical models that let an adaptive assessment select items intelligently, is only as trustworthy as the sample it’s built on. Try to calibrate against 10, 20 or 30 responses, and you get a model that looks precise but is actually fitting noise. It will produce numbers. Those numbers will not mean what they appear to mean.

We would rather tell a community “we don’t have enough data yet” than publish a calibrated model built on a sample too small to support one. That is true even when a partner organisation is eager to see adaptive results sooner, and even when a smaller sample would technically produce output.

Producing output isn’t the bar. Producing output that means what we say it means is the bar.

What happens before the gate is reached

None of this slows down what Lux already offers. The current fixed-form engine (the 19 layers, the versioned scoring, the reliability evidence already established across 603,322 responses in the source sample) is real, tested, and available today. What’s gated on n≥500 per cohort is specifically the next layer: population-specific calibration and, eventually, adaptive item selection for that population.

Community cohorts accumulate real, consenting responses over time. As a cohort clears the n≥500 threshold, that population’s data becomes eligible for calibration work, and every study built on it says plainly what it found, including when the honest answer is “not yet enough to conclude that.” That is the same standard we hold the rest of our public claims to, applied to the research programme itself: real evidence, checked before it’s trusted, and never rounded up to sound more finished than it is.

See it for yourself

Read the standard Lux is actually built to

Everything above is the discipline behind Lux, not a claim apart from it. See the full picture on the assessment itself, or get in touch if you're considering a pilot, funding a study, or just have a question.

Explore LuxGet in touch