Issue 01 Evidence and methods

The evidence behind Lux, in plain terms

A clear account of Lux's scoring architecture, the evidence available today and the questions that remain open.

In brief

What to know before you read.

  • Lux uses a versioned 19-layer scoring engine spanning 301 fields.
  • Regression tests and fixtures help detect unintended computational changes.
  • Source-bank reliability figures do not establish Lux-specific reliability, validity or fitness for use.

Most people never see what happens between answering an assessment’s questions and receiving a report. That gap matters. It contains the scoring decisions, technical controls and evidence limits that determine what the result can reasonably mean.

What “19 layers” means in practice

Lux’s scoring is built as a 19-layer engine spanning 301 fields. The layers turn raw responses into progressively more meaningful structure, from item-level scoring to the narrative language in a participant report. Separating that work into layers makes each step easier to inspect. Scoring changes can be checked through automated regression tests and selected reference fixtures.

That matters because psychological scoring is easy to get subtly wrong in ways that are hard to notice later. A transposed weight or a rounding error can pass through several steps without announcing itself.

Testing layers against committed fixtures, and versioning the whole pathway, helps us detect unintended computational changes before release.

What the reliability figures actually say

The complete ten-item Big Five item banks used as source/reference material have documented internal consistency of α .800 to .898 in a 603,322-response source/reference dataset. Internal consistency describes how closely responses to items within each bank relate in that dataset.

That is source-bank internal-consistency evidence. It does not establish that a bank measures one underlying construct, that Lux’s proprietary fields are valid, that Lux itself is reliable, or that Lux is fit for a particular use. The 603,322 responses are not Lux-participant responses. We state what the figures support and stop there.

Why we lead with this instead of hiding it

Publishing the mechanics, available evidence and limits gives readers a clearer basis for judging the work. Versioning and testing make the scoring system more inspectable. They do not replace psychological validation.

Lux is moving through an active, staged evidence-development programme. Test–retest work will assess score stability as repeat-response data becomes available. Later stages will examine convergent and discriminant validity for proprietary Lux fields, representative norms, intended-population fairness and clinical utility using suitable data and pre-defined study designs. This work is active, but it has not yet produced completed Lux-specific validation results. The full evidence position is the source of record for those boundaries and will change as the programme produces findings that can support new claims.

Source notes

Follow the evidence trail.

  1. Lux evidence and claim limits

    The complete public account of current evidence and active research.

  2. How Lux works

    The assessment purpose, reporting model and intended role.