How do you build a testing tool that’s fair for every learner – not just accurate on average? We asked the statistics to answer.
Deliverable T4.3 of the HITS Project lays out the methodology behind the HITS psychometric tool, built on Item Response Theory (IRT). Instead of relying on a single reliability score, the model shows precisely where a test measures well, adapts to each learner over time, and stays open to revision rather than handing out permanent labels.
Four ideas build the model:
- What IRT measures – a person’s hidden trait and each item’s difficulty and discrimination, through models like Rasch, 2PL and the Graded Response Model.
- Why it matters for HITS – items adapt in real time, cultural bias is flagged through Differential Item Functioning, and estimates update over time so no early struggle ever defines a learner for good.
- How it’s built, step by step – from preparing the data to fitting and comparing models, checking fit and fairness, and refining with anchor items.
- How theory becomes working code – a reproducible Python/R workflow the team can run, review and recalibrate.
And four more show how it’s kept honest:
- The validation checklist – one coherent construct, solid item fit, fairness across language, gender and country, and reliability read as precision across the whole scale.
- The pitfalls to avoid – samples too small, near-duplicate items, calibration left static for years, wording that assumes one culture fits all.
- A real pilot, in numbers – 420 students, 3 countries, 3 items removed for cultural bias, and 18% of learners moving upward after 6 months of recalibration.
- The principle behind it all – precision alone never guarantees fairness; the data is there to support understanding, never to classify.
Across all eight, one direction emerges: measurement should adapt to learners, not the other way around.