Each item has a stable source ID, subject, topic, difficulty, four distinct options, a keyed answer and an explanation.
Trust the question before trusting the score.
ElevenKey uses layered checks and is explicit about the difference between a well-formed item, an editorially approved item and an empirically calibrated item.
Five gates before confidence
Automatic checks reject missing answers, duplicate options, unusable explanations, repeated stems and answer-position imbalance.
Production packs include a separate answer ledger or rule-based solution pass before private staging.
Questions are staged privately, then approved and published through a separate action. Replacements can retire weaker versions atomically.
Attempt accuracy and response time are monitored once enough children have seen an item. Extreme results trigger review rather than silent score changes.
What automatic checks can establish
- Required fields and four valid options are present.
- The answer index is valid.
- The explanation meets minimum depth.
- Stem repetition and option-length clues are flagged.
- Difficulty, topic and answer positions are balanced.
What they cannot establish
- That the keyed answer is certainly true.
- That wording is fair to every child.
- That an item matches a real exam’s difficulty.
- That an explanation teaches the best method.
- That performance predicts a school offer.
Human review status
ElevenKey’s private production process records approval and publication separately. However, not every live item has yet been reviewed by a named external 11+ teacher or assessment specialist. Until that external programme is complete, ElevenKey describes its bank as internally verified—not independently certified.
How empirical calibration works
An item begins in collecting evidence. After at least ten recorded attempts, accuracy and average response time can flag it as balanced, unusually easy, unusually hard or a timing anomaly. This is an early quality signal, not a full psychometric validation.
Calm Readiness Score methodology
The score combines seven factors: unified accuracy (30%), target-subject coverage (18%), consistency over 14 days (14%), mistake recovery (13%), calm answer pace (10%), reading evidence (10%) and mini-mock evidence (5%).
Scores are capped when evidence is thin: fewer than five answers cannot appear above the starting band, and fewer than twenty cannot appear fully secure. Evidence labels progress from Starting estimate to Well supported.
What readiness does not mean
It is not a standardised age score, percentile, school cut-off prediction, diagnosis or guarantee. Real admission outcomes depend on exam-day conditions, cohort standardisation, school rules and information that may not be public.