Each entry states the design decision, the research that informed it, and the conclusions that research does not support. The final line matters as much as the first two.
A composite score with sub-scores, rather than a single number
Process overlap theory explains how a general factor can emerge from partially overlapping processes without assuming a single mental substance. M.I. therefore combines working intelligence, recall, adaptation, lucidity and empathy rather than claiming to measure a single trait.
- Kovács, K., & Conway, A. R. A. (2016). Process Overlap Theory: A Unified Account of the General Factor of Intelligence. Psychological Inquiry, 27(3), 151–177.DOI 10.1080/1047840X.2016.1153946
- Hao, H., Conway, A. R. A., Kovács, K., & Snijder, J.-P. (2025). Simulating the process overlap theory of intelligence. Personality and Individual Differences, 233, 112865.DOI 10.1016/j.paid.2024.112865
- WAIS-5, Pearson (2024) — official documentation, not independent research.
What this does not supportThe separation of domains draws on WAIS-5, but M.I. reproduces neither its questions nor its norms and claims no equivalence with it.
Items that require holding, manipulating, inhibiting and switching rules
Factor models of working memory distinguish storage, executive attention and updating, with the latter two carrying most of the relationship with intelligence. Working-intelligence items are built around those three operations rather than span alone.
- Hao, H. et al. (2025). The latent structure of working memory: A large sample factor model of working memory capacity. Cognitive, Affective, & Behavioral Neuroscience, 25, 1378–1399.DOI 10.3758/s13415-025-01310-3
- Past reflections, present insights: A systematic review and new empirical research into the working memory capacity–fluid intelligence relationship (2025). Intelligence, 108, 101874.DOI 10.1016/j.intell.2024.101874
What this does not supportWorking memory and fluid intelligence are strongly related but not interchangeable, so immediate-manipulation tasks remain separate from reasoning tasks.
Encoding separated from recall by intervening items
The sequence of exposure, an intervening task, free recall and then recognition comes from classical cognitive neuropsychology. It distinguishes what was encoded, retained and retrieved rather than treating memory as a single score.
- Classical episodic memory paradigms (encoding, interference, recall, recognition).
- WAIS-5 / WMS-5 documentation (Pearson, 2024) for the intelligence / memory separation.
What this does not supportUsing a similar task structure does not transfer clinical norms: M.I. does not produce a memory index that can be interpreted neuropsychologically.
Sequences where you learn from feedback, not only where you succeed
Dynamic testing measures initial performance, the effect of help or feedback, and subsequent improvement. That is the logic of the adaptation sequences: discover a rule, receive feedback after an error, try again, transfer to a variant, face an unannounced switch.
- Boosman, H. et al. (2016). Dynamic testing of learning potential in adults with cognitive impairments. Journal of Neuropsychology, 10(2), 186–210.DOI 10.1111/jnp.12063
- Dixon, C., Oxley, E., Gellert, A. S., & Nash, H. (2023). Dynamic assessment as a predictor of reading development. Reading and Writing, 36, 673–698.DOI 10.1007/s11145-022-10312-3
What this does not supportDynamic assessment methods remain highly heterogeneous, and the Dixon review concerns children's reading: its application here is conceptual, not a validation.
Pace relative to each item’s reference time
Response time mixes ability, caution, perceptual encoding and motor execution. On items where timing is interpretable, credit is therefore multiplied by a factor bounded from 0.75 to 1.25 and centred on that item’s reference time. Tasks with fixed timing are not pace-adjusted.
- Kyllonen, P. C., & Zu, J. (2016). Use of Response Time for Measuring Cognitive Ability. Journal of Intelligence, 4(4), 14.DOI 10.3390/jintelligence4040014
- Theisen, M. et al. (2021). Age differences in diffusion model parameters: A meta-analysis. Psychological Research, 85, 2012–2021.
What this does not supportWithout a diffusion model fitted to the data, the test cannot yet separate slowness, caution and genuine difficulty. The ±25% factor is an experimental design choice that still needs validation, not an established psychometric property.
Recording the correction of an answer, not just the first click
Looking only at a person’s first choice discards information contained in later corrections. The assessment records the first selection, changes of mind, the time taken to correct an answer and declared confidence. Confidence is scored with a quadratic rule designed so that honestly reporting one’s belief maximises the expected score.
- Markovitch, B., Evans, N. J., & Birk, M. V. (2024). The value of error-correcting responses for cognitive assessment in games. Scientific Reports, 14, 20657.DOI 10.1038/s41598-024-71762-z
What this does not supportThe confidence weight and interpretation of answer changes have not been calibrated on human data. The rule’s mathematical property validates neither its weight nor the construct it measures.
A "no correct answer is offered" control
Language models recognise that a correct answer is missing less reliably than they select an ordinary one. That finding motivated the control that lets a participant report an ill-posed question instead of answering anyway.
- Groot & Colombo (2024). Large Language Models lack essential metacognition for reliable medical reasoning.
What this does not supportThe study concerns medical reasoning in language models, not human intelligence. The mechanism remains experimental in this assessment.
A short, visual, interactive format
Game-based assessment can produce useful cognitive data at scale, with measurable convergent validity and test-retest reliability.
- Leutner, F., Codreanu, S.-C., Brink, S., & Bitsakis, T. (2023). Game based assessments of cognitive ability in recruitment. Frontiers in Psychology, 13, 942662.DOI 10.3389/fpsyg.2022.942662
- Bhargava, Y., Kottapalli, A., & Baths, V. (2024). Validation and comparison of virtual reality and 3D mobile games for cognitive assessment against ACE-III. Scientific Reports, 14, 23918.DOI 10.1038/s41598-024-75065-1
What this does not supportGamification does not turn a test into a valid instrument. These studies also show that age-based norms and validation remain essential.
A score displayed as 100 ± 15, computed from standardised scores
You cannot average a response time, a count of correct answers and a memory score directly. Each measure is first converted to a z score, then combined. The display follows the convention M.I. = 100 + 15z.
- Andrade, C. (2021). Z Scores, Standard Scores, and Composite Test Scores Explained. Indian Journal of Psychological Medicine, 43(6), 555–557.DOI 10.1177/02537176211046525
What this does not supportThis justifies the scaling, not the weighting of domains. The weights are currently equal by default and will have to be estimated from the test's own data.