The Meta Intelligence IndexMethod

Evidence base and predictions

What informed this assessment, and what remains to be proven.

This page separates two things that online cognitive tests often conflate: the research that informed each design decision and the evidence that this assessment measures what it claims to measure. The first is set out below, decision by decision. The second still needs to be established, and this page explains what findings would show that our hypotheses are wrong.

Current status

What M.I. does not have yet.

  • No validation study has been run on this assessment. No correlation with an existing instrument has been measured.
  • The items are original. They draw on established principles but reproduce neither the questions, the norms, nor the psychometric properties of any validated battery.
  • The five domains currently have equal weight by default. No evidence supports those weights yet.
  • Norms start from a simulated distribution and move toward observed data as people take the test.
  • The sample is self-selected and the test is unproctored.
Attempts included in the norms
0
Active norm set
development-3
Share of the reference still based on simulation
100%

These three figures are read live, not entered manually. The percentile places a score within the active simulated reference used during development; it does not indicate a person’s rank in the human population.

Form 7 addendum — The preregistered protocol below was written for the original four domains. Empathy was added later as a fifth family, and bank 3.0 now balances all five families. These changes have not been retrofitted into the original hypotheses: form 7 analyses will be published separately and versioned.

What informed each decision

Each entry states the design decision, the research that informed it, and the conclusions that research does not support. The final line matters as much as the first two.

A composite score with sub-scores, rather than a single number

Process overlap theory explains how a general factor can emerge from partially overlapping processes without assuming a single mental substance. M.I. therefore combines working intelligence, recall, adaptation, lucidity and empathy rather than claiming to measure a single trait.

  • Kovács, K., & Conway, A. R. A. (2016). Process Overlap Theory: A Unified Account of the General Factor of Intelligence. Psychological Inquiry, 27(3), 151–177.DOI 10.1080/1047840X.2016.1153946
  • Hao, H., Conway, A. R. A., Kovács, K., & Snijder, J.-P. (2025). Simulating the process overlap theory of intelligence. Personality and Individual Differences, 233, 112865.DOI 10.1016/j.paid.2024.112865
  • WAIS-5, Pearson (2024) — official documentation, not independent research.

What this does not supportThe separation of domains draws on WAIS-5, but M.I. reproduces neither its questions nor its norms and claims no equivalence with it.

Items that require holding, manipulating, inhibiting and switching rules

Factor models of working memory distinguish storage, executive attention and updating, with the latter two carrying most of the relationship with intelligence. Working-intelligence items are built around those three operations rather than span alone.

  • Hao, H. et al. (2025). The latent structure of working memory: A large sample factor model of working memory capacity. Cognitive, Affective, & Behavioral Neuroscience, 25, 1378–1399.DOI 10.3758/s13415-025-01310-3
  • Past reflections, present insights: A systematic review and new empirical research into the working memory capacity–fluid intelligence relationship (2025). Intelligence, 108, 101874.DOI 10.1016/j.intell.2024.101874

What this does not supportWorking memory and fluid intelligence are strongly related but not interchangeable, so immediate-manipulation tasks remain separate from reasoning tasks.

Encoding separated from recall by intervening items

The sequence of exposure, an intervening task, free recall and then recognition comes from classical cognitive neuropsychology. It distinguishes what was encoded, retained and retrieved rather than treating memory as a single score.

  • Classical episodic memory paradigms (encoding, interference, recall, recognition).
  • WAIS-5 / WMS-5 documentation (Pearson, 2024) for the intelligence / memory separation.

What this does not supportUsing a similar task structure does not transfer clinical norms: M.I. does not produce a memory index that can be interpreted neuropsychologically.

Sequences where you learn from feedback, not only where you succeed

Dynamic testing measures initial performance, the effect of help or feedback, and subsequent improvement. That is the logic of the adaptation sequences: discover a rule, receive feedback after an error, try again, transfer to a variant, face an unannounced switch.

  • Boosman, H. et al. (2016). Dynamic testing of learning potential in adults with cognitive impairments. Journal of Neuropsychology, 10(2), 186–210.DOI 10.1111/jnp.12063
  • Dixon, C., Oxley, E., Gellert, A. S., & Nash, H. (2023). Dynamic assessment as a predictor of reading development. Reading and Writing, 36, 673–698.DOI 10.1007/s11145-022-10312-3

What this does not supportDynamic assessment methods remain highly heterogeneous, and the Dixon review concerns children's reading: its application here is conceptual, not a validation.

Pace relative to each item’s reference time

Response time mixes ability, caution, perceptual encoding and motor execution. On items where timing is interpretable, credit is therefore multiplied by a factor bounded from 0.75 to 1.25 and centred on that item’s reference time. Tasks with fixed timing are not pace-adjusted.

  • Kyllonen, P. C., & Zu, J. (2016). Use of Response Time for Measuring Cognitive Ability. Journal of Intelligence, 4(4), 14.DOI 10.3390/jintelligence4040014
  • Theisen, M. et al. (2021). Age differences in diffusion model parameters: A meta-analysis. Psychological Research, 85, 2012–2021.

What this does not supportWithout a diffusion model fitted to the data, the test cannot yet separate slowness, caution and genuine difficulty. The ±25% factor is an experimental design choice that still needs validation, not an established psychometric property.

Recording the correction of an answer, not just the first click

Looking only at a person’s first choice discards information contained in later corrections. The assessment records the first selection, changes of mind, the time taken to correct an answer and declared confidence. Confidence is scored with a quadratic rule designed so that honestly reporting one’s belief maximises the expected score.

  • Markovitch, B., Evans, N. J., & Birk, M. V. (2024). The value of error-correcting responses for cognitive assessment in games. Scientific Reports, 14, 20657.DOI 10.1038/s41598-024-71762-z

What this does not supportThe confidence weight and interpretation of answer changes have not been calibrated on human data. The rule’s mathematical property validates neither its weight nor the construct it measures.

A "no correct answer is offered" control

Language models recognise that a correct answer is missing less reliably than they select an ordinary one. That finding motivated the control that lets a participant report an ill-posed question instead of answering anyway.

  • Groot & Colombo (2024). Large Language Models lack essential metacognition for reliable medical reasoning.

What this does not supportThe study concerns medical reasoning in language models, not human intelligence. The mechanism remains experimental in this assessment.

A short, visual, interactive format

Game-based assessment can produce useful cognitive data at scale, with measurable convergent validity and test-retest reliability.

  • Leutner, F., Codreanu, S.-C., Brink, S., & Bitsakis, T. (2023). Game based assessments of cognitive ability in recruitment. Frontiers in Psychology, 13, 942662.DOI 10.3389/fpsyg.2022.942662
  • Bhargava, Y., Kottapalli, A., & Baths, V. (2024). Validation and comparison of virtual reality and 3D mobile games for cognitive assessment against ACE-III. Scientific Reports, 14, 23918.DOI 10.1038/s41598-024-75065-1

What this does not supportGamification does not turn a test into a valid instrument. These studies also show that age-based norms and validation remain essential.

A score displayed as 100 ± 15, computed from standardised scores

You cannot average a response time, a count of correct answers and a memory score directly. Each measure is first converted to a z score, then combined. The display follows the convention M.I. = 100 + 15z.

  • Andrade, C. (2021). Z Scores, Standard Scores, and Composite Test Scores Explained. Indian Journal of Psychological Medicine, 43(6), 555–557.DOI 10.1177/02537176211046525

What this does not supportThis justifies the scaling, not the weighting of domains. The weights are currently equal by default and will have to be estimated from the test's own data.

Lucidity is a hypothesis, not an established capacity

The category-error, ambiguous-pronoun and false-precision items are an original construction. They rest on classical principles — logical validity independent of premise plausibility, detecting a category error, telling absurd information from insufficient information, inhibiting the urge to answer anyway — but no study has validated this construct in this form.

What must be shown before calling it a dimension

  • That these items correlate with one another.
  • That they stay relatively stable over time.
  • That they do not simply measure education, familiarity with philosophy or scepticism.
  • That they predict a relevant external behaviour.
  • That they form a distinct, or at least useful, factor.

Until those five points are established, "lucidity" is a working name for a research hypothesis. This page will say so for as long as that remains true.

The working hypothesis

A conventional IQ score is based primarily on correct responses. Our hypothesis is that two additional signals may explain differences that accuracy alone does not capture.

  1. CalibrationThe gap between declared certainty and actual accuracy. Two people with the same score may differ greatly on this measure, and that difference matters when decisions are made under uncertainty.
  2. Pace regulationThe ratio between time spent on an obvious question and on an anomaly. Slowing down in the right place is the signal; answering fast everywhere is not.
  3. What follows if we are wrongIf those two signals add nothing beyond accuracy, then M.I. is simply an accuracy test with extra steps. We will state that here rather than continue to claim otherwise.

Predictions to test

Written before the data were collected so they could not be adjusted afterwards. Each states what would falsify it, and all can be tested using data the assessment already collects.

P1

The four dimensions correlate positively without collapsing into one.

Measure
Pairwise correlations between dimensions across the first retained attempts.
Falsified if
Any zero or negative correlation, or a mean correlation above 0.75 — at which point the four dimensions are one.

P2

The lucidity items form a coherent set.

Measure
Internal consistency of the category-error, ambiguity and false-precision items, and their correlation with declared education level.
Falsified if
Low internal consistency, or a correlation with education higher than the consistency among the items themselves.

P3

Calibration is distinct from accuracy.

Measure
Correlation between the calibration score and normalised accuracy.
Falsified if
An absolute correlation above 0.30: calibration would then measure nothing new.

P4

Slowing down for an anomaly tracks lucidity, not recall.

Measure
Correlation between the anomaly-to-obvious time ratio and each dimension.
Falsified if
Association with lucidity below 0.15, or a stronger association with recall.

P5

Overconfidence decreases as accuracy rises.

Measure
Mean gap between declared confidence and accuracy, compared between the bottom and top accuracy quartiles.
Falsified if
A gap between quartiles smaller than 10 points.

P6

Changing an initial answer provides additional information.

Measure
Contribution of changes of mind and correction delay to predicting the score, beyond accuracy alone.
Falsified if
No measurable contribution once accuracy is accounted for.

P7

Retaking improves the score, which is why retakes are excluded.

Measure
Mean difference between a device's first attempt and its later ones.
Falsified if
A mean gain below 4 M.I. points.

P8

Norms stabilise as the sample grows.

Measure
Movement of each dimension's mean between successive norm sets, beyond a thousand attempts.
Falsified if
Movement above 0.01 in raw score from one set to the next.

P9

Language models produce a different profile from humans.

Measure
Comparison of the four dimensions between machine API runs and the human sample.
Falsified if
A model profile within half a standard deviation of the human profile on all four dimensions.

These predictions will be assessed publicly, including the ones that fail. A prediction only published when it succeeds is worth nothing.

Planned, but not yet in place

These are documented because they shape what comes next. They are not in the current test, and this page will not pretend otherwise.

  1. Adaptive item selectionChoosing later questions from earlier answers improves precision, but assumes item difficulty has been correctly estimated. Calibration on a small sample would bias the test. The current version is therefore deliberately non-adaptive: everyone receives a plan of the same length. (Sorrel, M. A., Barrada, J. R., de la Torre, J., & Abad, F. J., 2020, PLOS ONE, 15(1), e0227196.)
  2. A stability score across sittingsBrief assessments repeated at different moments are feasible on a phone and would support a stability index. Today that is only a recommendation to sit the test a second time, not a feature. (Fifield, K. et al., 2025, Assessment, 32(8), 1175–1194.)

The bar to clear

The Standards for Educational and Psychological Testing do not provide test questions; they describe what must be in place before an instrument can be regarded as a sound test. This project should be assessed against those standards. An assessment does not become valid simply by accumulating users.

  • Validity
  • Reliability
  • Fairness
  • Norms
  • Documentation
  • Interpretation
  • Item security
  • Appropriate use of results

AERA, APA & NCME — Standards for Educational and Psychological Testing.

Why this page exists

An online cognitive test can simply display a number and look serious. We would rather publish what guided each decision, the conclusions that work does not support, and dated predictions you can hold us to. Each completed assessment brings us closer to answering those questions.