The Meta Intelligence IndexMethod

Grounding and predictions

What informed this assessment, and what remains to be proven.

This page separates two things online cognitive tests almost always conflate: the literature that guided a design decision, and the evidence that the instrument measures what it claims. The first is set out below, decision by decision. The second is still to be produced, and this page states how we will know we were wrong.

Current status

What M.I. does not have yet.

  • No validation study has been run on this assessment. No correlation with an existing instrument has been measured.
  • The items are original. They draw on established principles but reproduce neither the questions, the norms, nor the psychometric properties of any validated battery.
  • The weights between the four domains are currently equal, by default. No data justifies them yet.
  • Norms start from a simulated distribution and move toward observed data as people take the test.
  • The sample is self-selected and the test is unproctored.
Attempts feeding the norms
0
Active norm set
auto-0
Share still from the starting assumption
100%

These three figures are read live, not written by hand. While the assumption's share stays high, a percentile is a reference point, not a measurement.

What informed each decision

Each entry states the design decision, the work that guided it, and what that work does not license. The last line matters as much as the first two.

A composite score with sub-scores, rather than a single number

Process overlap theory explains how a general factor can emerge from partially overlapping processes, without assuming a single mental substance. M.I. is therefore built as a blend — processing, memory, adaptation, lucidity — not as the measurement of one thing.

  • Kovács, K., & Conway, A. R. A. (2016). Process Overlap Theory: A Unified Account of the General Factor of Intelligence. Psychological Inquiry, 27(3), 151–177.DOI 10.1080/1047840X.2016.1153946
  • Hao, H., Conway, A. R. A., Kovács, K., & Snijder, J.-P. (2025). Simulating the process overlap theory of intelligence. Personality and Individual Differences, 233, 112865.DOI 10.1016/j.paid.2024.112865
  • WAIS-5, Pearson (2024) — official documentation, not independent research.

What this does not licenseThe separation of domains draws on WAIS-5, but M.I. reproduces neither its questions nor its norms and claims no equivalence with it.

Items that require holding, manipulating, inhibiting and switching rules

Factor models of working memory distinguish storage, executive attention and updating, with the latter two carrying most of the relationship with intelligence. Working-intelligence items are built around those three operations rather than span alone.

  • Hao, H. et al. (2025). The latent structure of working memory: A large sample factor model of working memory capacity. Cognitive, Affective, & Behavioral Neuroscience, 25, 1378–1399.DOI 10.3758/s13415-025-01310-3
  • Past reflections, present insights: A systematic review and new empirical research into the working memory capacity–fluid intelligence relationship (2025). Intelligence, 108, 101874.DOI 10.1016/j.intell.2024.101874

What this does not licenseWorking memory and fluid intelligence are strongly related without reducing to one another, so immediate-manipulation tasks stay separate from reasoning tasks.

Encoding separated from recall by intervening items

Exposure, interference, free recall then recognition comes from classical cognitive neuropsychology. It distinguishes what was encoded, retained, then retrieved, rather than treating memory as a single score.

  • Classical episodic memory paradigms (encoding, interference, recall, recognition).
  • WAIS-5 / WMS-5 documentation (Pearson, 2024) for the intelligence / memory separation.

What this does not licenseReusing a task structure transfers no clinical norm: M.I. produces no memory index interpretable in a neuropsychological sense.

Sequences where you learn from feedback, not only where you succeed

Dynamic testing measures initial performance, the effect of help or feedback, then the progression. That is the logic of the adaptation sequences: discover a rule, receive feedback after an error, try again, transfer to a variant, face an unannounced switch.

  • Boosman, H. et al. (2016). Dynamic testing of learning potential in adults with cognitive impairments. Journal of Neuropsychology, 10(2), 186–210.DOI 10.1111/jnp.12063
  • Dixon, C., Oxley, E., Gellert, A. S., & Nash, H. (2023). Dynamic assessment as a predictor of reading development. Reading and Writing, 36, 673–698.DOI 10.1007/s11145-022-10312-3

What this does not licenseDynamic assessment methods remain highly heterogeneous, and the Dixon review concerns children's reading: its application here is conceptual, not a validation.

Speed rewarded on the obvious, never penalised on the ambiguous

A response time mixes ability, caution, perceptual encoding and motor execution. That is why raw speed never accounts for more than 10% of a dimension, and why slowing down in front of an anomaly is not treated as a failure.

  • Kyllonen, P. C., & Zu, J. (2016). Use of Response Time for Measuring Cognitive Ability. Journal of Intelligence, 4(4), 14.DOI 10.3390/jintelligence4040014
  • Theisen, M. et al. (2021). Age differences in diffusion model parameters: A meta-analysis. Psychological Research, 85, 2012–2021.

What this does not licenseWithout a diffusion model fitted to the data, the test cannot yet separate slowness, caution and genuine difficulty. It therefore avoids concluding from duration alone.

Recording the correction of an answer, not just the first click

Reducing a person to their first choice loses the information in the correction. The test records the first click, changes of mind, the correction delay and declared confidence.

  • Markovitch, B., Evans, N. J., & Birk, M. V. (2024). The value of error-correcting responses for cognitive assessment in games. Scientific Reports, 14, 20657.DOI 10.1038/s41598-024-71762-z

What this does not licenseThese behaviours are recorded, but their weight in the score is still to be estimated on real data.

A "no correct answer is offered" control

Language models recognise that a correct answer is missing less reliably than they select an ordinary one. That finding motivated the control that lets a participant report an ill-posed question instead of answering anyway.

  • Groot & Colombo (2024). Large Language Models lack essential metacognition for reliable medical reasoning.

What this does not licenseThe study concerns medical reasoning in language models, not human intelligence. The mechanism stays experimental in this test.

A short, visual, interactive format

Game-based assessment can produce useful cognitive data at scale, with measurable convergent validity and test-retest reliability.

  • Leutner, F., Codreanu, S.-C., Brink, S., & Bitsakis, T. (2023). Game based assessments of cognitive ability in recruitment. Frontiers in Psychology, 13, 942662.DOI 10.3389/fpsyg.2022.942662
  • Bhargava, Y., Kottapalli, A., & Baths, V. (2024). Validation and comparison of virtual reality and 3D mobile games for cognitive assessment against ACE-III. Scientific Reports, 14, 23918.DOI 10.1038/s41598-024-75065-1

What this does not licenseGamification does not turn a test into a valid instrument. This work is a reminder that age norms and validation remain indispensable.

A score displayed as 100 ± 15, computed from standardised scores

You cannot average a response time, a count of correct answers and a memory score directly. Each measure is first converted to a z score, then combined. The display follows the convention M.I. = 100 + 15z.

  • Andrade, C. (2021). Z Scores, Standard Scores, and Composite Test Scores Explained. Indian Journal of Psychological Medicine, 43(6), 555–557.DOI 10.1177/02537176211046525

What this does not licenseThis justifies the scaling, not the weighting of domains. The weights are currently equal by default and will have to be estimated from the test's own data.

Lucidity is a hypothesis, not an established capacity

The category-error, ambiguous-pronoun and false-precision items are an original construction. They rest on classical principles — logical validity independent of premise plausibility, detecting a category error, telling absurd information from insufficient information, inhibiting the urge to answer anyway — but no study has validated this construct in this form.

What must be shown before calling it a dimension

  • That these items correlate with one another.
  • That they stay relatively stable over time.
  • That they do not simply measure education, familiarity with philosophy, or suspicion.
  • That they predict a relevant external behaviour.
  • That they form a distinct, or at least useful, factor.

Until those five points are established, "lucidity" is a working name for a research hypothesis. This page will say so for as long as that remains true.

The working hypothesis

A conventional IQ test mostly aggregates accuracy. Our bet is that observing two further signals explains variance that accuracy alone leaves out.

  1. CalibrationThe gap between declared certainty and actual accuracy. Two people with the same score can differ entirely on this, and that gap is what bears on a decision made under uncertainty.
  2. Pace regulationThe ratio between time spent on an obvious question and on an anomaly. Slowing down in the right place is the signal; answering fast everywhere is not.
  3. What follows if we are wrongIf those two signals add nothing to accuracy, then M.I. is an accuracy test with extra steps, and this page will have to say so rather than carry on.

Predictions to check

Written before the data exists, so they cannot be adjusted afterwards. Each states what would falsify it. All are measurable with data the product already collects.

P1

The four dimensions correlate positively without collapsing into one.

Measure
Pairwise correlations between dimensions across the first retained attempts.
Falsified if
Any zero or negative correlation, or a mean correlation above 0.75 — at which point the four dimensions are one.

P2

The lucidity items form a coherent set.

Measure
Internal consistency of the category-error, ambiguity and false-precision items, and their correlation with declared education level.
Falsified if
Low internal consistency, or a correlation with education higher than the consistency among the items themselves.

P3

Calibration is not a duplicate of accuracy.

Measure
Correlation between the calibration score and normalised accuracy.
Falsified if
An absolute correlation above 0.30: calibration would then measure nothing new.

P4

Slowing down for an anomaly tracks lucidity, not recall.

Measure
Correlation between the anomaly-to-obvious time ratio and each dimension.
Falsified if
Association with lucidity below 0.15, or a stronger association with recall.

P5

Overconfidence decreases as accuracy rises.

Measure
Mean gap between declared confidence and accuracy, compared between the bottom and top accuracy quartiles.
Falsified if
A gap between quartiles smaller than 10 points.

P6

Correcting a first answer carries information of its own.

Measure
Contribution of changes of mind and correction delay to predicting the score, beyond accuracy alone.
Falsified if
No measurable contribution once accuracy is accounted for.

P7

Retaking improves the score, which is why retakes are excluded.

Measure
Mean difference between a device's first attempt and its later ones.
Falsified if
A mean gain below 4 M.I. points.

P8

Norms stabilise as the sample grows.

Measure
Movement of each dimension's mean between successive norm sets, beyond a thousand attempts.
Falsified if
Movement above 0.01 in raw score from one set to the next.

P9

Language models produce a different profile from humans.

Measure
Comparison of the four dimensions between machine API runs and the human sample.
Falsified if
A model profile within half a standard deviation of the human profile on all four dimensions.

These predictions will be assessed publicly, including the ones that fail. A prediction only published when it succeeds is worth nothing.

Planned, but not yet in place

These are documented because they shape what comes next. They are not in the current test, and this page will not pretend otherwise.

  1. Adaptive item selectionChoosing later questions from earlier answers improves precision, but assumes item difficulty has been correctly estimated. Calibration on a small sample would bias the test. The current version is therefore deliberately non-adaptive: everyone receives a plan of the same length. (Sorrel, M. A., Barrada, J. R., de la Torre, J., & Abad, F. J., 2020, PLOS ONE, 15(1), e0227196.)
  2. A stability score across sittingsBrief assessments repeated at different moments are feasible on a phone and would support a stability index. Today that is only a recommendation to sit the test a second time, not a feature. (Fifield, K. et al., 2025, Assessment, 32(8), 1175–1194.)

The bar to clear

The Standards for Educational and Psychological Testing supply no questions: they set out what has to be in place before calling something a test in earnest. This is the reference this project should be judged against, and it is a reminder that a test does not become sound by accumulating users.

  • Validity
  • Reliability
  • Fairness
  • Norms
  • Documentation
  • Interpretation
  • Item security
  • Appropriate use of results

AERA, APA & NCME — Standards for Educational and Psychological Testing.

Why this page exists

An online cognitive test can simply display a number and look serious. We would rather publish what guided each decision, what that work does not license, and dated predictions you can hold us to. Every attempt brings those answers closer.