Full transparency

How our tests work

No proprietary black box. Here is precisely what we measure, how the score is computed, and where the limits are.

What the tests contain

Every item on IQ Lab is original - written for this site - but each item format was chosen for its research pedigree as a strong indicator of general cognitive ability:

The quick test presents 16 items across four domains (~10–15 minutes); the full test presents 40 items across five (~30–40 minutes). Within each domain, items ascend in difficulty.

The scoring model

Your score is a deviation IQ in the Wechsler convention: mean 100, standard deviation 15, computed with item response theory (IRT) - the same family of models behind modern clinical and educational tests. The computation, in order:

  1. Each item is modeled with a three-parameter logistic curve (Birnbaum, 1968; Embretson & Reise, 2000): a difficulty parameter, a discrimination parameter set by item format, and a guessing floor equal to chance on that item (25% on a 4-option item, ~17% on 6 options, 0% for typed digit-span answers). A lucky guess on a hard item therefore moves your estimate far less than a solved easy item pattern would.
  2. Item difficulties are calibrated so the model reproduces each item's a-priori pass-rate estimate - the fraction of adults expected to solve it, assigned from rule complexity and the difficulty gradients reported for comparable item types in the psychometric literature. These calibrations are provisional until validated on real response data.
  3. Your ability estimate is Bayesian (expected a posteriori): the model combines a population prior with the likelihood of your full right/wrong pattern (Bock & Mislevy, 1982). Which items you solved matters, not just how many. Then IQ = 100 + 15·ability, reported within 60–145, the range where a brief instrument retains any precision.
  4. The confidence interval is individualized: it comes from the spread of your own posterior (90% CI roughly ±13–17 points on the quick test and ±9–12 on the full test), tighter where the test is most informative and honestly wider near the extremes.
  5. Response times are recorded per item. Results showing a rapid-guessing pattern (many answers faster than reading time allows; Wise & Kong, 2005) are flagged as likely underestimates rather than silently reported.

What this is - and isn't

✔️
It is a research-based estimate: real item formats, a transparent normal-curve scoring model, honest uncertainty bands, and a domain profile instead of a single naked number.
✖️
It is not a clinical instrument. Our norms are provisional (model-derived, not census-normed), conditions are unproctored, and motivation measurably moves online scores (Duckworth et al., 2011). No online test - ours included - qualifies you for Mensa, a diagnosis, or an accommodation.

Privacy & calibration data

Everything runs in your browser. Your name, answers, and results are stored only in this browser's local storage - there are no accounts, no trackers, and no advertising analytics. Deleting a result (or clearing site data) removes it permanently.

Optional calibration contribution. Real tests earn their accuracy by measuring how items actually perform. If you leave the consent box on the test's start screen checked, then after you finish, the site sends one fully anonymous record: which items you got right or wrong, per-item response times, the test form, and the resulting score. That's the entire payload - no name, no free text, no account, and the server stores no IP addresses, cookies, or device fingerprints, so records cannot be linked to you or to each other across visits. Uncheck the box and nothing ever leaves your browser. As responses accumulate, item difficulties are re-estimated from real data and the scoring model's provisional norms are replaced with measured ones - this page will state when that has happened.

Key sources for this page

  1. Raven, J. (2000). The Raven's Progressive Matrices: Change and stability over culture and time. Cognitive Psychology, 41(1), 1–48.
  2. Carroll, J. B. (1993). Human Cognitive Abilities. Cambridge University Press.
  3. McGrew, K. S. (2009). CHC theory and the human cognitive abilities project. Intelligence, 37(1), 1–10.
  4. Shepard, R. N., & Metzler, J. (1971). Mental rotation of three-dimensional objects. Science, 171(3972), 701–703.
  5. Deary, I. J. (2012). Intelligence. Annual Review of Psychology, 63, 453–482.
  6. Birnbaum, A. (1968). Some latent trait models and their use in inferring an examinee's ability. In Lord & Novick, Statistical Theories of Mental Test Scores. Addison-Wesley.
  7. Embretson, S. E., & Reise, S. P. (2000). Item Response Theory for Psychologists. Erlbaum.
  8. Bock, R. D., & Mislevy, R. J. (1982). Adaptive EAP estimation of ability in a microcomputer environment. Applied Psychological Measurement, 6(4), 431–444.
  9. Wise, S. L., & Kong, X. (2005). Response time effort: A new measure of examinee motivation in computer-based tests. Applied Measurement in Education, 18(2), 163–183.
  10. Duckworth, A. L., et al. (2011). Role of test motivation in intelligence testing. PNAS, 108(19), 7716–7720.
  11. Wechsler, D. (2008). WAIS-IV Technical and Interpretive Manual. Pearson.
🧠
Read the broader evidence on the Science page, or take the test now.