What the tests contain
Every item on IQ Lab is original - written for this site - but each item format was chosen for its research pedigree as a strong indicator of general cognitive ability:
- Figural matrices (find the tile completing a 3×3 pattern) - the format of the Raven's Progressive Matrices, one of the most g-loaded and culture-portable formats known (Raven, 2000; Carroll, 1993). Our matrices use rules established in the literature: attribute progression, distribution-of-three (Latin squares), figure combination, and rotation.
- Verbal reasoning - analogies, classification, vocabulary, and syllogisms, the core of verbal-comprehension scales in the Wechsler tradition. Syllogisms use nonsense terms ("Blorks," "Fleems") so logic, not world knowledge, decides the answer.
- Quantitative reasoning - number-series and quantitative logic items, indicators of fluid/quantitative reasoning in the CHC model (McGrew, 2009).
- Spatial rotation - deciding whether a shape is a rotation or a mirror image, after Shepard & Metzler's (1971) classic mental-rotation paradigm. All shapes are chiral, so every mirrored distractor is genuinely wrong - no rotation can make it match.
- Working memory (full test only) - forward digit span, administered digit-by-digit, exactly one attempt per sequence, in the Wechsler tradition. Span tasks index the memory component of fluid cognition (Deary, 2012).
The quick test presents 16 items across four domains (~10–15 minutes); the full test presents 40 items across five (~30–40 minutes). Within each domain, items ascend in difficulty.
The scoring model
Your score is a deviation IQ in the Wechsler convention: mean 100, standard deviation 15, computed with item response theory (IRT) - the same family of models behind modern clinical and educational tests. The computation, in order:
- Each item is modeled with a three-parameter logistic curve (Birnbaum, 1968; Embretson & Reise, 2000): a difficulty parameter, a discrimination parameter set by item format, and a guessing floor equal to chance on that item (25% on a 4-option item, ~17% on 6 options, 0% for typed digit-span answers). A lucky guess on a hard item therefore moves your estimate far less than a solved easy item pattern would.
- Item difficulties are calibrated so the model reproduces each item's a-priori pass-rate estimate - the fraction of adults expected to solve it, assigned from rule complexity and the difficulty gradients reported for comparable item types in the psychometric literature. These calibrations are provisional until validated on real response data.
- Your ability estimate is Bayesian (expected a posteriori): the model combines a
population prior with the likelihood of your full right/wrong pattern (Bock & Mislevy, 1982). Which
items you solved matters, not just how many. Then
IQ = 100 + 15·ability, reported within 60–145, the range where a brief instrument retains any precision. - The confidence interval is individualized: it comes from the spread of your own posterior (90% CI roughly ±13–17 points on the quick test and ±9–12 on the full test), tighter where the test is most informative and honestly wider near the extremes.
- Response times are recorded per item. Results showing a rapid-guessing pattern (many answers faster than reading time allows; Wise & Kong, 2005) are flagged as likely underestimates rather than silently reported.
What this is - and isn't
- Practice effects are real. Retaking inflates scores by several points; your first attempt is the most informative. If you retest, wait a few weeks.
- Conditions matter. Sleep deprivation, alcohol, interruptions, and low effort all depress scores. Test rested, in quiet, in one sitting.
- One sitting is one sample. Even clinical IQ has a ±5-point band around any single administration. Interpret gaps between your domain scores loosely - short scales are noisy.
Privacy & calibration data
Everything runs in your browser. Your name, answers, and results are stored only in this browser's local storage - there are no accounts, no trackers, and no advertising analytics. Deleting a result (or clearing site data) removes it permanently.
Optional calibration contribution. Real tests earn their accuracy by measuring how items actually perform. If you leave the consent box on the test's start screen checked, then after you finish, the site sends one fully anonymous record: which items you got right or wrong, per-item response times, the test form, and the resulting score. That's the entire payload - no name, no free text, no account, and the server stores no IP addresses, cookies, or device fingerprints, so records cannot be linked to you or to each other across visits. Uncheck the box and nothing ever leaves your browser. As responses accumulate, item difficulties are re-estimated from real data and the scoring model's provisional norms are replaced with measured ones - this page will state when that has happened.
Key sources for this page
- Raven, J. (2000). The Raven's Progressive Matrices: Change and stability over culture and time. Cognitive Psychology, 41(1), 1–48.
- Carroll, J. B. (1993). Human Cognitive Abilities. Cambridge University Press.
- McGrew, K. S. (2009). CHC theory and the human cognitive abilities project. Intelligence, 37(1), 1–10.
- Shepard, R. N., & Metzler, J. (1971). Mental rotation of three-dimensional objects. Science, 171(3972), 701–703.
- Deary, I. J. (2012). Intelligence. Annual Review of Psychology, 63, 453–482.
- Birnbaum, A. (1968). Some latent trait models and their use in inferring an examinee's ability. In Lord & Novick, Statistical Theories of Mental Test Scores. Addison-Wesley.
- Embretson, S. E., & Reise, S. P. (2000). Item Response Theory for Psychologists. Erlbaum.
- Bock, R. D., & Mislevy, R. J. (1982). Adaptive EAP estimation of ability in a microcomputer environment. Applied Psychological Measurement, 6(4), 431–444.
- Wise, S. L., & Kong, X. (2005). Response time effort: A new measure of examinee motivation in computer-based tests. Applied Measurement in Education, 18(2), 163–183.
- Duckworth, A. L., et al. (2011). Role of test motivation in intelligence testing. PNAS, 108(19), 7716–7720.
- Wechsler, D. (2008). WAIS-IV Technical and Interpretive Manual. Pearson.