Scientific basis
A century of research heritage
The scientific measurement of intelligence began in 1904, when Spearman used statistics to pin down a general factor he called g. Cattell divided intelligence into fluid and crystallized abilities in the 1940s, and in 1993 Carroll re-analyzed hundreds of datasets to build a three-stratum hierarchy. The model that pulls this line of work together is CHC. It is the standard taxonomy that today’s major cognitive assessments share, and the closest thing psychometrics has to a consensus reached over a century.
The task formats in MOSAIC-48 rest on the same tested lineage. Tasks that ask you to hold digits and sequences in mind have been standard tools of memory research since the late 19th century. Matrix reasoning became the reference format for measuring reasoning with minimal cultural loading after the 1930s, and the color–word interference task of 1935 remains one of the most replicated paradigms in attention research. Our version, though, does not measure interference through response time the way the original does; it scores accuracy — whether you picked the dimension the instruction asked for, the ink color or the word’s meaning. The lineage is shared; the measurement is not. MOSAIC-48 was designed on top of theory and task formats like these — tested over decades, and in some cases well past the hundred-year mark.
None of that theory, and none of those task formats, is a closed chapter. Through the 2020s researchers have kept re-examining the CHC broad-ability structure in large samples, alongside validity work on open item banks, checks on the data quality of unproctored online testing, and a rethinking of how attention control should be measured. That body of work is all still being tested.
Measurement model
MOSAIC-48 takes up that lineage without reducing cognitive ability to a single score, measuring eight domains at the broad-ability level the hierarchical CHC model proposes. Fluid reasoning, visual-spatial processing, quantitative reasoning, working memory and processing speed line up reasonably well with CHC broad abilities. Verbal reasoning has to be read more carefully. Its items ask nothing about vocabulary or general knowledge — crystallized ability, Gc — and instead measure propositional logic and argument analysis, that is, reasoning carried by language, so it cannot simply be equated with Gc. There are no vocabulary or general-knowledge items in the bank at all, which means the eight domains do not cover the whole of the CHC broad abilities. Learning efficiency and attention control are a slightly different case. Recent research has reported the centrality of the learning and control complexes, and we added the two domains to reflect that.
In large-scale online studies, the weight of the evidence falls on models that break individual differences into several components — reasoning, short-term memory, verbal ability — rather than on a general factor alone. So the results do not stop at a single composite cognitive index; they come with an eight-axis profile.
Item design
The items for all eight domains were written from scratch. We drew on the item-type taxonomy of the open ICAR project, the probability and frequency formats of the Berlin Numeracy Test and the intuition-then-verification structure of the Cognitive Reflection Test, but nothing was lifted from them — each item is generated from attribute-combination rules.
Figural items are defined by combinations of shape, count, fill, rotation and marking, together with the rules that transform them. Distractors are not drawn at random. Each one comes from a predictable reasoning error: applying only one of the rules, reading the rotation the wrong way round, or missing a change in count. Every distractor carries an error code and is filed under a shared classification that cuts across items, so when the same error recurs on several items it surfaces as a response tendency in the result.
Reasoning, verbal and quantitative items are interleaved, so that no item announces which ability it is measuring. Working memory (shown, then removed), processing speed (under a time limit) and learning efficiency (learn now, reuse after a delay) run differently enough that they get their own sessions.
Scoring
Turning responses into scores involves no black-box weighting. We use arithmetic you can follow and check yourself instead. A domain raw score is difficulty-weighted accuracy (easy 1.0, medium 1.5, hard 2.0), which we convert to a z-score under an assumed internal reference distribution and then present in the familiar 100±15 standard-score format. The z-score is clipped at ±2.0, so domain scores and the composite index are only ever displayed between 70 and 130, and percentiles only ever fall between the top 2% and the top 98%.
The composite cognitive index approximates a general factor: the mean of the eight domain z-scores, multiplied by a re-standardization factor of 1.25. Averaging several domains shrinks the variance, and 1.25 is what restores it once the domains are assumed to correlate at about r = 0.6. That is a different calculation from the attenuation correction of psychometrics, which corrects a correlation for unreliability. The wider the spread between domains, the wider the estimated range, and percentiles assume a standard normal distribution.
Working-memory items whose display was cut short because the tab was hidden are dropped from scoring, since the stimulus was never there to be seen. They leave both the numerator and the denominator, and the accuracy and sub-skill figures shown in the result as well. The number of items we will drop is capped, and a session past that cap is scored in full instead. Either way, a session with any interrupted item carries a reliability caution.
Processing speed is the one domain where response time enters an ability score directly. Its z-score is the accuracy z plus 0.3 times a speed z, and speed is measured as the time taken relative to the item’s limit. That reference distribution — a median at 60% of the limit — is an assumption made before any pilot data, and the speed component moves the domain score by no more than nine points either way. At exactly the reference speed the score is the same as under accuracy alone. An item that runs past its limit counts as wrong for accuracy and as using the full limit for speed.
Items whose timing cannot be trusted are counted conservatively. One whose tab was hidden only briefly — half a second or less — keeps its measured time; one hidden for longer, or interrupted and resumed, is not dropped from the speed sample but counted as having used the full limit, so that damaging the sample can never pay off. If more than half the processing-speed items have to be counted that way, the speed component still applies, but the result carries a reliability caution. In the other seven domains response time does not enter the score at all.
There are three supplementary indices — accuracy, cognitive efficiency and processing stability — and only the last two use response time. Accuracy is computed from correct answers alone, across the five core domains (fluid reasoning, visual-spatial processing, quantitative reasoning, verbal reasoning and working memory).
Every domain score is paired with an estimate of how much it would move on a retake, computed from the number of items and their difficulty mix. When the difference between domains stays inside that margin of measurement error, we do not commit to a profile type or a ranking of strengths; the two domains are shown side by side as candidate strengths instead.
Only once the normative sample is large enough can we estimate item difficulty and discrimination with two-parameter IRT and move to standardized scores based on the actual distribution.
A 96-item form is also available, measuring the same eight domains more finely at twelve items per domain. Its norms are not yet its own, though: it borrows the 48-item form’s provisional reference distribution, and they will be replaced once a pilot sample is collected.
Reliability and limitations
One thing needs to be stated plainly. This assessment was developed in-house from measurement principles in the research literature; it is not an accredited instrument that has been through standardization and norming. The scores are estimates against an internal reference distribution. The test and its scores must not be used for medical or clinical diagnosis, nor for selection decisions in settings such as education and hiring.
The unsupervised online setting imposes limits of its own. Device performance, screen size and the surroundings you take the test in can all affect the result. Some attention-control items require color perception and put test takers with a color vision deficiency at a disadvantage; with red–green deficiency in particular, there are items whose answer cannot reliably be told apart. While you take the test we record response times and whether the tab was hidden. Unusually fast responses, repeated interruptions and resumes, long spells away from the screen while an item was on it, nearly every timed processing-speed item running past its limit, any working-memory item whose display was interrupted, more than half the processing-speed items having to be counted at the full time limit because their timing could not be trusted, or more than thirty minutes between the immediate and delayed recall blocks all raise a reliability caution alongside the result.
References
- Zorowitz, S., Chierchia, G., Blakemore, S.-J., & Daw, N. D. (2024). An item response theory analysis of the matrix reasoning item bank (MaRs-IB). Behavior Research Methods, 56(3), 1104–1122.
- Uittenhove, K., Jeanneret, S., & Vergauwe, E. (2023). From lab-testing to web-testing in cognitive research: Who you test is more important than how you test. Journal of Cognition, 6(1), 13.
- McGrew, K. S., Schneider, W. J., Decker, S. L., & Bulut, O. (2023). A psychometric network analysis of CHC intelligence measures: Implications for research, theory, and interpretation of broad CHC scores "beyond g". Journal of Intelligence, 11(1), 19.
- Burgoyne, A. P., Tsukahara, J. S., Mashburn, C. A., Pak, R., & Engle, R. W. (2023). Nature and measurement of attention control. Journal of Experimental Psychology: General, 152(8), 2369–2402.
- Dworak, E. M., Revelle, W., Doebler, P., & Condon, D. M. (2021). Using the International Cognitive Ability Resource as an open source tool to explore individual differences in cognitive ability. Personality and Individual Differences, 169, 109906.
- Schneider, W. J., & McGrew, K. S. (2018). The Cattell–Horn–Carroll theory of cognitive abilities. In D. P. Flanagan & E. M. McDonough (Eds.), Contemporary intellectual assessment (4th ed.). Guilford Press.
- Condon, D. M., & Revelle, W. (2014). The International Cognitive Ability Resource: Measurement reliability and validity. Intelligence, 43, 52–64.
- Hampshire, A., Highfield, R. R., Parkin, B. L., & Owen, A. M. (2012). Fractionating human intelligence. Neuron, 76(6), 1225–1237.
- Cokely, E. T., Galesic, M., Schulz, E., Ghazal, S., & Garcia-Retamero, R. (2012). Measuring risk literacy: The Berlin Numeracy Test. Judgment and Decision Making, 7(1), 25–47.
- Frederick, S. (2005). Cognitive reflection and decision making. Journal of Economic Perspectives, 19(4), 25–42.
- Carroll, J. B. (1993). Human cognitive abilities: A survey of factor-analytic studies. Cambridge University Press.
- Stroop, J. R. (1935). Studies of interference in serial verbal reactions. Journal of Experimental Psychology, 18(6), 643–662.
- Spearman, C. (1904). "General intelligence," objectively determined and measured. American Journal of Psychology, 15(2), 201–292.