Research note · 31 August 2026
MOSAIC-48 pilot results: first 88 completed sessions
What the first anonymous sessions show about performance and response time—and how those findings are shaping the service.
Observation period: –
Pilot results
- Completed sessions
- 88
- Observed test locales
- 7
- Mean session accuracy
- 80.8%
- Timing responses retained
- 95.3%
Across 88 completed sessions, mean session accuracy was 80.8%. Each session carries the same weight in this average.
10 formal sessions were rescored with the current code. Their mean internal score was 112, the median was 113, and 6 passed the current reliability check.
Sample and method
- 88 completed attempts are represented: 78 mini-form sessions and 10 formal-form sessions across 7 observed test locales. Each attempt is counted separately.
- Accepted sessions used a supported form and bank revision, matched registered item IDs and revisions, and contained a complete answer set with a valid locale.
- Only aggregate results are public. Session records remain private and are not linked to names or accounts.
How results were calculated
- Answer accuracy and formal scores were recalculated with the current scoring code. The headline accuracy is the same-weight mean of completed-session accuracy.
- The assessment method documents the internal reference, scoring rules, and reliability checks used for this snapshot.
Response-time analysis
The timing distribution retains 1,079 of 1,132 item responses. The completed-session count is unchanged.
| Filter | Responses |
|---|---|
| Compromised timing | 7 |
| Rapid-guess rule | 25 |
| Statistical outlier rule | 21 |
| Retained | 1,079 |
Current item-specific NT10 threshold: 10% of expected time, bounded from 1.5 to 10 seconds.
Tukey 1.5×IQR fence on log response time within combined mini/formal locale-by-part groups; groups smaller than 8 received no statistical-outlier decision.
How to read these results
This snapshot describes early use of the current MOSAIC-48 forms. Formal scores use the service’s provisional internal reference, not a population norm. A later representative study will evaluate norms and validity.