How Are Online Intelligence Tests Changing the Way People Measure Cognitive Skills?

Intelligence Tests

Cognitive measurement used to be scarce, expensive and normally given once to a test-taker by an experienced examiner in a quiet room. After the test, it would take some time to score the test and to write a report about the test-taker’s cognitive abilities. This model is still in place for clinical purposes as well as for high-stakes selection purposes. However, for most other purposes, there is now an abundance of cognitive data that is generated in an almost inexpensive and automated fashion online, even without the involvement of the test-taker’s institution.

Those numbers and reports, generated by instrumentation typically under the control of and used by the researcher, have meaning; they can be “owned” by someone. Three implications of online testing follow.

What unsupervised delivery actually does to the measurement

Online testing is not simply equivalent to a paper test that is filled out on a screen. For one, test timing at the item level is more precise than ever before. Response time data can be used in a manner that was previously impossible, such as processing time in the case of speeded test items. Additionally, the adaptive test that is administered online in half the time of its paper equivalent maintains equivalent reliability at the midpoint of the distribution.

The losses, however, are equally real. Online testing is conducted in virtual isolation. Because a test-taker cannot be verified to be who he claims to be, as little as his best effort, and in what environment he is taking the test, online testing is inherently compromised. There is considerable noise in timed responses, particularly when there are millisecond-level differences, due to variability in test-takers’ computers and devices; and a single pause of even a few seconds to answer a call or to get up from the computer to get something from another room can severely contaminate a block of speeded test items.

Which constructs travel well and which do not

  • Matrix reasoning and figural analogies transfer cleanly, since item exposure is the main threat and large calibrated banks reduce it.
  • Working memory span tasks hold up reasonably well when scored on accuracy rather than raw speed.
  • Verbal comprehension is vulnerable, because lookup is trivial and hard to detect in an untimed format.
  • Processing speed measures are the most device sensitive and should be interpreted within, not across, hardware categories.
  • Anything requiring examiner judgment, such as qualitative error analysis, does not transfer at all.

Comparing delivery models on the criteria that matter

The key trade-off between the available tests is between the level of test control versus cost and use. The following table summarizes the primary characteristics of the tests on each of the key dimensions.

 

Delivery model Identity assurance Cost per candidate Best suited to
In person, examiner administered High High Clinical diagnosis, forensic work, accommodation decisions
Remote proctored, live or recorded Moderate to high Moderate Certification, selection at final stages, licensure
Unsupervised with verification retest Moderate Low Early stage screening where a confirmation step follows
Fully open self administered None Very low Self assessment, practice, research at scale, learning feedback

Why people take these tests when nothing is riding on the result

A large part of online cognitive tests is taken voluntarily. This would have been considered uncivilized 30 years ago. Currently people like to get a check on their cognitive functioning from time to time, and sitting a couple of well validated online intelligence tests is a reasonable way to get a sense of where they are in terms of preparation for a test or study, or even to get a sense of improvement over time after, for example, having been ill with something that affected attention.

Designing feedback that teaches rather than labels

Self-assessment has limited value unless the resulting report tells the test-taker something concrete that he or she can do with the information. Reporting a single composite score to the test-taker does little to help. Reporting subscale scores and describing in the report what relative weaknesses in spatial reasoning (e.g., rotation vs. Surface feature extraction) or in verbal reasoning (e.g., analogy vs. Matrix reasoning) correspond to in terms of the test-taker’s real-world abilities (e.g., planning studies, giving presentations) is far more valuable.

Confidence intervals help to report scores in “report format” in consumer-facing reports instead of just raw scores and technical reports; instead of a single score, a report can include a number of subscales with corresponding scores and standard deviations.

Reading scores from uncontrolled conditions

Treat unsupervised testing results with the respect due a hypothesis tested with data subject to large error. Such results are useful for establishing the existence of very large effects and nothing else.

  1. Discount any single administration that shows unusual response time patterns, such as near-instant answers on difficult items.
  2. Require two sessions before treating a result as a baseline for tracking change over time.
  3. Compare against a norm group matched on delivery mode, never against paper norms collected under supervision.
  4. Assume a practice effect on repeat administration and use alternate forms where the bank allows it.
  5. Separate the screening decision from the confirmatory one, and never let an unsupervised score make the final call.

What this means for practice

The big question in the field of online testing has shifted from whether or not online testing is valid to what kinds of questions can be answered with different types of testing. Given that volume can be handled with cheap online testing and that control is expensive, online testing should be the means by which most testing is done, with expensive assessment used sparingly in instances in which errors could have grave consequences for the test-taker.

Scroll to Top