Cognitive Test Development
2026-09-16
What we’ve already touched on
CFIT and the Progressive Matrices were designed to reduce the influence of language and cultural background on scores. This design does not fully eliminate cultural influence, and we discuss this issue in depth today.
Generally, there are three issues that often get confused with one another: bias, cultural loading, and fairness.
Key point for today
A test can be highly culturally loaded without being proven statistically biased, and conversely can be free of cultural loading yet still biased. These are two different dimensions.
Most major intelligence and cognitive ability tests, when retested, are not shown to be statistically biased in this sense — but that doesn’t mean the test is automatically free of cultural loading or fair when used for cross-group comparisons.
Connection to Meeting 6
This is a further chapter in the history of intelligence-test misuse we already discussed in Meeting 6. This issue is closely tied to decisions (school placement) made on the basis of test results.
DIF is a statistical indicator, but it is not automatically proof of “bias.” An item flagged for DIF still needs to be reviewed qualitatively to confirm the cause is genuinely unfair content, rather than simply a real difference in ability between groups.
A sound bias investigation starts with item-level DIF, moves on to factor invariance, and ends with a predictive-bias test at the total-score level.
Definition of fairness
“A fair test does not advantage or disadvantage some individuals because of characteristics irrelevant to the intended construct.” — Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014, p. 50)
Fairness covers more than just statistical test results: test design, administration procedures, and score interpretation must all give test takers an equal chance to demonstrate their ability.
An assumption that’s often not met
When we compare a test taker’s score to a norm, we assume the test taker’s level of acculturation is comparable to that of the standardization sample. When a test taker’s background and experience differ greatly from the norm sample, using that norm to evaluate or predict their performance may not be appropriate.
| Approach | Brief description | Main limitation |
|---|---|---|
| Modified/adapted tests | Removing/simplifying some items or instructions, dropping time limits | Violates standardized procedure → introduces uncontrolled error |
| Translator/interpreter | Items are translated directly during administration | Items remain tied to the original culture; standardization is still violated |
| Native-language tests | Tests actually developed and standardized in the test taker’s language | Norms often come from monolingual speakers in another country, not representative of bilingual test takers in the destination country |
| Non-verbal tests | Items designed to minimize language demands | Still culturally loaded through visual stimuli and gestural instructions |
A survey of school psychologists found that 88% chose non-verbal tests, 40% used a translator, and 20% used native-language tests when assessing culturally and linguistically diverse individuals.
So, far fewer actually check score validity directly.
Empirical finding
A study of three popular non-verbal tests commonly used to identify gifted children found a large score gap between English language learner (ELL) and non-ELL students.
Conclusion: non-verbal tests don’t automatically “level the playing field” for children from different cultural backgrounds.
Rather than simply assuming a test is “culture-free” or not, this framework maps the degree of cultural loading and language demand for each subtest into a 2×2 matrix:
| Low language demand | Moderate language demand | High language demand | |
|---|---|---|---|
| Low cultural loading | Least-affected subtests | ||
| Moderate cultural loading | Analogical reasoning subtests | ||
| High cultural loading | Spatial memory subtests | Most-affected subtests |
Managing bias in a test, not eliminating it entirely
This framework helps practitioners choose the subtests whose cultural and language demands best match the test taker’s background.
This isn’t a way to eliminate cultural influence entirely — under this framework, every test is culturally loaded to some degree.
But what can be pursued is managing that loading deliberately when choosing instruments and interpreting scores.
That said, this doesn’t mean we MUST stop using these tests — rather, it’s a reason to read test manuals critically and to honestly report their limitations when interpreting scores. This also aligns with the Indonesian Psychological Code of Ethics discussed in Meetings 1 and 6.
Case
A psychologist uses the CFIT (normed in the United States and Europe) to assess children in a village rarely exposed to standardized psychological tests, including multiple-choice formats and timed tasks. The psychologist concludes that the low scores indicate low cognitive ability.
Discuss with your group:
Any questions?
These slides were prepared using and Quarto with a template from UNAIR Theme.
American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA.
Franklin, T. (2018). Best practices in multicultural assessment of cognition. In R. S. McCallum (Ed.), Handbook of nonverbal assessment (2nd ed., pp. 46–53). Springer.
Maller, S. J., & Pei, L.-K. (2018). Best practices in detecting bias in cognitive tests. In R. S. McCallum (Ed.), Handbook of nonverbal assessment (2nd ed., pp. 29–45). Springer.
McCallum, R. S. (2018). Context for nonverbal assessment of intelligence and related abilities. In R. S. McCallum (Ed.), Handbook of nonverbal assessment (2nd ed., pp. 12–28). Springer.
Ortiz, S. O., Piazza, N., Ochoa, S. H., & Dynda, A. M. (2018). Testing with culturally and linguistically diverse populations: New directions in fairness and validity. In D. P. Flanagan & E. M. McDonough (Eds.), Contemporary intellectual assessment: Theories, tests, and issues (4th ed., pp. 684–735). Guilford Press.