Cognitive Test Development
2026-09-16
Final project: each group will develop one complete set of cognitive tests, from item specifications through the test booklet and manual, supported by evidence of validity and reliability. For this, you are asked to form groups of at most 5 members.
Before the midterm (Meetings 1–7): building the conceptual foundation. What a cognitive test is, the stages of developing one, what a taxonomy of instructional objectives is, and the fundamental differences among achievement tests, aptitude tests, and intelligence tests.
After the midterm (Meetings 9–15): working on the project. Writing HOTS items, developing specifications and a blueprint, collecting pilot-test data, analyzing items, estimating reliability and validity, and developing norms and the manual.
An elementary school wants to know whether its new literacy program is working: is this year’s 5th-grade students’ reading ability better than that of the students who entered last year, after the same material was taught?
A startup is screening 500 applicants for a programming internship before seeing a single line of code they’ve written, because they want to know who has the greatest potential to learn coding quickly.
A neuropsychologist needs to establish a baseline for a patient’s general reasoning ability after a head injury, before being able to assess whether there has been a decline in cognitive function.
All three alike require a cognitive test, because they measure a person’s maximal performance, not their tendencies or preferences.
But the three ask conceptually different questions: what has already been mastered (situation 1), what could potentially be mastered (situation 2), and how large a person’s general capacity is (situation 3).
If the wrong type of test is used — for instance, using an achievement test to predict potential — the conclusions can be completely wrong.
These three situations will keep coming up throughout the course in the form of achievement tests (Meeting 4), aptitude tests (Meeting 5), and intelligence tests (Meeting 6).
From the Psychometrics course, you have already learned the following:
Important to remember
You have already learned how to evaluate whether a test performs well (validity, reliability). This course will help you learn how to build one from scratch, but specifically for the cognitive domain.
Classic definition
Psychometrics is the branch of science concerned with the theory and technique of measurement of educational and psychological attributes (Kline, 1986).
Kline (1986) formulated two main roles of psychometrics:
Validity is a property of the test
International standards (AERA, APA, & NCME, 2014) emphasize that test quality does not stop at the measurement tool itself — validity is a property of the interpretation and use of scores, not merely a property of the test itself.
The implication
Because the attribute is latent, no test measures it perfectly. Every test is a sample of behavior that we generalize to a broader construct. This is why validity and reliability are always central issues in psychological measurement.
| Term | Brief definition |
|---|---|
| Test | An objective, standardized measurement of a sample of behavior (Anastasi & Urbina, 1997) |
| Scale | An instrument for identifying psychological constructs/attributes, usually non-cognitive |
| Inventory | A tool for estimating and assessing behavior, interests, or preferences |
| Questionnaire | A set of questions about a topic, answered based on participants’ self-report |
The rule we’ll use throughout the semester
If there is a right-or-wrong answer that can be objectively scored from performance → it’s a test. If what’s being measured is a tendency or preference with no right-or-wrong answer → it’s usually a scale or inventory.
(ability test → cognitive test)
(personality test → non-cognitive)
Why this matters for this course
This entire course lives in the left-hand column. The principles for writing items for maximal performance (with an objective answer key) are fundamentally different from the principles for writing personality-scale items — don’t mix them up later when writing items in Meeting 7.
Classify the following three examples:
Answer key
A legacy that was not neutral
The results of the US Army tests were later misused by Carl Brigham (1923) to support immigration policies that discriminated against certain ethnic groups. However, this conclusion was later retracted by Brigham himself. This is why the Indonesian Psychological Code of Ethics, which we’ll discuss later, strictly limits who is permitted to use psychological tests.
Why this SAT anecdote matters
This is proof that distinguishing “achievement test” from “aptitude test” is not merely a classroom definition exercise. Even a nationally administered test can carry the wrong label for decades.
| Type | Measures | Nature | Example research question |
|---|---|---|---|
| Intelligence test | General potential for solving problems and adapting to new situations | Relatively stable across time and context | “How large is this student’s general learning capacity?” |
| Aptitude test (aptitude) | Capacity to learn a specific skill in the future | Predictive — judged by its ability to predict future performance | “Does this candidate have the potential to become a good programmer?” |
| Achievement test (achievement) | What has already been learned/mastered at present | Retrospective — judged by how well it matches the content taught | “How well has the student mastered the Chapter 3 material?” |
Intelligence
Aptitude
Achievement
Ethical note
Most of the instruments above are commercially licensed psychological tests with restricted use. In this course we only discuss how they work and their underlying principles.
Example of an original non-verbal item (not taken from any actual test): what pattern fills the cell marked with a question mark?
| ● | ● ● | ● ● ● |
| ● ● | ● ● ● | ● ● ● ● |
| ● ● ● | ● ● ● ● | ? |
Why this is an example of maximal performance
There is a logical rule (number of dots = row number + column number − 1), an objectively correct answer (C), a clear stimulus, and no language is required — the hallmark of a non-verbal ability test item, well-suited to measuring reasoning that is relatively free of cultural bias.
In Indonesia, psychological tests are not free for just anyone to use. The Indonesian Psychological Code of Ethics (HIMPSI, 2010) divides them into 4 categories, based on the qualifications needed to administer, interpret, and report their results:
Relevance to this course’s CPL
This course’s CPL (Graduate Learning Outcomes) explicitly names the authority to use Category A and B tests. This is the ethical boundary on what kinds of tests you may develop and use as a prospective psychology graduate.
Imagine a journal abstract states:
“…participants completed a timed assessment of non-verbal reasoning ability to control for variation in general cognitive ability before the experimental manipulation was administered…”
Practice for Assignment 1
In Week 6, you will review 3 journal articles that use a cognitive test as their research instrument. Practicing this kind of identification is a core skill assessed in that assignment.
Next: the stages of developing a cognitive test
If today we talked about what a cognitive test is, Meeting 2 will cover how to develop one. We’ll start with needs analysis and construct definition, through to an initial blueprint.
Any questions?
These slides were prepared using and Quarto with a template from UNAIR Theme.
AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association.
Anastasi, A., & Urbina, S. (1997). Psychological testing (7th ed.). Prentice Hall.
Anastasi, A., & Urbina, S. (2016). Tes psikologi (Edisi ke-7). PT Indeks.
Azwar, S. (1987). Tes prestasi: Fungsi dan pengembangan pengukuran prestasi belajar. Liberty.
Azwar, S. (2019). Konstruksi tes: Kemampuan kognitif. Pustaka Pelajar.
Binet, A., & Simon, T. (1905). Méthode nouvelle pour le diagnostic du niveau intellectuel des anormaux. L’Année Psychologique, 11, 191–244.
Brigham, C. C. (1923). A study of American intelligence. Princeton University Press.
DuBois, P. H. (1970). A history of psychological testing. Allyn & Bacon.
Himpunan Psikologi Indonesia (HIMPSI). (2010). Kode etik psikologi Indonesia.
Kline, P. (1986). A handbook of test construction: Introduction to psychometric design. Methuen.
Terman, L. M. (1916). The measurement of intelligence. Houghton Mifflin.
Yoakum, C. S., & Yerkes, R. M. (1920). Army mental tests. Henry Holt.