Psychological Scale Construction
2026-09-14
When we use a Likert scale, we quietly accept three things:
All items carry equal weight. Item 1 and item 7 are each worth a maximum of 5 points.
Items are abstract self-descriptions. “I am a careful person.”
Each item is answered independently. Agreeing with one item does not force you to reject another.
Three formats, each of which gives up one of these — and pays a price for it.
| Format | What it gives up | What it gains |
|---|---|---|
| Thurstone | Equal item weights | Items calibrated across the continuum |
| SJT | Abstract self-description | Concrete context and judgement |
| Forced-choice | Independent responding | Protection from social desirability |
All three are conceptually attractive, but far harder to build than Likert, and two of them break assumptions we established in Session 3.
A tuning fork is built to vibrate at one frequency. Put it near a tone of that frequency and it vibrates; at any other frequency it stays still.
Imagine an array of tuning forks arranged from low to high frequency. To identify a tone’s frequency, you just see which fork vibrates.
A Thurstone scale works the same way: each item is “tuned” to respond to a particular level of the attribute.
Contrast this with Likert, where every item is essentially the same detector and only the degree of the answer varies.
Generate many statements about the attitude object — usually dozens or hundreds.
Ask a panel of judges to sort each statement into 11 piles, from most opposed to most in favour.
Take the median judge rating for each statement.
Look at the spread of judge ratings (the interquartile range). Statements the judges disagreed about are ambiguous, and get discarded.
Select from what remains so that the scale values are evenly spread across the continuum (\(\theta\)).
We have established each item’s “difficulty” before a single participant has answered. CTT cannot do that.
participants are only asked to agree or disagree with each statement. There are no degrees.
A participant’s score is the mean scale value of the statements they endorsed.
So someone endorsing statements with scale values 8.5 and 10.5 scores 9.5 — nothing is summed.
A fundamental difference from Likert
In Likert, the information is in the degree of the answer. In Thurstone, the information is in which items were endorsed — the answer itself is only yes or no.
The workload is heavy. You need a judge panel before any participant data is collected.
Finding items that genuinely “resonate” at a specific level is much harder than it sounds.
Nunnally (1978), as cited by DeVellis, judged that the practical problems often outweigh the advantages, unless a researcher genuinely needs that kind of calibration.
And there is a theoretical consequence
Like Guttman in Session 4, Thurstone items are not equally weighted. So the parallel and tau-equivalent models from Session 3 do not apply, and \(\alpha\) cannot simply be used.
Notice the resemblance
The CVI procedure is Thurstone’s logic applied to content relevance, rather than to position on a continuum.
participants are given a realistic scenario and asked to evaluate several possible actions.
What is measured is not a self-description but judgement in a concrete situation.
Rooted in personnel selection testing since the 1940s, and revived after Motowidlo et al. (1990) described it as a low-fidelity simulation.
Situation. You are leading a five-person group assignment. Two days before the deadline, one member tells you they have not started their section at all because of a family emergency.
How effective is each of the following actions?
| Action | Response (1: not at all effective; 5: very effective) | |
|---|---|---|
| A | Do that section yourself so the deadline is met | 1–5 |
| B | Redistribute the section to members who have finished | 1–5 |
| C | Ask the lecturer for an extension | 1–5 |
| D | Submit as is and note who did not contribute | 1–5 |
Measures behavioural tendency.
Closer to personality.
More open to impression management.
Measures knowledge of effective action.
Closer to cognitive ability.
More resistant to impression management.
Ployhart & Ehrhart (2003) showed that changing the response instruction changes the test’s psychometric properties — both its reliability and the construct it measures.
Consensus-based — the key comes from the average judgement of a large sample.
Expert-based — a panel of practitioners decides which actions are effective.
Theory-based — the key is derived from theory about the construct.
The first two require a panel, exactly like Thurstone and like the CVI in Session 11.
Concrete context. participants are not asked to generalise about themselves but to react to a specific situation.
Face relevance, which makes it easier to accept in a selection setting.
Reasonable prediction of performance. McDaniel et al. (2001) show SJTs predict job performance adequately.
SJTs are almost always multidimensional. A single scenario can involve knowledge, personality, and reasoning at once.
As a result, internal consistency is often low.
This violates the unidimensionality assumption from Session 2, and means the CTT model from Session 3 needs an extra layer of complexity.
A mistake people make often
Reporting Cronbach’s alpha for an SJT and concluding the test is poor because the number is low.
\(\alpha\) assumes every item is driven by one shared latent variable. In an SJT that assumption is not even intended to hold. The wrong coefficient produces the wrong conclusion, not a bad instrument.
Recall the sixth assumption from Session 2: participants are willing to report their internal, psychological state honestly.
In a job selection setting that assumption is plainly unrealistic. Everyone has a reason to look good.
With a Likert format, a participant can simply rate every attractive-sounding statement highly.
The core idea
If participants are forced to choose between two equally attractive statements, they can no longer endorse everything. They must reveal which one describes them better.
Choose the statement that describes you best:
| (a) | I like to finish my work on time. |
| (b) | I like to help other people solve their problems. |
Both statements are equally positive. Neither is obviously the “better” answer.
Both measure different constructs — roughly conscientiousness and agreeableness here.
Variants use blocks of three or four statements with “most like me” and “least like me” instructions.
The statements in a pair must be matched for attractiveness beforehand — which requires its own preliminary study.
Without that matching, the format solves nothing: participants will still pick whichever sounds better.
Matching attractiveness is itself a calibration task.
Every time a participant picks one statement, they automatically do not pick the other.
As a result, the total across all scales is constant for every participant.
Data with this property are called ipsative: the scores are meaningful within a person, not between people.
With ipsative data, a score answers “of all these things, which stands out most in me?” — not “how much of this trait do I have compared with other people?”
Because the total is forced to be constant, correlations between scales are forced negative — even if the underlying constructs are entirely unrelated.
For an ipsative instrument with \(k\) scales, the average intercorrelation is:
\[\bar{r} = -\frac{1}{k-1}\]
| Number of scales (\(k\)) | Forced average correlation |
|---|---|
| 3 | −0.50 |
| 5 | −0.25 |
| 10 | −0.11 |
Those numbers come from the format, not from the participants’ psychology.
Between-person comparison becomes invalid. A conscientiousness score of 8 for person A cannot be compared directly with 8 for person B.
Factor analysis becomes misleading, because the correlation structure was distorted by the format before any analysis began.
Classical reliability coefficients become hard to interpret — the CTT model in Session 3 assumes scores are free to vary, and here they are not.
Ipsative instruments should not be used for selection, ranking, or comparison between individuals.
Brown & Maydeu-Olivares (2011) showed that with IRT modelling, normative scores can be recovered from forced-choice responses.
The approach is called Thurstonian IRT, because it builds on Thurstone’s model of paired comparisons.
So the format remains usable — provided the analysis uses the right model rather than simple summing.
Thurstone’s 1920s idea turns out to be the solution to this problem.
Thurstone — still realistic, but technically harder than building a Likert scale.
SJT — not realistic. Writing good scenarios and building a scoring key is a project in itself.
Forced-choice — not realistic. The correct analysis of ipsative data is outside the scope of this course.
So why cover them?
Because you will encounter all three in practice — especially SJTs and forced-choice instruments, which are widely used in assessment and recruitment.
What you should take from today is not the ability to build them, but the ability to judge whether such an instrument is being used properly.
Attitude object: using AI to complete coursework.
Rate each statement from 1 (strongly opposed) to 11 (strongly in favour). Judge where the statement sits, not your own opinion.
Compute the median class rating for each statement. That is its scale value.
Compute the interquartile range for each. Which statement produced the most disagreement?
Which statement would you discard, and why?
Are the surviving statements evenly spread across the continuum, or bunched on one side?
Notice that you have just run the same procedure you will use for the CVI in Session 11 — only the criterion differs.
Look at these three forced-choice pairs. Which are well matched, and which are not?
Pair 1 (a) I am an organised person. (b) I am an outgoing person.
Pair 2 (a) I like to plan things in advance. (b) I sometimes put things off.
Pair 3 (a) I enjoy working alone. (b) I enjoy working in a team.
For any pair that is not well matched, write down how you would fix it.
Thurstone calibrates items in advance using a judge panel, so items have known positions on the continuum.
Thurstone scale values are the ancestor of the IRT difficulty parameter, and its judging procedure is the ancestor of CVI.
SJTs measure judgement in context and are almost always multidimensional — so a low \(\alpha\) is not a sign of failure.
Response instructions change what an SJT measures. “Would do” and “should do” are not the same thing.
Forced-choice protects against social desirability, but produces ipsative data.
Ipsative data forces negative correlations between scales and makes between-person comparison invalid.
After five sessions on concepts and formats, we start computing: reliability.
Preparation
Read DeVellis & Thorpe (2022), Chapter 3. To see it in practice, the Reliability chapter of Bikos lays the coefficients out side by side.
Questions?
These slides were made using and Quarto, with a template from UNAIR Theme.
Brown, A., & Maydeu-Olivares, A. (2011). Item response modeling of forced-choice questionnaires. Educational and Psychological Measurement, 71(3), 460–502.
DeVellis, R. F., & Thorpe, C. T. (2022). Scale development: Theory and applications (5th ed.). SAGE.
Edwards, A. L. (1959). Edwards Personal Preference Schedule manual. Psychological Corporation.
Hicks, L. E. (1970). Some properties of ipsative, normative, and forced-choice normative measures. Psychological Bulletin, 74(3), 167–184.
McDaniel, M. A., Morgeson, F. P., Finnegan, E. B., Campion, M. A., & Braverman, E. P. (2001). Use of situational judgment tests to predict job performance. Journal of Applied Psychology, 86(4), 730–740.
Motowidlo, S. J., Dunnette, M. D., & Carter, G. W. (1990). An alternative selection procedure: The low-fidelity simulation. Journal of Applied Psychology, 75(6), 640–647.
Nunnally, J. C. (1978). Psychometric theory (2nd ed.). McGraw-Hill.
Ployhart, R. E., & Ehrhart, M. G. (2003). Be careful what you ask for: Effects of response instructions on the construct validity and reliability of situational judgment tests. International Journal of Selection and Assessment, 11(1), 1–16.
Thurstone, L. L. (1928). Attitudes can be measured. American Journal of Sociology, 33(4), 529–554.
Thurstone, L. L., & Chave, E. J. (1929). The measurement of attitude. University of Chicago Press.
Whetzel, D. L., & McDaniel, M. A. (2009). Situational judgment tests: An overview of current research. Human Resource Management Review, 19(3), 188–202.