Psychological Scale Construction
2026-08-25
All of the above share the same shape: we take these assumptions to be true without ever actually checking them.
Psychological measurement is indirect. We work through proxies.
The definition of measurement is itself contested, and Michell’s objection — whether psychological attributes are genuinely quantitative — has not really been resolved.
Most psychological data are ordinal but are treated as interval.
What decides whether your item set is a scale is theory, and software output doesn’t tell you anything except the numbers.
Because these assumptions never appear in a research report. Only the results do.
Because software will not warn you when an assumption is violated. It will simply return a number.
Because in Sessions 9–15 you will make technical decisions that all rest on these assumptions.
Remember, measurement is never truly value-free or objective. It is always theory-laden, so measurement results must be read within the framework of the assumptions that come with them.
No measurement is assumption-free. Even a thermometer assumes that mercury expands linearly with temperature.
Assumptions are the price we pay for being able to measure anything at all.
The problem is not that assumptions exist. The problem is assumptions nobody noticed making.
The distinction that matters
“I am assuming this attribute is continuous, and here is why” is a scientific statement.
“I summed the items because that is how it is done” is not.
The assumption holds. No problem.
The assumption is violated, but the conclusion barely changes. A method like this is called robust to that violation.
The assumption is violated, and the conclusion changes completely. This is the case we need to deal with.
| # | Assumption | Checkable with data? |
|---|---|---|
| 1 | The attribute varies quantitatively | Very hard |
| 2 | Its distribution is normal | Yes |
| 3 | The attribute is continuous, not categorical | Partly |
| 4 | All items measure one and the same thing | Yes |
| 5 | Items work the same for every group | Yes, with enough data |
| 6 | Respondents can and will answer accurately | Partly |
Note
Note that the most fundamental assumption (number 1) is also the hardest to check.
Imagine a 10-item anxiety scale, each item scored 1–5. You sum them into a single score from 10 to 50.
In summing, you have already assumed at least three things:
Every item carries equal weight. Item 1 counts exactly as much as item 7.
The distances between response categories are equal. The step from 1→2 is treated as the step from 4→5.
The scores are addable. Someone scoring 40 is treated as having “twice” the anxiety of someone scoring 20 — or at minimum, a 10-point difference means the same thing everywhere on the scale.
The item “I sometimes feel nervous before a presentation” and the item “I often feel I am going to die from panic” clearly indicate different levels of severity.
But under ordinary summing (unit weighting), both are worth a maximum of 5 points.
Our scale therefore treats those two statements as equivalent indicators.
But…
This is precisely what separates Classical Test Theory from Item Response Theory. IRT lets each item have its own difficulty and discrimination. We will get a brief look at it in Session 3.
This course uses CTT, so unit weighting is an assumption you will be making. What matters is that you know you are making it.
Adolphe Quetelet (1835) introduced l’homme moyen, “the average man”. He applied the normal curve — at the time a model of observational error in astronomy — to human characteristics.
The implication: the normal curve moved from being a model of error to being a model of people.
Galton then extended it to psychological traits, and normality has been built into psychometrics ever since.
Worth noticing
There was never strong evidence that psychological attributes must be normally distributed. What there was, was a habit inherited from a different context.
In scale construction we discard items that almost everyone endorses — where nearly everyone answers “5/strongly agree” (e.g., “I believe God truly exists”) — or that almost nobody endorses — where nearly everyone answers “1/strongly disagree” (e.g., “Premarital sex is a moral obligation everyone must fulfil”).
The items that survive are the ones that split respondents down the middle.
The result: the distribution of total scores becomes more symmetric and piles up in the centre.
Note the order of operations
We select items so the distribution comes out normal, and then treat that normality as a finding about people. Much of it is a consequence of how we built the scale.
IPIP Big Five data (data/data.csv in this course repository), 50 items, 1–5 scale, n = 19,719 respondents (the raw row count; Session 6 uses n = 19,718 after listwise deletion).
First, three individual items:
| Item | Content | Mean | Skewness | Kurtosis |
|---|---|---|---|---|
| N3 | I worry about things | 3.84 | −0.83 | −0.16 |
| O2 | I have difficulty understanding abstract ideas | 2.15 | +0.78 | −0.23 |
| E1 | I am the life of the party | 2.63 | +0.21 | −0.96 |
Note
The items are clearly not normal. N3 piles up at the agreement end (69% answer 4 or 5); O2 piles up at the opposite end (67% answer 1 or 2).
Total score per dimension (10 items, range 10–50), same dataset:
| Dimension | Mean | SD | Skewness | Kurtosis |
|---|---|---|---|---|
| Extraversion | 30.11 | 9.22 | −0.04 | −0.71 |
| Neuroticism | 30.97 | 8.62 | −0.08 | −0.60 |
| Conscientiousness | 33.47 | 7.31 | −0.09 | −0.38 |
Skewness is close to zero — the distributions have become symmetric.
But kurtosis is negative in every case: the distributions are flatter than a normal curve.
The individual items are skewed in opposite directions. When summed, those skews cancel each other out.
The result looks symmetric, but that is not evidence that the attribute is normal. It is an effect of aggregation.
And the distribution is still not quite normal: it is flatter, with respondents spread more evenly across the middle than a normal curve would predict.
The main point
A total-score distribution that “looks normal” is not evidence that the underlying psychological attribute is normally distributed. It is largely a consequence of adding up many items.
Micceri (1989) examined 440 large-sample distributions from achievement tests and psychometric instruments.
The title of the paper: “The unicorn, the normal curve, and other improbable creatures.”
His finding: not one of those distributions met the criteria for normality. Most showed asymmetry, heavy tails, or multimodality.
Why this matters for your project
Norms and standard scores depend on this assumption. The percentiles, z-scores, and T-scores you will construct in Session 14 are derived by assuming a particular distributional shape.
If the empirical distribution is not that shape, your norms will misplace people — and worst of all at the extremes, which is exactly where decisions get made.
People differ in degree.
There is no natural boundary; any cut-off is one we impose ourselves.
Example: height.
There really are two naturally distinct groups.
The boundary is not the researcher’s invention.
Example: pregnant / not pregnant.
The question is not “which is more convenient” but which is true for this construct.
That is an empirical question, and there is a method for it: taxometrics, developed by Paul Meehl.
Haslam, Holland, & Kuppens (2012) reviewed 177 taxometric articles (311 findings) on personality and psychopathology constructs.
Their conclusion: most constructs appear dimensional. Once confounds were controlled, they estimated only about 14% of findings were genuinely taxonic.
That means most of the “disorders” we treat as categories are better described as matters of degree.
Connect this to Session 1
This is also the problem baked into diagnostic and classification systems for mental disorders: a categorical classification system applied to constructs whose evidence points to a continuum. It is also what motivated the Research Domain Criteria (RDoC) project.
If your construct is dimensional but you force categories (“high/medium/low”), you throw away information and create a boundary that does not exist in nature.
If your construct is taxonic but you scale it, your middle scores describe a person who does not actually exist.
Most psychological scales — including the one you are about to build — assume dimensional. That is a common assumption, but it is still one that needs to be tested.
Recall the logic of a scale from Session 1: the latent construct is the cause of the item responses.
If one common cause drives ten items, those ten items must correlate with each other.
Those inter-item correlations are the only observable trace of a variable we cannot observe.
If we are assuming a “common cause” is what makes the items correlate, those items should not correlate with each other for any other reason (local independence).
The consequence
Because of this, inter-item correlations can be used to estimate how strongly each item relates to its latent construct — even though we never measure the construct directly.
Unidimensionality: every item in a scale is driven by one latent variable.
If two constructs are actually mixed into one scale, the total score becomes a blend with no clear meaning.
Two people can obtain exactly the same total score for entirely different reasons.
A common example
A “psychological wellbeing” scale that mixes items about mood with items about financial situation. The two may correlate, but they are not one thing.
A score of 40 might mean “calm mind, struggling financially” or “financially secure, anxious”. Same number, very different people.
Once we account for the latent variable, the items should no longer be related to each other.
In other words: the entire reason two items correlate should come from the latent construct, not from anything else.
The most common violation: two items that say nearly the same thing.
An example of a violation
“I often feel sad” and “I frequently feel down”.
These will correlate strongly — but partly because they are almost the same sentence, not only because of depression. Your reliability will look good for the wrong reason.
We will cover how to avoid this when we write items in Session 10.
We assume the same item means the same thing to men and women, to students and workers, to respondents from different ethnic backgrounds (e.g., Javanese and Sulawesi respondents).
If that fails, comparing scores between groups is not legitimate — like comparing temperatures read off two thermometers with different scales (e.g., Kelvin and Celsius).
In psychometrics this property is called measurement invariance; a violation at the item level is called differential item functioning (DIF).
How would you test this?
The formal test requires more complex technical skills and a large sample, so we will not cover it in this course. What you need now is awareness that the assumption is there.
Heine, Lehman, Peng, & Greenholtz (2002) demonstrated a serious problem with cross-cultural Likert comparisons.
When someone answers “I am a conscientious person”, they judge themselves relative to the people around them — not against a universal standard.
As a result, respondents from a culture with a higher local standard may rate themselves lower, even when their behaviour fits the construct better.
The implication is uncomfortable
Cross-national comparisons of mean self-report scores can come out backwards relative to reality. And nothing in the data looks wrong.
Many scales in use in Indonesia are adaptations of English-language instruments.
Back-translation ensures the words match. It does not ensure the meaning matches, still less that the item functions the same way for Indonesian respondents.
If your group chooses to adapt an existing scale — one written in a foreign language and validated empirically in another country — that is not a translation job, it is a psychometric one.
A question your group must answer
Who is the target population for your scale, and does your item make sense for every subgroup within it? We will use this question again when setting the pilot sample in Session 12.
A self-report scale presumes two things at once:
Respondents are able to observe their internal states and report them.
Respondents are willing to report them honestly.
Both can fail, and they fail in different ways.
Nisbett & Wilson (1977), in a paper titled “Telling more than we can know”, showed that people often have no access to the mental processes actually driving their behaviour.
Asked for their reasons, they still answer — but the answer is often a plausible lay theory rather than a report of what actually happened.
Implication for item writing
Items asking about concrete behaviour (“How many times a week do you …”) are generally more dependable than items asking about causes (“Why do you …”).
We will apply this principle when writing items in Session 10.
Respondents tend to answer in ways that make them look good — sometimes deliberately, often not.
The effect is strongest for constructs with moral weight: honesty, prejudice, religious observance, substance use, sexual behaviour.
Guaranteeing anonymity helps, but does not remove it.
What your group should anticipate
If your construct has a socially “correct” answer, your score distribution will pile up at one end (e.g., a ceiling or floor effect). You will struggle to distinguish respondents from one another because your data will have low variance.
Acquiescence: the tendency to agree with any statement, regardless of content.
Extreme responding: the tendency to choose the endpoints (1 or 5).
Midpoint responding: the tendency to sit at 3 and avoid committing.
All three differ across cultures, so they compound the invariance problem from the previous assumption.
One partial fix, with a cost of its own
Acquiescence is usually dampened by including reverse-worded items.
But reverse-worded items bring their own problems: some respondents misread them, and they often cluster together in factor analysis as though they formed a separate dimension. We weigh this up in Session 10.
Some respondents complete a questionnaire without reading it — especially long, online, unincentivised ones.
Meade & Craig (2012) showed the proportion can be large enough to change the results of factor analysis and reliability estimates.
This is not only about respondent attitude; it is a consequence of instrument design that overburdens people.
What you can do in your pilot study
Record completion time, watch for flat response patterns (all “3”s), and fix your data-cleaning rules before you look at the results. We will design this together in Session 12.
Everything we discussed today applies equally to well-established, widely used instruments.
The wrong conclusion to draw is “psychological measurement is useless.”
The right conclusion: psychological measurement is useful to the extent that we are honest about its limits.
What distinguishes good work
Not a scale free of assumptions — there is no such thing. Rather a scale whose assumptions are stated, checked as far as they can be checked, and reported honestly.
Psychometrics is one of the most successful branches of psychology . . . [yet] the average psychologist is not aware of the developments in the field.
Borsboom argues that psychometrics and substantive psychology have drifted apart.
The tools for checking these assumptions have existed for a long time, but are rarely used by researchers who are not psychometricians.
One aim of this course is to narrow that gap for you.
Shape of the construct — dimensional or categorical? On what grounds?
Dimensionality — do your items measure one thing? How did you check?
Distribution — what shape is your score distribution actually? Do not assume it; look.
Groups — who is your population, and are there subgroups for whom items may function differently?
Sources of response bias — is your construct prone to social desirability? What did you do about it?
These five points will form part of your final report.
Quantitative — summing items assumes equal weights and equal intervals.
Normal — an inherited assumption, not a finding. A symmetric total-score distribution is largely an effect of aggregation.
Continuous — most psychological constructs appear dimensional, but this still needs stating.
One and the same thing — unidimensionality and local independence determine whether your total score means anything.
The same for every group — an identical item does not always mean the same thing for everyone.
Accurate respondents — people have limited access to themselves, and reasons not to be candid.
We move to the central framework of this course: the latent variable.
Preparation
Read DeVellis & Thorpe (2022), Chapter 2. If you have time, Chapter 3 of Petersen on constructs.
Notes
Borsboom, D. (2006). The attack of the psychometricians. Psychometrika, 71(3), 425–440.
DeVellis, R. F., & Thorpe, C. T. (2022). Scale development: Theory and applications (5th ed.). SAGE.
Haslam, N., Holland, E., & Kuppens, P. (2012). Categories versus dimensions in personality and psychopathology: A quantitative review of taxometric research. Psychological Medicine, 42(5), 903–920.
Heine, S. J., Lehman, D. R., Peng, K., & Greenholtz, J. (2002). What’s wrong with cross-cultural comparisons of subjective Likert scales? The reference-group effect. Journal of Personality and Social Psychology, 82(6), 903–918.
Meade, A. W., & Craig, S. B. (2012). Identifying careless responses in survey data. Psychological Methods, 17(3), 437–455.
Meehl, P. E. (1995). Bootstraps taxometrics: Solving the classification problem in psychopathology. American Psychologist, 50(4), 266–275.
Micceri, T. (1989). The unicorn, the normal curve, and other improbable creatures. Psychological Bulletin, 105(1), 156–166.
Nisbett, R. E., & Wilson, T. D. (1977). Telling more than we can know: Verbal reports on mental processes. Psychological Review, 84(3), 231–259.
Petersen, I. T. Principles of psychological assessment: With applied examples in R. https://isaactpetersen.github.io/Principles-Psychological-Assessment/
Quetelet, A. (1835). Sur l’homme et le développement de ses facultés. Bachelier.