Psychological Scale Construction
2026-08-25
Session 9 produced the construct specification: definition, boundary, estimand, facets, and indicators.
Session 10 produced the blueprint and the item list, with a response format.
Today that item list is judged by other people, before any respondent completes it.
In Session 7, validity was treated as a body of evidence rather than a single number. One source of that evidence is test content.
The CVI is a way of making content evidence computable and reportable.
Content experts — people who know the construct: lecturers, researchers, or practitioners in the field.
Method experts — people who know scale construction and can judge whether an item is well written.
Representatives of the target population — people whose characteristics resemble those of your respondents.
Note
These three give different kinds of feedback. A panel made up entirely of academics tends to miss readability problems that a prospective respondent spots immediately.
At least 3. Below that, agreement means very little.
5 to 10 is the usual and safer range.
The size of the panel determines the pass criterion for each item, so decide the number first.
An odd number is more practical
With an even number, agreement can split exactly down the middle, leaving you with no basis for a decision.
The most commonly forgotten item
Sending the item list without the construct definition. Raters then judge relevance against their own understanding of the construct — and the result cannot be interpreted.
| Rating | Meaning |
|---|---|
| 1 | Not relevant |
| 2 | Somewhat relevant, needs major revision |
| 3 | Quite relevant, needs minor revision |
| 4 | Highly relevant |
Four points, with no midpoint. Raters have to commit: relevant or not.
For the computation, ratings of 3 and 4 count as “relevant”; 1 and 2 count as “not relevant”.
A comment box for each item. A number alone does not tell you what to fix.
One question at the end: “Is there any important aspect of this construct that the items above do not cover?”
Why that last question matters
The CVI only checks whether the items that exist are relevant. It cannot check whether part of the construct is not covered at all.
That is construct underrepresentation from Session 7. The only way to check it through a panel is to ask directly.
We provide an Excel template for computing I-CVI, S-CVI Universal Agreement (UA), S-CVI Average (Ave), and the chance-corrected modified kappa (κ*).
These concepts are explained in the next few slides.
\[\text{I-CVI} = \frac{\text{number of raters giving a 3 or 4}}{\text{total number of raters}}\]
Computed for each item, one value per item.
It ranges from 0 to 1.
Example: of 6 raters, 5 give a 3 or 4. Then I-CVI = 5/6 = 0.83.
| Item | Raters judging it relevant | I-CVI |
|---|---|---|
| A1 | 6 | 1.00 |
| A2 | 6 | 1.00 |
| A3 | 5 | 0.83 |
| A4 | 6 | 1.00 |
| A5 | 4 | 0.67 |
| A6 | 5 | 0.83 |
| A7 | 6 | 1.00 |
| A8 | 3 | 0.50 |
With 3 to 5 raters: I-CVI must be 1.00. Every rater must agree.
With 6 to 8 raters: I-CVI must be at least 0.83.
With 9 or more raters: I-CVI must be at least 0.78.
In the example above, with 6 raters, two items fail: A5 (0.67) and A8 (0.50).
Note
Note that the criterion is stricter when there are fewer raters. The reason is on the next slide.
If raters answered at random, some agreement would still appear by chance.
With few raters, that chance is substantial. With 3 raters, the probability that all three happen to rate an item “relevant” is 12.5%.
So I-CVI on its own can look better than it should.
The probability of chance agreement:
\[p_c = \binom{N}{A} \times 0.5^N\]
Then:
\[\kappa^* = \frac{\text{I-CVI} - p_c}{1 - p_c}\]
where \(N\) is the number of raters and \(A\) the number judging the item relevant.
| Item | \(A\) | I-CVI | \(p_c\) | \(\kappa^*\) | Verdict |
|---|---|---|---|---|---|
| A1 | 6 | 1.00 | 0.016 | 1.00 | excellent |
| A3 | 5 | 0.83 | 0.094 | 0.82 | excellent |
| A5 | 4 | 0.67 | 0.234 | 0.56 | fair |
| A8 | 3 | 0.50 | 0.313 | 0.27 | poor |
Note
The usual benchmarks: \(\kappa^* > 0.74\) excellent, \(0.60\)–\(0.74\) good, \(0.40\)–\(0.59\) fair, below that poor.
Note A8: an I-CVI of 0.50 looks like “half agreed”, but once corrected, the agreement is barely better than guessing.
S-CVI/Ave — the average of all the I-CVIs. It describes average relevance.
S-CVI/UA — the proportion of items on which all raters agreed (universal agreement). Considerably stricter.
For the example above:
\[\text{S-CVI/Ave} = \frac{6.83}{8} = 0.85 \qquad \text{S-CVI/UA} = \frac{4}{8} = 0.50\]
0.85 looks adequate.
0.50 looks poor.
Both come from exactly the same table.
Polit & Beck (2006) wrote a paper specifically about this confusion: many reports cite “the CVI” without saying which one.
What your group must do
State explicitly which computation you used. Writing “CVI = 0.85” without qualification leaves the number uninterpretable.
The one usually reported is S-CVI/Ave, with a benchmark of at least 0.90.
Revise first, do not delete straight away. Read the raters’ comments. The problem is often the wording rather than the idea.
Delete if the item really measures something else, or if the facet is already better represented by another item.
Keep it with a stated reason if it is conceptually essential and has no replacement. Write that reason into the report.
Warning
Recall the attenuation paradox from Session 7. Dropping items purely to produce tidier numbers can narrow content coverage, which does more harm than good.
If A5 and A8 are revised and re-rated, or removed:
The six remaining items have I-CVIs of 1.00 — 1.00 — 0.83 — 1.00 — 0.83 — 1.00
\(\text{S-CVI/Ave} = 5.67 / 6 = \mathbf{0.94}\)
That now passes the 0.90 benchmark.
Note
Report both: the result before revision and after, together with the decision taken on each item. What is assessed is the process, not only the final number.
Completeness. The CVI judges the items that exist, not the items that should exist. For that you have to ask the panel directly.
Clarity for respondents. Expert raters read items differently from prospective respondents. Clarity is checked through cognitive interviewing, as covered in Session 7.
Whether the construct itself makes sense. If your construct definition is flawed, the panel will rate relevance against that flawed definition.
Tip
The CVI is one source of evidence, not a verdict. It occupies one row of the validity evidence table from Session 7 — not the whole table.
30 minutes. Each group prepares:
25 minutes. Groups swap and act as raters for one another.
Each group member rates 10 items belonging to another group using the 4-point scale.
Collect the ratings and compute: I-CVI for each item, \(\kappa^*\), and both S-CVI/Ave and S-CVI/UA.
Identify which items fail, and write one revision suggestion for each.
Tip
This exercise uses classmates as raters, so the result is not genuine validity evidence. The point is to make sure the computation is correct before the real panel data arrives.
Preparation
Bring today’s revised item list. If ratings from your actual panel have come back, bring those too.
Notes
AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association.
Davis, L. L. (1992). Instrument review: Getting the most from a panel of experts. Applied Nursing Research, 5(4), 194–197.
Haynes, S. N., Richard, D. C. S., & Kubany, E. S. (1995). Content validity in psychological assessment: A functional approach to concepts and methods. Psychological Assessment, 7(3), 238–247.
Lynn, M. R. (1986). Determination and quantification of content validity. Nursing Research, 35(6), 382–385.
Polit, D. F., & Beck, C. T. (2006). The content validity index: Are you sure you know what’s being reported? Research in Nursing & Health, 29(5), 489–497.
Polit, D. F., Beck, C. T., & Owen, S. V. (2007). Is the CVI an acceptable indicator of content validity? Appraisal and recommendations. Research in Nursing & Health, 30(4), 459–467.
Rubio, D. M., Berg-Weger, M., Tebb, S. S., Lee, E. S., & Rauch, S. (2003). Objectifying content validity: Conducting a content validity study in social work research. Social Work Research, 27(2), 94–104.
Waltz, C. F., & Bausell, R. B. (1981). Nursing research: Design, statistics, and computer analysis. F. A. Davis.