Item-Writing Workshop: Blueprint and Choosing a Response Format

Psychological Scale Construction

Rizqy Amelia Zein & Dian Kartika Amelia Arbi

Department of Psychology, Universitas Airlangga

2026-08-25

Outline

  • Building a blueprint: how many items per facet, and why that many
  • Item-writing principles, with faulty examples and their repairs
  • Reverse-worded items: what the evidence says, and when to use them
  • Settling the response format for the whole scale
  • Assembling the questionnaire: introduction, instructions, layout
  • Workshop and peer review between groups

From construct specification to items

What you bring today

  • The construct specification from Session 9: definition, boundary, estimand, dimensions, facets, and indicators.

  • Today those indicators become items, and the items become a questionnaire.

Warning

If your group does not yet have a construct specification, do that first today. Writing items without a construct definition only produces work that has to be redone.

The blueprint

What a blueprint does

  • It links facets of the construct to the number of items representing each one.

  • Its job is to keep content coverage aligned with your definition, rather than with whichever facet is easiest to write items for.

Note

Without a blueprint, the same thing happens almost every time: the facet that is easiest to picture gets ten items, while the difficult one gets two.

An example blueprint

Construct: social anxiety in undergraduates.

Facet Weight Indicators Items written Final target
Fear of negative evaluation 40% 3 12 6
Avoidance of social situations 35% 3 10 5
Physiological symptoms 25% 2 8 4
Total 100% 8 30 15
  • Weights come from your construct definition, not from preference.

  • Note the last two columns: you write considerably more than you will use.

Why write a surplus

  • Some items will be dropped at expert review (Session 11), others at item analysis (Session 13).

  • If you write exactly 15 items for a 15-item target, you have no room to discard anything.

  • A common rule of thumb: write two to three times your target number.

How many items in the end

Recall from Session 6: reliability is driven by the number of items and the mean inter-item correlation. For a scale with moderate inter-item correlations, 10–20 items is usually enough.

Adding items indefinitely gives shrinking returns and increases respondent burden.

Spread the items out

  • Recall the discussion of item difficulty in Session 3: items should be spread across the range of the construct.

  • For each facet, write some items that are easy to endorse (indicating a low level) and some that are hard to endorse (indicating a high level).

  • If every item describes a severe level, respondents at moderate levels will not be distinguished from one another.

Tip

This is also what prevents the ceiling and floor effects that will show up immediately in the score distribution in Session 13.

Item-writing principles

One item, one idea

  • A double-barrelled item contains two ideas at once, so a respondent who agrees with only one half has no correct answer available.

  • The words “and”, “because”, and “if” are frequent signs of a double-barrelled item.

Problematic Repair
“I feel nervous and sweaty when presenting.” Split into two items.
“I avoid campus events because I fear being judged.” “I avoid campus events.” + a separate item about the reason.

Short and plain

  • A long item forces the respondent to hold the start of the sentence in mind while reading the end.

  • Avoid double negatives. “I do not mind not being invited to discussions” is hard for anyone to process.

  • Avoid ambiguous pronoun references and misplaced modifiers.

Note

DeVellis suggests a reading level around grades 5–7 for instruments used with the general population. With undergraduate respondents you have a little more room, but short sentences are still better.

Concrete behaviour, not its causes

  • Recall Session 2: Nisbett & Wilson showed that people often have no access to the real reasons behind their behaviour.

  • Items asking what happens are generally more dependable than items asking why.

Problematic Repair
“I feel anxious because I fear others’ judgement.” “I worry that others judge me negatively.”
“I am introverted, so I avoid crowds.” “I avoid crowded events.”

The time frame comes from the estimand

  • Your estimand sentence from Session 9 already names a time frame. That time frame has to appear in the items.

  • “In the past month, I have …” measures something different from “I usually …”.

  • Make sure the time frame is the same across all items in a scale.

Warning

Mixing time frames within one scale adds a source of variation you did not intend, and it usually appears as an extra factor in the analysis in Session 13.

How strongly to word the stem

  • From Session 4: for a Likert format, statements should be fairly strong but not extreme, because degree is already carried by the response options.

  • A statement that is too mild will be endorsed by nearly everyone, and stops distinguishing anyone.

  • A statement that is too extreme will be rejected by nearly everyone, with the same result.

Avoid loaded words

  • Words such as “reasonable”, “should”, “excessive”, or “fail” steer the answer.

  • Items should also avoid making respondents feel judged, particularly for constructs vulnerable to social desirability (Session 2).

Problematic Repair
“I worry excessively about others’ opinions.” “I often think about what others think of me.”
“I fail to control my anxiety.” “I find it hard to calm myself when anxious.”

Reverse-worded items

Why they were introduced

  • The idea comes from Likert (1932) himself: to reduce the tendency to agree with anything, half the items were worded in one direction and half in the other.

  • If a respondent is simply agreeing, the pattern becomes visible: they endorse both the forward and the reversed items.

Note

So far the idea is sound. The difficulty is in the execution.

What the evidence shows

  • Scales with all items worded in one direction produce more accurate responses than mixed-worded scales.

  • Scores on forward and reversed items are not symmetric: strongly disagreeing with a positive statement is not the same as strongly agreeing with its negative counterpart.

  • Reversed items often form their own factor in factor analysis. What appears there is the wording, not the construct.

And respondents stop noticing

  • Respondent performance declines roughly 12 minutes after starting a survey.

  • After that, many stop registering the word “not” in an item — even when it is bolded, underlined, or capitalised.

If you use them anyway

  • Place reversed items early in the questionnaire.
  • Make sure they are genuine polar opposites, not just a “not” added to the sentence.
  • Check their effect during analysis: see whether they cluster on their own.

The “not” trap

  • “I feel anxious when presenting” and “I do not feel anxious when presenting” are not two ends of one continuum.

  • Not feeling anxious is not the same as feeling calm. There is a great deal in between.

  • If you want a reversed item, reverse the content, not the grammar: “I feel calm when presenting.”

For some constructs, one-directional items are more appropriate

For constructs that are themselves negative — anxiety, depression — negatively worded statements are the natural form. Forcing reversals here only makes the sentences convoluted.

Response format

The decisions you already know

From Session 4, settle these now for the whole scale:

  • How many options. The 5–7 range is sensible for most purposes.
  • Midpoint or not. Use one if a middle position is genuinely meaningful for your construct.
  • Labels. Label every option, not only the endpoints.
  • Anchors. Agreement, frequency, or item-specific anchors.

On the midpoint

Chyung et al. (2017) suggest reframing the question: not “should there be a midpoint” but “when is a midpoint appropriate”. A midpoint is appropriate when it is a genuine middle position, not somewhere to hedge.

If what you expect is genuine uncertainty, provide a separate “don’t know” option rather than folding it into the midpoint.

Consistency and order

  • Use one response format across all items in a scale. Switching midway increases burden and complicates summing.

  • Present the response options in ascending order, from lowest to highest.

  • Keep the direction consistent from beginning to end.

Assembling the questionnaire

Its parts

  • Title — reflects the content, concise, and not off-putting.

  • Introductory statement — brief purpose, confidentiality and consent, and the approximate time required.

  • Instructions — complete, unambiguous, including how to submit responses.

  • Items — grouped and numbered.

  • Closing statement — thanks, and next steps if any.

This is not a formality

  • Pershing & Pershing (2001) examined 50 training-evaluation forms used at a well-regarded medical school.

  • 72% had no introductory statement at all. 78% had no closing statement. 30% had no instructions, and another 54% had minimal ones.

  • Only 8% were professional in appearance.

Warning

A cluttered questionnaire reduces respondent engagement, and that feeds directly into the reliability and validity of your scores.

Layout

  • Leave enough white space; do not crowd the page.
  • Number items and use section headings so respondents can see their progress.
  • Insert a break every 4–6 items, or shade alternate rows.
  • For online administration, check that it reads well on a phone screen.

Demographic questions

Put them at the end, unless they are needed to screen respondents at the start. Demographic questions at the beginning make some respondents feel identified before they know what the questionnaire is about.

Workshop

Part 1: writing items

  • 15 minutes — Complete your group’s blueprint: facets, weights, indicators, and the number of items to write.

  • 40 minutes — Write the items.

  • 10 minutes — Settle the response format and draft the instructions.

Part 2: peer review

Swap item lists with another group. 25 minutes.

For each item, check:

  1. Does it contain more than one idea?
  2. Are there double negatives or unclear pronoun references?
  3. Does it ask about behaviour or about reasons?
  4. Is the time frame the same as in the other items?
  5. Are there loaded words steering the answer?
  6. Could people at different levels of the construct answer this identically? If so, it distinguishes nothing.

Tip

Return the list with written notes, not a general verdict. “Item 7 is double-barrelled: nervous and sweaty” is far more useful than “the items are unclear”.

What you must produce

  • A blueprint table, with weights and item counts.
  • An item list revised after peer review, each item tagged with its facet.
  • The response format and its labels.
  • A draft questionnaire with introduction and instructions.

Warning

Bring all of it to Session 11. The expert panel cannot judge item relevance without reading your construct definition and blueprint.

Any questions❓

Notes

References

Bikos, L. H. ReCentering psych stats: Psychometrics. https://lhbikos.github.io/ReC_Psychometrics/

Chyung, S. Y., Barkin, J. R., & Shamsy, J. A. (2018). Evidence-based survey design: The use of negatively worded items in surveys. Performance Improvement, 57(3), 16–25.

Chyung, S. Y., Kennedy, M., & Campbell, I. (2018). Evidence-based survey design: The use of ascending or descending order of Likert-type response options. Performance Improvement, 57(9), 9–16.

Chyung, S. Y., Roberts, K., Swanson, I., & Hankinson, A. (2017). Evidence-based survey design: The use of a midpoint on the Likert scale. Performance Improvement, 56(10), 15–23.

DeVellis, R. F., & Thorpe, C. T. (2022). Scale development: Theory and applications (5th ed.). SAGE.

Krosnick, J. A., & Presser, S. (2010). Question and questionnaire design. In P. V. Marsden & J. D. Wright (Eds.), Handbook of survey research (2nd ed., pp. 263–313). Emerald.

Likert, R. (1932). A technique for the measurement of attitudes. Archives of Psychology, 22(140), 1–55.

Pershing, J. A., & Pershing, J. L. (2001). Ineffective reaction evaluation. Performance Improvement Quarterly, 14(1), 73–90.

Weijters, B., Baumgartner, H., & Schillewaert, N. (2013). Reversed item bias: An integrative model. Psychological Methods, 18(3), 320–334.