Writing Items to Measure Higher-Order Thinking Skills (HOTS)

Cognitive Test Development

2026-09-16

Agenda

  1. From LOTS to HOTS
  2. Definition and characteristics of HOTS items
  3. Writing items at the Analyzing level
  4. Writing items at the Evaluating level
  5. Writing items at the Creating level
  6. HOTS in achievement tests vs. aptitude tests
  7. Common mistakes in writing HOTS items
  8. Activity: reviewing a HOTS item

LOTS vs. HOTS

Recall from Meeting 3: the six levels of cognitive process

Meeting 3 covered the revised Bloom’s taxonomy (Anderson & Krathwohl, 2001) with six cognitive process levels, arranged in order from simplest:

Remembering → Understanding → Applying → Analyzing → Evaluating → Creating

LOTS vs. HOTS

The first three levels are usually called lower-order thinking skills (LOTS); the last three are called higher-order thinking skills (HOTS). Today we focus on the Analyzing, Evaluating, and Creating levels.

Why LOTS-level items alone aren’t enough

Items that measure only the Remembering level require test-takers to recall information, without needing to show that they actually understand, interpret, or can apply that information to solve a problem they have not encountered in exactly the same context in which the information was taught.

The higher the cognitive level being measured, the more complex the thinking demanded of test-takers — so it’s more than just a higher difficulty level for the item.

Definition and characteristics of HOTS items

What is HOTS

Definition

HOTS covers thinking processes that go beyond remembering and understanding: analyzing, evaluating (making a judgement), and creating — including critical thinking, creative thinking, and problem-solving.

Two characteristics of a good HOTS item

  1. There is an introductory stimulus — text, graph, table, case — that gives test-takers something to think through.
  2. A context or problem that the test-taker has not encountered in exactly the same form before, so the item cannot be answered simply by relying on memory.

A common mistake

A long, complicated-looking stimulus does not automatically make an item measure HOTS. If the answer can be found simply by matching a sentence in the stimulus, without needing to process that information, the cognitive process actually being demanded is still Remembering.

Writing items at the Analyzing level

Three types of analysis process

  • Analyzing elements — recognizing assumptions that are not explicitly stated, distinguishing facts from opinions, distinguishing conclusions from the facts that support them.
  • Analyzing relationships — identifying relationships between ideas, recognizing cause-and-effect relationships, distinguishing relevant arguments from irrelevant ones.
  • Analyzing organizational principles — recognizing the form, pattern, or structure implicit in a text or communication.

Example item: analyzing relationships

Stimulus

“A report found that ice cream sales and the number of drowning cases at public swimming pools both increased over the year. The report’s author concluded that buying ice cream increases the risk of drowning.”

“The main flaw in this conclusion is…”

  1. The ice cream sales data are not accurate enough B) Both are likely influenced by the summer season C) There are too few drowning cases to analyze D) Public swimming pools are not representative

B — correlation between two variables does not prove that one causes the other; both can be explained by a third variable (the summer season). Answering this requires analyzing a cause-and-effect relationship.

Writing items at the Evaluating level

Two types of evaluation criteria

  • Internal criteria — the accuracy, consistency, and logical sequencing of the work or argument itself.
  • External criteria — standards outside the work, such as efficiency, fitness for a particular purpose, or comparison with other alternatives.

Example item: evaluating

Stimulus

“The following two study designs both test whether training X improves score Y.”

Design 1: Measures score Y in the same group, before and after training, with no comparison group.

Design 2: Measures score Y in a group that received the training and a group that did not, both measured at the same time.

“Based on the criterion of ability to control for alternative explanations, which design is stronger?”

  1. Design 1 B) Design 2 C) Equally strong D) Cannot be determined

Answer key

B — a design with a comparison group is better able to control for alternative explanations such as time effects or maturation, which cannot be ruled out in Design 1.

Writing items at the Creating level

Three forms of product demanded

Bloom et al.’s (1956) original taxonomy called this level Synthesis; the revised taxonomy folded it into Creating (Meeting 3). Its three original sub-categories are still useful as a guide for writing items:

  • Producing a unique communication — a short story, an essay, a composition.
  • Devising a plan or a set of proposed operations — an experimental proposal, an intervention design.
  • Deriving a set of abstract relations — formulating a hypothesis, constructing a new conceptual scheme.

The limits of multiple choice for this level

Multiple-choice items are poorly suited to measuring Creating in its pure form, because this level emphasizes the test-taker’s ability to produce original ideas. In practice, this level is more often measured through essays or project-based assessment than through objective tests.

For tests that must remain standardized with objectively scorable items (such as aptitude tests), a more realistic approach is to ask test-takers to recognize the best synthesis among several options provided.

HOTS in achievement tests vs. aptitude tests

Different paths to the same level

  • Aptitude tests — items are designed to measure reasoning processes, not content mastery (Meetings 5–7). Because of this, a good aptitude test item is almost automatically at the Analyzing or Evaluating level; for example, the induction and general sequential reasoning items discussed in Meeting 7.
  • Achievement tests — items are derived from Specific Instructional Objectives (TIK) tied to particular content (Meeting 4). Moving from LOTS to HOTS here has to be deliberate: a TIK item that merely tests rote recall needs to be rewritten as one that demands case analysis or application to a new situation.

Relevance for your final project

The cognitive test you develop for your final project (Meeting 14) will have stronger content validity if some of its items are deliberately designed at the HOTS level, in line with the proportions in your blueprint (Meetings 2 and 3).

Common mistakes in writing HOTS items

Two mistakes that keep happening

  • Pseudo-HOTS — the stimulus is made long and looks complex, but the answer can be found directly in the stimulus without any need to process new information. The cognitive process actually demanded is still Remembering.
  • Irrelevant language difficulty — the stimulus is written in language that is too complicated, so the item ends up testing test-takers’ literacy skills more than the intended cognitive process.

The principles for writing stems and options from Meeting 7 — positive wording, one core idea, options of comparable length — still apply in full to HOTS items. A higher cognitive level is no excuse for ignoring these basic principles.

Activity: reviewing a HOTS item

Item to review

“Read the following paragraph: ‘Cultural bias in a cognitive test occurs when a test produces systematic measurement error against a particular cultural group.’ Based on the paragraph above, the definition of cultural bias is…”

  1. Random measurement error affecting all test-takers B) Systematic measurement error against a particular cultural group C) A difference in mean scores between cultural groups D) A mismatch between the test norms and the population

Discuss with a classmate: does this item really measure the Analyzing level?

Answer key

No — answer (B) can be found by directly matching a sentence in the stimulus, without needing to process or connect any information. This is an example of pseudo-HOTS: the stimulus looks as though it requires analytical skills, but the cognitive process actually demanded is still Remembering.

Summary

  1. HOTS covers the three highest levels of the revised Bloom’s taxonomy: Analyzing, Evaluating, and Creating.
  2. A good HOTS item has a stimulus and context that the test-taker has not encountered in exactly the same form before.
  3. The sub-categories of Bloom et al.’s (1956) original taxonomy — analyzing elements/relationships/organizational principles, internal/external evaluation criteria, the three forms of synthesis products — remain useful as a practical guide for writing items.
  4. Creating in its pure form is hard to measure with multiple choice; aptitude tests usually approximate it through pattern-recognition items.
  5. Aptitude test items tend to sit at the HOTS level almost automatically; achievement test items need to be deliberately raised from the existing TIK.
  6. Watch out for pseudo-HOTS — a complicated stimulus whose answer can still be found through simple recall.

For Meeting 10

Next up: structured practice

Meeting 10 begins a series of structured, guided practice sessions on developing a cognitive test, starting with item specifications and building the blueprint for your group’s final project.

Thank you!😊

Any questions?

These slides were prepared using and Quarto with a template from UNAIR Theme.

References

Anderson, L. W., & Krathwohl, D. R. (Eds.). (2001). A taxonomy for learning, teaching, and assessing: A revision of Bloom’s taxonomy of educational objectives. Longman.

Bloom, B. S., Engelhart, M. D., Furst, E. J., Hill, W. H., & Krathwohl, D. R. (1956). Taxonomy of educational objectives: The classification of educational goals. Handbook I, Cognitive domain. David McKay Company.

Brookhart, S. M. (2010). How to assess higher-order thinking skills in your classroom. ASCD.

Haladyna, T. M., Downing, S. M., & Rodriguez, M. C. (2002). A review of multiple-choice item-writing guidelines for classroom assessment. Applied Measurement in Education, 15(3), 309–334.

Krathwohl, D. R. (2002). A revision of Bloom’s taxonomy: An overview. Theory Into Practice, 41(4), 212–218. https://www.jstor.org/stable/1477405?seq=1