Towards more reliable survey items: An item-level stability protocol

Sara Dolnicar et al.

Annals of Tourism Research2026https://doi.org/10.1016/j.annals.2026.104148article
AJG 4ABDC A*
Weight
0.50

What the paper says

Tourism researchers rely heavily on self-report data. The validity of their insights depends on reliability of measures, yet conventional test–retest reliability (1) overlooks item-level stability by relying on aggregate scale coefficients that can mask unstable questions; and (2) ignores how response options affect stability. We introduce the Item-Level Stability Protocol, which evaluates item-level test–retest reliability across answer options. We demonstrate its value using two-wave longitudinal survey data ( N = 3193). Results show test–retest reliabilities can be substantially increased; an internal validation experiment achieves on average 0.10 higher Pearson correlation values. The new protocol is paradigm-agnostic, complements existing psychometric methods, and is simple to implement. It helps tourism scholars and practitioners identify optimal survey questions and response options for increased reliability. • Tourism research relies on high quality self-report data for insights. • Current test–retest reliability methods have two key limitations. • A new protocol assesses item-level reliability and answer option effects. • It can be implemented easily with a small longitudinal sample. • It helps optimise items and answer options to improve reliability.

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1016/j.annals.2026.104148

Or copy a formatted citation

@article{sara2026,
  title        = {{Towards more reliable survey items: An item-level stability protocol}},
  author       = {Sara Dolnicar et al.},
  journal      = {Annals of Tourism Research},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1016/j.annals.2026.104148},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Towards more reliable survey items: An item-level stability protocol

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.