State of the psychometric methods: patient-reported outcome measure development and refinement using item response theory

被引：56

作者：

Stover, Angela M. ^{[1
,2
]}

McLeod, Lori D. ^{[3
]}

Langer, Michelle M. ^{[2
,4
,5
]}

Chen, Wen-Hung ^{[3
]}

Reeve, Bryce B. ^{[1
,6
]}

机构：

[1] Univ N Carolina, Dept Hlth Policy & Management, 1101-G McGavran Greenberg Hall CB 7411, Chapel Hill, NC 27599 USA

[2] Univ N Carolina, Lineberger Comprehens Canc Ctr, Sch Med, 101 Manning Dr, Chapel Hill, NC 27599 USA

[3] RTI Hlth Solut, 3040 Cornwallis Rd, Res Triangle Pk, NC 27709 USA

[4] Northwestern Univ, Med Social Sci, 625 N Michigan Ave Suite 2700, Chicago, IL 60611 USA

[5] Northwestern Univ, Feinberg Sch Med, 625 N Michigan Ave Suite 2700, Chicago, IL 60611 USA

[6] Duke Univ, Sch Med, Dept Populat Hlth Sci & Pediat, Ctr Hlth Measurement, 2200 West Main St,Suite 720A, Durham, NC 27707 USA

来源：

JOURNAL OF PATIENT-REPORTED OUTCOMES | 2019年 / 3卷 / 01期

关键词：

Item response theory; Scale construction; Scale evaluation; Measurement; PROMIS (R); GOODNESS-OF-FIT; INFORMATION-SYSTEM PROMIS(R); INSTRUMENT DEVELOPMENT; LIMITED-INFORMATION; LATENT ABILITY; IRT MODEL; DEPRESSION; IMPACT; ORGANIZATION; VALIDATION;

D O I：

10.1186/s41687-019-0130-5

中图分类号：

R19 [保健组织与事业（卫生事业管理）];

学科分类号：

摘要：

Background: This paper is part of a series comparing different psychometric approaches to evaluate patient-reported outcome (PRO) measures using the same items and dataset. We provide an overview and example application to demonstrate 1) using item response theory (IRT) to identify poor and well performing items; 2) testing if items perform differently based on demographic characteristics (differential item functioning, DIF); and 3) balancing IRT and content validity considerations to select items for short forms. Methods: Model fit, local dependence, and DIF were examined for 51 items initially considered for the Patient-Reported Outcomes Measurement Information System (R) (PROMIS (R)) Depression item bank. Samejima's graded response model was used to examine how well each item measured severity levels of depression and how well it distinguished between individuals with high and low levels of depression. Two short forms were constructed based on psychometric properties and consensus discussions with instrument developers, including psychometricians and content experts. Calibrations presented here are for didactic purposes and are not intended to replace official PROMIS parameters or to be used for research. Results: Of the 51 depression items, 14 exhibited local dependence, 3 exhibited DIF for gender, and 9 exhibited misfit, and these items were removed from consideration for short forms. Short form 1 prioritized content, and thus items were chosen to meet DSM-V criteria rather than being discarded for lower discrimination parameters. Short form 2 prioritized well performing items, and thus fewer DSM-V criteria were satisfied. Short forms 1-2 performed similarly for model fit statistics, but short form 2 provided greater item precision. Conclusions: IRT is a family of flexible models providing item- and scale-level information, making it a powerful tool for scale construction and refinement. Strengths of IRT models include placing respondents and items on the same metric, testing DIF across demographic or clinical subgroups, and facilitating creation of targeted short forms. Limitations include large sample sizes to obtain stable item parameters, and necessary familiarity with measurement methods to interpret results. Combining psychometric data with stakeholder input (including people with lived experiences of the health condition and clinicians) is highly recommended for scale development and evaluation.

引用

页数：16

共 50 条

[1] State of the psychometric methods: patient-reported outcome measure development and refinement using item response theory
Angela M. Stover
Lori D. McLeod
Michelle M. Langer
Wen-Hung Chen
Bryce B. Reeve
Journal of Patient-Reported Outcomes, 3
[2] Using item response theory to develop and refine patient-reported outcome measures
Nguyen, Tam H.
Lee, Christopher S.
Kim, Miyong T.
EUROPEAN JOURNAL OF CARDIOVASCULAR NURSING, 2022, 21 (05) : 509 - 515
[3] An Introduction to Item Response Theory for Patient-Reported Outcome Measurement
Nguyen, Tam H.
Han, Hae-Ra
Kim, Miyong T.
Chan, Kitty S.
PATIENT-PATIENT CENTERED OUTCOMES RESEARCH, 2014, 7 (01) : 23 - 35
[4] An Introduction to Item Response Theory for Patient-Reported Outcome Measurement
Tam H. Nguyen
Hae-Ra Han
Miyong T. Kim
Kitty S. Chan
The Patient - Patient-Centered Outcomes Research, 2014, 7 : 23 - 35
[5] Reliability, Validity, and Efficiency of an Item Response Theory-Based Balance Confidence Patient-Reported Outcome Measure
Deutscher, Daniel
Kallen, Michael A.
Werneke, Mark W.
Mioduski, Jerome E.
Hayes, Deanna
PHYSICAL THERAPY, 2023, 103 (07):
[6] Development of a Patient-Reported Outcome Measure for Mohs Reconstruction
Kavanagh, Kaitlin J.
Christophel, J. Jared
FACIAL PLASTIC SURGERY & AESTHETIC MEDICINE, 2020, 22 (04) : 274 - 280
[7] Development and validation of a patient-reported outcome measure for stroke patients
Yanhong Luo
Jie Yang
Yanbo Zhang
Health and Quality of Life Outcomes, 13
[8] Item development for a patient-reported measure of compassionate healthcare in action
Chatburn, Eleanor
Marks, Elizabeth
Maddox, Lucy
HEALTH EXPECTATIONS, 2024, 27 (01)
[9] Development and validation of a patient-reported outcome measure for stroke patients
Luo, Yanhong
Yang, Jie
Zhang, Yanbo
HEALTH AND QUALITY OF LIFE OUTCOMES, 2015, 13
[10] Using Classical Test Theory, Item Response Theory, and Rasch Measurement Theory to Evaluate Patient-Reported Outcome Measures: A Comparison of Worked Examples
Petrillo, Jennifer
Cano, Stefan J.
McLeod, Lori D.
Coon, Cheryl D.
VALUE IN HEALTH, 2015, 18 (01) : 25 - 34

← 1 2 3 4 5 →