Evaluating Equating Transformations in IRT Observed-Score and Kernel Equating Methods

被引：7

作者：

Leoncio, Waldir ^{[1
,2
]}

Wiberg, Marie ^{[3
]}

Battauz, Michela ^{[4
]}

机构：

[1] Univ Padua, Dept Stat Sci, Padua, Italy

[2] Univ Oslo, Ctr Educ Measurement, Ctr Biostat & Epidemiol, Oslo, Norway

[3] Umea Univ, Umea Sch Business Econ & Stat, Dept Stat, Umea, Sweden

[4] Univ Udine, Dept Econ & Stat, Udine, Italy

来源：

APPLIED PSYCHOLOGICAL MEASUREMENT | 2023年 / 47卷 / 02期

关键词：

equating; item response theory; classical test theory; psychometrics; simulation; statistics; TRUE-SCORE; ITEM; MODELS;

D O I：

10.1177/01466216221124087

中图分类号：

O1 [数学]; C [社会科学总论];

学科分类号：

03 ; 0303 ; 0701 ; 070101 ;

摘要：

Test equating is a statistical procedure to ensure that scores from different test forms can be used interchangeably. There are several methodologies available to perform equating, some of which are based on the Classical Test Theory (CTT) framework and others are based on the Item Response Theory (IRT) framework. This article compares equating transformations originated from three different frameworks, namely IRT Observed-Score Equating (IRTOSE), Kernel Equating (KE), and IRT Kernel Equating (IRTKE). The comparisons were made under different data-generating scenarios, which include the development of a novel data-generation procedure that allows the simulation of test data without relying on IRT parameters while still providing control over some test score properties such as distribution skewness and item difficulty. Our results suggest that IRT methods tend to provide better results than KE even when the data are not generated from IRT processes. KE might be able to provide satisfactory results if a proper pre-smoothing solution can be found, while also being much faster than IRT methods. For daily applications, we recommend observing the sensibility of the results to the equating method, minding the importance of good model fit and meeting the assumptions of the framework.

引用

页码：123 / 140

页数：18

共 32 条

[1]

ANDERSSON B, 2013, J STAT SOFTW, V55, P1, DOI DOI 10.18637/JSS.V055.I06

[2] ITEM RESPONSE THEORY OBSERVED-SCORE KERNEL EQUATING [J].

Andersson, Bjorn ;

Wiberg, Marie .

PSYCHOMETRIKA, 2017, 82 (01) :48-66

[3] equateIRT: An R Package for IRT Test Equating [J].

Battauz, Michela .

JOURNAL OF STATISTICAL SOFTWARE, 2015, 68 (07) :1-22

[4]

Braun H.I., 1982, TEST EQUATING, P9

[5] PROBLEMS RELATED TO THE USE OF CONVENTIONAL AND ITEM RESPONSE THEORY EQUATING METHODS IN LESS THAN OPTIMAL CIRCUMSTANCES [J].

COOK, LL ;

PETERSEN, NS .

APPLIED PSYCHOLOGICAL MEASUREMENT, 1987, 11 (03) :225-244

[6]

Dorans N.J., 1994, Technical issues related to the introduction of the new SAT and PSAT/NMSQT, (RM-94-10), P91

[7]

GONZALEZ J, 2017, METHOD EDUC MEAS, P1

[8] A Note on the Poisson's Binomial Distribution in Item Response Theory [J].

Gonzalez, Jorge ;

Wiberg, Marie ;

von Davier, Alina A. .

APPLIED PSYCHOLOGICAL MEASUREMENT, 2016, 40 (04) :302-310

[9] AN INVESTIGATION OF CLASSIFICATION CONSISTENCY INDEXES ESTIMATED UNDER ALTERNATIVE STRONG TRUE SCORE MODELS [J].

HANSON, BA ;

BRENNAN, RL .

JOURNAL OF EDUCATIONAL MEASUREMENT, 1990, 27 (04) :345-359

[10]

Harris D.J., 1993, APPL MEAS EDUC, V6, P195, DOI [DOI 10.1207/S15324818AME0603_3, 10.1207/s15324818ame0603_3]

← 1 2 3 4 →