Guo Wenjing, Wind Stefanie A
The University of Alabama, Tuscaloosa, AL, USA.
Appl Psychol Meas. 2023 Mar;47(2):91-105. doi: 10.1177/01466216231151705. Epub 2023 Jan 12.
In standalone performance assessments, researchers have explored the influence of different rating designs on the sensitivity of latent trait model indicators to different rater effects as well as the impacts of different rating designs on student achievement estimates. However, the literature provides little guidance on the degree to which different rating designs might affect rater classification accuracy (severe/lenient) and rater measurement precision in both standalone performance assessments and mixed-format assessments. Using results from an analysis of National Assessment of Educational Progress (NAEP) data, we conducted simulation studies to systematically explore the impacts of different rating designs on rater measurement precision and rater classification accuracy (severe/lenient) in mixed-format assessments. The results suggest that the complete rating design produced the highest rater classification accuracy and greatest rater measurement precision, followed by the multiple-choice (MC) + spiral link design and the MC link design. Considering that complete rating designs are not practical in most testing situations, the MC + spiral link design may be a useful choice because it balances cost and performance. We consider the implications of our findings for research and practice.
在独立表现评估中,研究人员探讨了不同评分设计对潜在特质模型指标对不同评分者效应的敏感性的影响,以及不同评分设计对学生成绩估计的影响。然而,文献对于不同评分设计在独立表现评估和混合格式评估中可能影响评分者分类准确性(严格/宽松)和评分者测量精度的程度几乎没有提供指导。利用对国家教育进展评估(NAEP)数据的分析结果,我们进行了模拟研究,以系统地探讨不同评分设计在混合格式评估中对评分者测量精度和评分者分类准确性(严格/宽松)的影响。结果表明,完整评分设计产生了最高的评分者分类准确性和最大的评分者测量精度,其次是多项选择题(MC)+螺旋链接设计和MC链接设计。鉴于完整评分设计在大多数测试情况下并不实用,MC+螺旋链接设计可能是一个有用的选择,因为它平衡了成本和性能。我们考虑了我们的研究结果对研究和实践的影响。