评分设计对大规模混合格式评估中评分者分类准确性和评分者测量精度的影响。

The Effects of Rating Designs on Rater Classification Accuracy and Rater Measurement Precision in Large-Scale Mixed-Format Assessments.

作者信息

Guo Wenjing, Wind Stefanie A

机构信息

The University of Alabama, Tuscaloosa, AL, USA.

出版信息

Appl Psychol Meas. 2023 Mar;47(2):91-105. doi: 10.1177/01466216231151705. Epub 2023 Jan 12.

DOI:10.1177/01466216231151705

PMID:36875294

原文链接:https://pmc.ncbi.nlm.nih.gov/articles/PMC9979195/

Abstract

In standalone performance assessments, researchers have explored the influence of different rating designs on the sensitivity of latent trait model indicators to different rater effects as well as the impacts of different rating designs on student achievement estimates. However, the literature provides little guidance on the degree to which different rating designs might affect rater classification accuracy (severe/lenient) and rater measurement precision in both standalone performance assessments and mixed-format assessments. Using results from an analysis of National Assessment of Educational Progress (NAEP) data, we conducted simulation studies to systematically explore the impacts of different rating designs on rater measurement precision and rater classification accuracy (severe/lenient) in mixed-format assessments. The results suggest that the complete rating design produced the highest rater classification accuracy and greatest rater measurement precision, followed by the multiple-choice (MC) + spiral link design and the MC link design. Considering that complete rating designs are not practical in most testing situations, the MC + spiral link design may be a useful choice because it balances cost and performance. We consider the implications of our findings for research and practice.

摘要

在独立表现评估中，研究人员探讨了不同评分设计对潜在特质模型指标对不同评分者效应的敏感性的影响，以及不同评分设计对学生成绩估计的影响。然而，文献对于不同评分设计在独立表现评估和混合格式评估中可能影响评分者分类准确性（严格/宽松）和评分者测量精度的程度几乎没有提供指导。利用对国家教育进展评估（NAEP）数据的分析结果，我们进行了模拟研究，以系统地探讨不同评分设计在混合格式评估中对评分者测量精度和评分者分类准确性（严格/宽松）的影响。结果表明，完整评分设计产生了最高的评分者分类准确性和最大的评分者测量精度，其次是多项选择题（MC）+螺旋链接设计和MC链接设计。鉴于完整评分设计在大多数测试情况下并不实用，MC+螺旋链接设计可能是一个有用的选择，因为它平衡了成本和性能。我们考虑了我们的研究结果对研究和实践的影响。

相似文献

The Effects of Rating Designs on Rater Classification Accuracy and Rater Measurement Precision in Large-Scale Mixed-Format Assessments.评分设计对大规模混合格式评估中评分者分类准确性和评分者测量精度的影响。

Appl Psychol Meas. 2023 Mar;47(2):91-105. doi: 10.1177/01466216231151705. Epub 2023 Jan 12.

Using Repeated Ratings to Improve Measurement Precision in Incomplete Rating Designs.在不完全评分设计中使用重复评分提高测量精度

J Appl Meas. 2018;19(2):148-161.

Exploring Incomplete Rating Designs With Mokken Scale Analysis.运用莫肯量表分析探索不完全评分设计

Educ Psychol Meas. 2018 Apr;78(2):319-342. doi: 10.1177/0013164416675393. Epub 2016 Oct 23.

Detecting Rater Biases in Sparse Rater-Mediated Assessment Networks.在稀疏评分者介导的评估网络中检测评分者偏差

Educ Psychol Meas. 2021 Oct;81(5):996-1022. doi: 10.1177/0013164420988108. Epub 2021 Jan 19.

Examining the Impacts of Rater Effects in Performance Assessments.审视评分者效应在绩效评估中的影响。

Appl Psychol Meas. 2019 Mar;43(2):159-171. doi: 10.1177/0146621618789391. Epub 2018 Aug 5.

Exploring the Combined Effects of Rater Misfit and Differential Rater Functioning in Performance Assessments.探索评分者不匹配和评分者差异功能在绩效评估中的综合影响。

Educ Psychol Meas. 2019 Oct;79(5):962-987. doi: 10.1177/0013164419834613. Epub 2019 Apr 2.

The Stabilizing Influences of Linking Set Size and Model-Data Fit in Sparse Rater-Mediated Assessment Networks.稀疏评分者介导评估网络中链接集大小与模型-数据拟合的稳定影响

Educ Psychol Meas. 2018 Aug;78(4):679-707. doi: 10.1177/0013164417703733. Epub 2017 Apr 12.

Does Sparseness Matter? Examining the Use of Generalizability Theory and Many-Facet Rasch Measurement in Sparse Rating Designs.稀疏性重要吗？审视通用izability理论和多面Rasch测量在稀疏评分设计中的应用。

Appl Psychol Meas. 2023 Sep;47(5-6):351-364. doi: 10.1177/01466216231182148. Epub 2023 Jun 7.

Examining rating scales using Rasch and Mokken models for rater-mediated assessments.使用拉施模型和莫肯模型检查评分量表以进行评分者介导的评估。

J Appl Meas. 2014;15(2):100-32.

Examining rating quality in writing assessment: rater agreement, error, and accuracy.审视写作评估中的评分质量：评分者一致性、误差与准确性。

J Appl Meas. 2012;13(4):321-35.

本文引用的文献

Detecting Rater Biases in Sparse Rater-Mediated Assessment Networks.在稀疏评分者介导的评估网络中检测评分者偏差

Educ Psychol Meas. 2021 Oct;81(5):996-1022. doi: 10.1177/0013164420988108. Epub 2021 Jan 19.

Automated language essay scoring systems: a literature review.自动化语言作文评分系统：文献综述

PeerJ Comput Sci. 2019 Aug 12;5:e208. doi: 10.7717/peerj-cs.208. eCollection 2019.

Exploring the Combined Effects of Rater Misfit and Differential Rater Functioning in Performance Assessments.探索评分者不匹配和评分者差异功能在绩效评估中的综合影响。

Educ Psychol Meas. 2019 Oct;79(5):962-987. doi: 10.1177/0013164419834613. Epub 2019 Apr 2.

Detecting Rater Effects under Rating Designs with Varying Levels of Missingness.在存在不同程度缺失值的评分设计下检测评分者效应。

J Appl Meas. 2018;19(3):243-257.

Using Repeated Ratings to Improve Measurement Precision in Incomplete Rating Designs.在不完全评分设计中使用重复评分提高测量精度

J Appl Meas. 2018;19(2):148-161.

文献检索

告别复杂PubMed语法，用中文像聊天一样搜索，搜遍4000万医学文献。AI智能推荐，让科研检索更轻松。

立即免费搜索

文件翻译

保留排版，准确专业，支持PDF/Word/PPT等文件格式，支持 12+语言互译。

免费翻译文档

深度研究

AI帮你快速写综述，25分钟生成高质量综述，智能提取关键信息，辅助科研写作。

立即免费体验