Forthcoming

Score Reliability, Classification Decision Accuracy, and Interrater Reliability of TOEFL Primary® Tests

Authors

  • Feifei Li ETS Research Institute Author

DOI:

https://doi.org/10.64634/ys398g44

Abstract

This study evaluates the score reliability, classification decision accuracy, and interrater reliability of TOEFL Primary® test scores. Using large operational datasets from Reading, Listening (Step 1 and Step 2), and Speaking tests, reliability was estimated via Cronbach’s alpha and standard error of measurement (SEM). Classification decision accuracy for CEFR levels was estimated using the Livingston and Lewis (1995) method. Interrater reliability for Speaking was evaluated through percentage agreement, correlation, and quadratic weighted kappa based on a double-scoring design. Results indicate strong internal consistency across all test sections (α = .83–.90) and relatively small SEMs (1.19–1.61), supporting score precision. Classification decision accuracy was high for Reading and Listening (approximately .78–.82) and somewhat lower for Speaking (.76), reflecting greater complexity in performance-based assessment. Interrater reliability for Speaking demonstrated moderate to strong agreement, with correlations and weighted kappa values ranging from .68 to .88 and adjacent agreement exceeding 96%. Overall, findings provide strong evidence supporting the reliability and interpretability of TOEFL Primary scores for reporting CEFR-aligned proficiency levels and making educational decisions about young English learners.

Suggested citation: Li, F. (In Press). Score Reliability, Classification Decision Accuracy, and Interrater Reliability of TOEFL Primary® Tests. ETS Research Report Series. https://doi.org/10.64634/ys398g44

Downloads

Published

2026-09-02

Issue

Section

Memorandum