سنجش روانی تربیتی

سنجش روانی تربیتی

داده‌های بیشتر روایی کمتر: پارادوکس ورود سوابق تحصیلی به معادلات پذیرش دانشجو

نوع مقاله : مقاله پژوهشی

نویسندگان
1 گروه روش ها و برنامه ریزی آموزشی و درسی، دانشکده روان‌شناسی و علوم تربیتی، دانشگاه تهران، تهران، ایران.
2 گروه علوم تربیتی، دانشکده علوم تربیتی و روانشناسی، دانشگاه شهید چمران اهواز، اهواز، ایران
چکیده
زمینه: آزمون سراسری ورود به دانشگاه در ایران به‌عنوان آزمونی سرنوشت‌ساز، نمره داوطلبان را از ترکیب کنکور (آزمون هنجار - مرجع) و سوابق تحصیلی (آزمون ملاک - مرجع) محاسبه می‌کند. هدف این پژوهش، ارزیابی انتقادی روایی این ادعا است که ترکیب این دو نمره، سنجشی جامع‌تر و معتبرتر از توانایی حقیقی داوطلبان فراهم می‌آورد.
روش: روش پژوهش مبتنی بر طرحی تحلیلی - شبیه‌سازی دو مرحله‌ای است. در فاز اول، با استفاده از چارچوب روایی استدلال‌محور کین، پیش‌فرض‌های نظری این ادعا نقد شد. در فاز دوم، شبیه‌سازی آماری برای نمایش کمّی تأثیر اثر سقف بر رتبه‌بندی داوطلبان انجام گرفت.
یافته‌ها: نتایج شبیه‌سازی نشان داد که اثر سقف، قدرت تمایزگذاری سوابق تحصیلی را در داوطلبان ممتاز (۱۰ درصد برتر) به‌شدت تضعیف می‌کند. هم‌زمان با نزدیک شدن میانگین نمرات جامعه به سقف مقیاس، همبستگی رتبه‌ای میان توانایی واقعی و نمره مشاهده‌شده در این گروه به سمت صفر میل می‌کند و رتبه‌بندی به‌شدت تحت تأثیر خطای تصادفی قرار می‌گیرد.
نتیجه‌گیری: نتیجه‌گیری این پژوهش نشان می‌دهد که ادعای بنیادین سیاست ترکیب دو آزمون، از نظر روایی قابل دفاع نیست. اثر سقف در امتحانات نهایی، این بخش از نمره کل را به منبعی اطلاعاتی ناکارآمد و پرنوسان (خطا) برای تمایزگذاری میان داوطلبان برتر تبدیل می‌کند. این امر نه‌تنها اعتبار کل سیستم رتبه‌بندی را در حساس‌ترین بخش رقابت تضعیف می‌کند، بلکه با وارد کردن خطای سیستماتیک، به کاهش عدالت در فرایند پذیرش دانشجو منجر می‌شود؛ بنابراین، ترکیب فعلی نمرات کنکور و سوابق تحصیلی عملی سؤال‌برانگیز در روش‌شناختی است که نیازمند بازنگری جدی در سیاست‌گذاری‌های نظام سنجش و پذیرش است تا از تضییع حقوق داوطلبان شایسته جلوگیری شود.
کلیدواژه‌ها

عنوان مقاله English

More Data, Less Validity: The Paradox of Incorporating Academic Records into Student Selection Equations

نویسندگان English

Ebrahim Khodaie 1
Mahdi Karvandi Renani 1
Mojtaba Jahanifar 2
1 Department of Tecching Methods and Educational and Curriculum Planning, Faculty of Psychology and Education, University of Tehran, Tehran, Iran
2 Department of Educational Sciences, Faculty of Education and Psychology, Shahid Chamran University of Ahvaz, Ahvaz, Iran.
چکیده English

Background: Iran’s high-stakes national university entrance examination calculates candidates' final scores by combining the Konkur (a norm-referenced test) and academic records (a criterion-referenced test). This study critically evaluates the claim that combining these scores provides a more comprehensive assessment of candidates' true abilities.
Methods: Using a two-stage analytical-simulation design, we first critiqued the theoretical assumptions of this policy using Kane’s argument-based approach to validation, focusing on score interpretation and decision consequences. Second, a statistical simulation quantified the impact of the ceiling effect on candidates’ rankings.
Results: Results revealed that the ceiling effect severely diminishes the discriminating power of academic records among top-tier candidates (the top 10%). As mean scores approach the scale's upper limit, the rank correlation between true ability and observed scores nears zero, rendering rankings highly susceptible to random error.
Conclusion: Findings indicate that the fundamental premise of combining these assessments is untenable from a validity perspective. The ceiling effect in final exams reduces this score component into an inefficient, error-prone source for differentiating elite candidates. This undermines the ranking system's credibility in the most competitive segment and compromises selection fairness by introducing systematic error. Consequently, combining Konkur scores and academic records is methodologically questionable, necessitating a serious revision of admission policies to safeguard deserving candidates' rights.

کلیدواژه‌ها English

Validity
National Entrance Examination
Ceiling Effect
Simulation
Final Exams
منابع
جهانی‌فر، م؛ خدایی، ا؛ یونسی، ج و موسوی، س. (1396). نقش هموار‌سازی در تبدیل غیر‌خطی نمره‌های خام به نمره‌های مقیاس نرمال. مطالعات اندازه‌گیری و ارزشیابی آموزشی. 7(20). 103-133.
کروندی رنانی، م؛ صالحی، ک و خدایی، ا. (1405). فراسوی عینیت مکانیکی: چالش‌های عینیت و راهکارهای نوین در پذیرش دانشجو. فصلنامه پژوهش و برنامه‌ریزی در آموزش عالی. e729749.
References
American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.
Borsboom, D., & Wijsen, L. D. (2016). Frankenstein’s validity monster: The value of keeping politics and science separated. Assessment in Education: Principles, Policy & Practice, 23(2), 281-283.
Borsboom, D., Mellenbergh, G. J., & van Heerden, J. (2004). The Concept of Validity. Psychological Review, 111(4), 1061-1071.
Crocker, L. & Algina, J. (1986) Introduction to Classical and Modern Test Theory. Harcourt, New York, 527.
Cronbach, L. J. (1988). Five perspectives on the validity argument. In Test validity. (pp. 3-17). Lawrence Erlbaum Associates, Inc.
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302.
Duhem, P. (2021). The Aim and Structure of Physical Theory. Princeton: Princeton University Press.
Jahanifar,M, Khodaie,E, Yunesi,J and Musavi,S A . (2018). Smoothing methods role in raw scores Non-linear transformation to Normalized scale scores. Educational Measurement and Evaluation Studies7(20), 103-133. (Persian)
Kane, M. (2012). All validity is construct validity. Or is it? Measurement: Interdisciplinary Research and Perspectives, 10(1-2), 66-70.
Kane, M. T. (2006). Validation. In R. L. Brennan (Ed.), Educational measurement (Fourth Edition ed., pp. 17-64).
Kane, M. T. (2013). Validating the Interpretations and Uses of Test Scores. Journal of Educational Measurement, 50(1), 1-73.
Kane, M. T. (2016). Validity as the evaluation of the claims based on test scores. Assessment in Education: Principles, Policy & Practice, 23(2), 309-311.
Karvandi Renani,M , Salehi,K and Khodaie,E . (2026). Moving Beyond Mechanical Objectivity: Addressing the Challenges and Exploring Novel Solutions in Iranian University AdmissionsQuarterly Journal of Research and Planning in Higher Education32(2), 165-187.
Kelley, T. L. (1927). Interpretation of educational measurements. World Book Co.
Kim, J.-E. (2025). What makes a killer question killer? A text mining analysis of high-difficulty questions in the Korean CSAT English section. English Teaching, 80(3), 3–22.
Kim, Y. (2023, December 17). South Korea to reform education system by removing “killer questions” from Suneung. The Herald Insight.
Lee, H.-J. (2023, June 26). 'Killer' CSAT questions require obscure knowledge. Korea JoongAng Daily. https://koreajoongangdaily.joins.com/2023/06/26/national/socialAffairs/CSAT-college-exam-killer-questions/20230626191501015.html
Majone, G. (1989). Evidence, Argument, and Persuasion in the Policy Process. New Haven, CT: Yale University Press.
Markus, K. A., Borsboom, D. (2013). Frontiers in test validity theory : measurement, causation, and meaning (1st edition ed.). Routledge.
Messick, S. (1989). Validity. In R. Linn (Ed.), Educational Measurement (3rd ed.) (pp. pp. 13–100). American Council on Education.
Newton, P. E., & Shaw, S. D. (2014). Validity in Educational & Psychological Assessment.
Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). New York: McGraw-Hill.
Yeung, J., & Seo, Y. (2023, July 1). South Korea is cutting ‘killer questions’ from an 8-hour exam some blame for a fertility rate crisis. CNN. https://edition.cnn.com/2023/07/01/asia/south-korea-college-exam-fertility-pressure-intl-hnk-dst
 

  • تاریخ دریافت 12 خرداد 1405
  • تاریخ بازنگری 14 خرداد 1405
  • تاریخ پذیرش 20 خرداد 1405