Evaluación de la validez de la Prueba Oral Automática Estandarizada de Inglés de la UCR en la Prueba de Seguimiento del MEP

Autores/as

DOI:

https://doi.org/10.15517/7exxmd35

Palabras clave:

Inteligencia Artificial, Evaluación de Lenguas Extranjeras, Validez de Pruebas, Evaluación del Habla

Resumen

El Programa de Evaluación de Lenguas Extranjeras (PELEx) de la Universidad de Costa Rica está a la vanguardia en la aplicación del reconocimiento automático de voz (ASR, por sus siglas en inglés) y la inteligencia artificial (IA) en la evaluación de lenguas extranjeras dentro de América Latina (LATAM), lo cual genera la necesidad de una investigación continua, recopilación de datos y evaluación de la confiabilidad para recopilar más evidencias de validez de estas pruebas. Este trabajo compara los resultados obtenidos durante la aplicación de la Prueba de Seguimiento del Ministerio de Educación Pública (MEP) de Costa Rica en 2023 por el PELEx de la Universidad de Costa Rica (UCR) a 7454 estudiantes de secundaria pública costarricenses, mediante un muestreo estadístico de las pruebas, en comparación con la calificación de un evaluador humano certificado del PELEx, con el fin de aportar evidencia de validez concurrente, es decir, el grado de coincidencia entre la calificación humana y la de la IA. Se ofrece un análisis de los puntajes y recomendaciones para investigaciones futuras orientadas a la validación continua y a la recopilación de evidencias de validez de constructo y confiabilidad de calificación del examen oral automatizado.

Descargas

Los datos de descarga aún no están disponibles.

Referencias

ACTFL. (2012). Oral proficiency interview. ACTFL. https://www.actfl.org/assessments/postsecondary-assessments/opi

Agrawal, A., Gans, J. S., & Goldfarb, A. (2023, June 12). The Turing transformation: Artificial intelligence, intelligence augmentation, and skill premiums. Brookings Institution. https://www.brookings.edu/articles/the-turing-transformation-artificial-intelligence-intelligence-augmentation-and-skill-premiums/

Aoun, J. (2017). Robot-proof: Higher education in the age of artificial intelligence. The MIT Press.

Araya Garita, W. (2021). Dominio lingüístico en inglés en estudiantes de secundaria para el año 2019 en Costa Rica. Revista de Lenguas Modernas, (34), 1–21. https://doi.org/10.15517/rlm.v0i34.43364

Brown, H. D., & Abeywickrama, P. (2019). Language assessment: Principles and classroom practices (3rd ed.). Pearson.

Cheng, J. (2018). Realtime scoring of an oral reading assessment on mobile devices. In Proceedings of Interspeech 2018(pp. 1621–1625). International Speech Communication Association. https://doi.org/10.21437/Interspeech.2018-34

Cizek, G. J., Rosenberg, S. L., & Koons, H. H. (2008). Sources of validity evidence for educational and psychological tests. Educational and Psychological Measurement, 68 (3), 397-412.

Council of Europe. (2020). Common European Framework of Reference for Languages: Learning, teaching, assessment—Companion volume. Council of Europe Publishing. https://www.coe.int/en/web/common-european-framework-reference-languages

Das, A. K., Nayak, J., Naik, B., Dutta, S., & Pelusi, D. (2022). Computational intelligence in pattern recognition proceedings of CIPR 2021. Springer.

Dodigovic, P. M. (2005). Artificial intelligence in second language learning: Raising error awareness. Multilingual Matters.

Duolingo, Inc. (2023, June 12). Assessing speaking on the Duolingo English Test (Duolingo Research Report DRR 23 03). https://web.archive.org/web/20240324111501/https://duolingo-testcenter.s3.amazonaws.com/media/resources/speaking-whitepaper.pdf

Economist Impact. (2022). Seizing the opportunity: The future of AI in Latin America. https://impact.economist.com/technology-innovation/seizing-opportunity-future-ai-latin-america

Edmett, A., Ichaporia, N., Crompton, H., & Crichton, R. (2024). Artificial intelligence and English language teaching: Preparing for the future (2nd ed.). British Council. https://www.teachingenglish.org.uk/sites/teacheng/files/2024-08/AI_and_ELT_Jul_2024.pdf

Educational Testing Service. (2020). Reliability and comparability of TOEFL iBT® scores (TOEFL Research Insight Series, Vol. 3). https://www.ets.org/content/dam/ets-india/pdfs/toefl/toefl-ibt-insight-s1v3.pdf

Evanini, K., Xie, S., & Zechner, K. (2013). Prompt-based content scoring for automated spoken language assessment. In Proceedings of the Eighth Workshop on Innovative Use of NLP for Building Educational Applications (pp. 157–162). Association for Computational Linguistics. https://aclweb.org/anthology/W/W13/W13-1721.pdf

Fan, J., Jin, Y., Kong, X., & Zhang, X. (2021). How do test takers prepare for computerized speaking tests? A comparative study of the PTE Academic and the CET-SET. University of Melbourne.

The Guardian. (2017, August 8). Computer says no: Irish vet fails oral English test needed to stay in Australia. The Guardian. https://www.theguardian.com/australia-news/2017/aug/08/computer-says-no-irish-vet-fails-oral-english-test-needed-to-stay-in-australia

Hao, K. (2020, April 2). This is how AI bias really happens-and why it’s so hard to fix. MIT Technology Review. https://www.technologyreview.com/2019/02/04/137602/this-is-how-ai-bias-really-happensand-why-its-so-hard-to-fix/

Johnstone, A., & Altmann, G. (1984). Automated speech recognition: A framework for research. Department of Artificial Intelligence, University of Edinburgh.

Kang, O., Rubin, D., & Kermad, A. (2019). The effect of training and rater differences on oral proficiency assessment. Language Testing, 36(4), 481–504. https://doi.org/10.1177/0265532219849522

Kunnan, A. J. (2000). Fairness and justice for all. In A. J. Kunnan (Ed.), Fairness and validation in language assessment (pp. 1-14). Cambridge University Press.

McNamara, T., & Roever, C. (2006). Language testing: The social dimension. Blackwell Publishing.

Mercer, S. H., & Cannon, J. E. (2022). Validity of automated learning progress assessment in English written expression for students with learning difficulties. Journal for Educational Research Online / Journal für Bildungsforschung Online, 14(1), 39–60. https://doi.org/10.31244/jero.2022.01.03

Mkrttchian, V. (2021). Impact of Artificial and Natural Intelligence Technologies With Avatar-Based Teaching, Learning, and Research in Russian Modern Universities. In S. Verma & P. Tomar (Eds.), Impact of AI Technologies on Teaching, Learning, and Research in Higher Education (pp. 204-221). IGI Global Scientific Publishing. https://doi.org/10.4018/978-1-7998-4763-2.ch013

Nakatsuhara, F., & Berry, V. (2021). Use of innovative technology in oral language assessment. Assessment in Education: Principles, Policy & Practice, 28(4), 343–349. https://doi.org/10.1080/0969594x.2021.2004530

Narasimhan, S. (2023, October 26). Challenges and opportunities of generative ai in assessments. HurixDigital. https://web.archive.org/web/20250420171101/https://www.hurix.com/blogs/challenges-and-opportunities-of-generative-ai-in-assessments/

OECD. (2022). The strategic and responsible use of artificial intelligence in the public sector of Latin America and the Caribbean (OECD Public Governance Reviews). OECD Publishing. https://doi.org/10.1787/1f334543-en

Pearson. (2024). PTE Academic score guide. https://web.archive.org/web/20250219130145/https://www.pearsonpte.com/ctf-assets/yqwtwibiobs4/5Sz9Ur4qbus8AEOQdetkAj/69f6c1f2e2870980740b10a2ea9b467f/pte-academic-test-taker-score-guide-nov-2024-v4.pdf

Pirrone, A. (2023, March 23). Resist AI by rethinking assessment. LSE Higher Education Blog. https://blogs.lse.ac.uk/highereducation/2023/03/23/resist-ai-by-rethinking-assessment/

PELEx. (n.d.). Prueba de Dominio Lingüístico (PDL) para secundaria – MEP. Universidad de Costa Rica. https://www.pelex.ucr.ac.cr/pruebasMEPDominioSecundaria.html

Quesada, A., Araya, W., & Fallas, J. A. (2023). La enseñanza y aprendizaje del inglés en secundaria pública costarricense del siglo XXI: Innovaciones, brechas y desafíos. Informe final para el Noveno Informe Estado de la Educación. Programa Estado de la Nación – CONARE.

Richardson, M., Leaton Gray, S., Popov, J., & Maddox, B. (2023). Computer-based tests and machine marking: Candidates’ perceptions and beliefs about the test taking experience. UCL Discovery. https://discovery.ucl.ac.uk/id/eprint/10177154/

Rizvi, S., Waite, J., & Sentance, S. (2023). Artificial Intelligence Teaching and learning in K-12 from 2019 to 2022: A systematic literature review. Computers and Education: Artificial Intelligence, 4, 100145. https://doi.org/10.1016/j.caeai.2023.100145

Rodrigo, M. M., Matsuda, N., Cristea, A. I., & Dimitrova, V. (Eds.). (2022). Artificial intelligence in education: 23rd international conference, AIED 2022, proceedings. Springer.

Roll, I., McNamara, D., Sosnovsky, S., Luckin, R., & Dimitrova, V. (Eds.). (2021). Artificial intelligence in education: 22nd international conference, AIED 2021, proceedings. Springer.

Shabtai, N. R. (Ed.) (2010). Advances in speech recognition. Sciyo.

Schilling, S. G. (2004). Conceptualizing the validity argument: An alternative approach. Measurement, 2, 178-182.

Speechace. (2023). Speechace speaking test. https://www.speechace.com/speaking-test/

Swiecki, Z., Khosravi, H., Chen, G., Martinez-Maldonado, R., Lodge, J. M., Milligan, S., Selwyn, N., & Gašević, D. (2022). Assessment in the age of Artificial Intelligence. Computers and Education: Artificial Intelligence, 3, 100075. https://doi.org/10.1016/j.caeai.2022.100075

Waibel, A., & Lee, K.-F. (1990). Readings in speech recognition. Morgan Kaufmann.

Wang, N., Rebolledo-Mendez, G., Matsuda, N., Santos, O. C., & Dimitrova, V. (Eds.). (2023). Artificial intelligence in education: 24th International Conference, AIED 2023, Tokyo, Japan, July 3–7, 2023, proceedings (Vol. 13916). Springer. https://doi.org/10.1007/978-3-031-36272-9

Weir, C. J. (2005). Language testing and validation: An evidence-based approach. Palgrave Macmillan.

Wet, F. de, Walt, C. van, & Niesler, T. (2007). Automatic large-scale oral language proficiency assessment. Interspeech 2007. https://doi.org/10.21437/interspeech.2007-90

Witzigmann, S. & Sachse, S. (2021). Diagnostic competencies of prospective teachers of French as a foreign language: judgement of oral language samples. Research in Subject-matter Teaching and Learning (RISTAL), 4(1), 71-87. https://doi.org/10.23770/rt1847

Yong, Q. (2020). Application analysis of artificial intelligence in oral English assessment. Journal of Physics: Conference Series, 1533(3), 032028. https://doi.org/10.1088/1742-6596/1533/3/032028

Young, J. R. (2023, October 5). As AI chatbots rise, more educators look to oral exams—With high-tech twist. EdSurge. https://www.edsurge.com/news/2023-10-05-as-ai-chatbots-rise-more-educators-look-to-oral-exams-with-high-tech-twist

Yu, Y., Han, L., Du, X., & Yu, J. (2022). An oral English evaluation model using artificial intelligence method. Mobile Information Systems, 2022, Article 3998886. https://doi.org/10.1155/2022/3998886

Zhai, C., & Wibowo, S. (2023). A systematic review on artificial intelligence dialogue systems for enhancing English as foreign language students’ interactional competence in the University. Computers and Education: Artificial Intelligence, 4, 100134. https://doi.org/10.1016/j.caeai.2023.100134

Publicado

2026-05-28

Número

Sección

Estudios sobre didáctica de lenguas extranjeras