A Comparative Study on ERNIE Bot 4.0 Turbo and ChatGPT 4o’s Performance in Evaluating First-Year Undergraduate Persuasive Essays
DOI:
https://doi.org/10.61173/nk1ywa21Keywords:
AI-assisted essay evaluation, ERNIE Bot 40 Turbo, ChatGPT 4o, Corpus AnalysisAbstract
This study compares the performance of two prominent AI language models, ERNIE Bot 4.0 Turbo and ChatGPT 4o, in evaluating first-year undergraduate persuasive essays within the social sciences domain. Drawing from the Louvain Corpus of Native English Essays, a comprehensive collection of academic writings by British and American university students, this study aims to examine the models’ capabilities in assessing the grammatical correctness, vocabulary usage, coherence, content depth, and writing style of the essays. This study adopts a structured evaluation framework based on IELTS writing criteria to assess the models’ performance. A 40 persuasive essays from the Louvain Corpus were evaluated by both AI models and compared with human raters’ evaluations to ensure validity. The findings reveal distinct differences in the assessment styles of the two models. ChatGPT 4o exhibits a more critical approach, pinpointing areas for improvement, such as lack of argument development, coherence issues, and grammatical errors. Conversely, ERNIE Bot 4.0 Turbo offers a more balanced assessment, acknowledging essays’ strengths and suggesting improvement areas. Notably, ERNIE Bot’s evaluation highlights potential biases in AI-based assessment systems, particularly in its unequal emphasis on viewpoints. This comparative examination offers valuable perspectives on the advantages and constraints of AI models in assessing scholarly compositions, underscoring the significance of amalgamating varied AI functionalities to establish more all-inclusive and efficient feedback systems for learners. By understanding these differences, researchers and educators can better utilize AI-assisted essay evaluation systems to enhance student learning experiences.
References
[1] Kaledio, P., Robert, A., & Frank, L. (2024). The impact of artificial intelligence on students’ learning experience. Social Science Research Network. https://doi.org/10.2139/ssrn.4716747
[2] Zhu, L., Mou, W., Lai, Y., Lin, J., & Luo, P. (2024). Language and cultural bias in AI: comparing the performance of large language models developed in different countries on Traditional Chinese Medicine highlights the need for localized models. Journal of Translational Medicine, 22(1). https://doi. org/10.1186/s12967-024-05128-4
[3] Gayed, J. M., Carlon, M. K. J., Oriola, A. M., & Cross, J. S. (2022). Exploring an AI-based writing Assistant’s impact on English language learners. Computers and Education. Artificial Intelligence, 3, 100055. https://doi.org/10.1016/ j.caeai.2022.100055
[4] Makarius, E. E., Mukherjee, D., Fox, J. D., & Fox, A. K. (2020). Rising with the machines: A sociotechnical framework for bringing artificial intelligence into the organization. Journal of Business Research, 120, 262–273. https://doi.org/10.1016/ j.jbusres.2020.07.045
[5] Su, J., Zhong, Y., & Ng, D. T. K. (2022). A meta-review of literature on educational approaches for teaching AI at the K-12 levels in the Asia-Pacific region. Computers and Education. Artificial Intelligence, 3, 100065. https://doi.org/10.1016/ j.caeai.2022.100065
[6] Chaudhry, I. S., Sarwary, S. A. M., Refae, G. A. E., & Chabchoub, H. (2023). Time to revisit existing student’s performance evaluation approach in higher education sector in a new era of CHATGPT — a case study. Cogent Education, 10(1). https://doi.org/10.1080/2331186x.2023.2210461
[7] Theodosiou, A. A., & Read, R. C. (2023). Artificial intelligence, machine learning and deep learning: Potential resources for the infection clinician. Journal of Infection, 87(4), 287–294. https://doi.org/10.1016/j.jinf.2023.07.006
[8] Mazzone, M., & Elgammal, A. (2019). Art, creativity, and the potential of artificial intelligence. Arts, 8(1), 26. https://doi. org/10.3390/arts8010026
[9] Dwivedi, Y. K., Kshetri, N., Hughes, L., Slade, E. L., Jeyaraj, Dean&Francis A., Kar, A. K., Baabdullah, A. M., Koohang, A., Raghavan, V., Ahuja, M., Albanna, H., Albashrawi, M. A., Al-Busaidi, A. S., Balakrishnan, J., Barlette, Y., Basu, S., Bose, I., Brooks, L., Buhalis, D., ... Wright, R. (2023). Opinion Paper: “So what if ChatGPT wrote it?” Multidisciplinary perspectives on opportunities, challenges and implications of generative conversational AI for research, practice and policy. International Journal of Information Management, 71, 102642. https://doi. org/10.1016/j.ijinfomgt.2023.102642
[10] BRITISH COUNCIL. (2023). IELTS Writing Task 1 Guidelines. Retrieved from https://www.chinaielts.org/pdf/ UOBDs_WritingT1.pdf
[11] BRITISH COUNCIL. (2023). IELTS Writing Task 2 Guidelines. Retrieved from https://www.chinaielts.org/pdf/ UOBDs_WritingT2.pdf
Downloads
Published
Issue
Section
License
Copyright (c) 2024 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
