Evaluating LLM-Based Institutional Information Chatbot Responses Using a Preliminary Human-Scored Analytic Rubric and Automatic Metrics
DOI:
https://doi.org/10.21831/elinvo.v11i1.91571Keywords:
Analytic Rubric, Automatic Metrics, Chatbot Evaluation, Human Evaluation, System Usability ScaleAbstract
Large language models (LLMs) are increasingly used for answering institutional information enquiries in higher education, yet the quality of responses is not straightforward to evaluate, as factual accuracy alone does not account for interactive qualities such as clarity, conversational flow, error handling and personalisation. This pilot study developed, preliminary examined a human-scored analytic rubric for assessing ChatGPT responses in a higher-education institutional-information setting and explored the alignment of selected rubric scores with reference-based automatic metrics. We designed a literature-informed rubric comprising 15 criteria across five conceptual domains. Of these, 14 criteria were operationalised through 42 rubric questions, and system usability was rated separately using the System Usability Scale. Seventy-five qualified students from one higher-education institution rated the chatbot responses using a four-point scale. Preliminary evidence at the item level was provided by item-total correlations and Cronbach’s alpha, while Pearson and Spearman correlations were used to investigate the alignment between human scores and reference-based metrics namely ROUGE-1, ROUGE-2, ROUGE-L and SacreBLEU for four content-oriented criteria. The results showed positive but partial agreement between human ratings and referenced-based metrics, with stronger agreement for clarity and up-to-date response than for accuracy and relevance. These findings suggest that reference-based metrics can complement, but not replace, human evaluation for insitutional information chatbot assessment. The study was confined to one institution and did not incorporate inter-rater reliability, expert validation or factor analysis. Thus, the rubric should be seen as a preliminary evaluation instrument rather than a fully validated scale.
References
[1] E. Adamopoulou and L. Moussiades, “Chatbots: History, technology, and applications,” Machine Learning with Applications, vol. 2, no. July, p. 100006, 2020, doi: 10.1016/j.mlwa.2020.100006.
[2] S. H. Tampubolon et al., ChatGPT di Era AI: Revolusi Komunikasi Berbasis Mesin. Yayasan Kita Menulis, 2025.
[3] N. A. M. Mokmin and N. A. Ibrahim, “The evaluation of chatbot as a tool for health literacy education among undergraduate students,” Educ. Inf. Technol. (Dordr)., vol. 26, no. 5, pp. 6033–6049, 2021, doi: 10.1007/s10639-021-10542-y.
[4] S. Gupta et al., “Comprehensive Framework for Evaluating Conversational AI Chatbots,” The Review of Contemporary Scientific and Academic Studies, vol. 5, no. 2, 2025, doi: 10.55454/rcsas.5.02.2025.003.
[5] M. Hmoud, H. Swaity, E. Anjass, and E. M. Aguaded-Ramírez, “Rubric Development and Validation for Assessing Tasks’ Solving via AI Chatbots,” Electronic Journal of e-Learning, vol. 22, no. 6 Special Issue, pp. 1–17, 2024, doi: 10.34190/ejel.22.6.3292.
[6] R. A. Sekarwati, A. Sururi, R. Rakhmat, M. Arifin, and A. Wibowo, “Survei Metode Pengukuran Aplikasi Chatbot Berbasis Media Sosial,” Gema Teknologi, vol. 21, no. 2, pp. 67–73, 2021.
[7] A. Parks, E. Travers, R. Perera-Delcourt, M. Major, M. Economides, and P. Mullan, “Is this chatbot safe and evidence-based? A call for the critical evaluation of generative AI mental health chatbots,” J. Particip. Med., vol. 17, p. e69534, 2025.
[8] M. Walker, D. Litman, C. A. Kamm, and A. Abella, “PARADISE: A framework for evaluating spoken dialogue agents,” in 35th Annual Meeting of the Association for Computational Linguistics and 8th Conference of the European Chapter of the Association for Computational Linguistics, 1997, pp. 271–280.
[9] G. Datta, N. Joshi, and K. Gupta, “Analysis of automatic evaluation metric on low-resourced language: BERTScore vs BLEU score,” in International Conference on Speech and Computer, 2022, pp. 155–162.
[10] Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu, “G-eval: NLG evaluation using gpt-4 with better human alignment,” in Proceedings of the 2023 conference on empirical methods in natural language processing, 2023, pp. 2511–2522.
[11] R. Bommasani, P. Liang, and T. Lee, “Holistic evaluation of language models,” Ann. N. Y. Acad. Sci., vol. 1525, no. 1, pp. 140–146, 2023.
[12] N. Cannavaro, “Aplikasi Chatbot untuk Layanan Akademik Menggunakan Platform RASA Open Source dengan Fitur Two Stage Fallback,” Jurnal Ilmu Komputer dan Informatika, vol. 3, no. 1, pp. 53–64, 2023, doi: 10.54082/jiki.73.
[13] T. Ait Baha, M. El Hajji, Y. Es-Saady, and H. Fadili, “The Power of Personalization: A Systematic Review of Personality-Adaptive Chatbots,” SN Comput. Sci., vol. 4, no. 5, 2023, doi: 10.1007/s42979-023-02092-6.
[14] M. E. S. Simaremare, C. Pardede, I. N. I. Tampubolon, D. A. Simangunsong, and P. E. Manurung, “The Penetration of Generative AI in Higher Education: A Survey,” 2024 IEEE Integrated STEM Education Conference, ISEC 2024, no. March, pp. 1–5, 2024, doi: 10.1109/ISEC61299.2024.10664825.
[15] J. R. Lechien, A. Maniaci, I. Gengler, S. Hans, C. M. Chiesa-Estomba, and L. A. Vaira, “Validity and reliability of an instrument evaluating the performance of intelligent chatbot: the Artificial Intelligence Performance Instrument (AIPI),” European Archives of Oto-Rhino-Laryngology, vol. 281, no. 4, pp. 2063–2079, 2024.
[16] D. Guo et al., “DeepSeek-Coder: when the large language model meets programming–the rise of code intelligence,” arXiv preprint arXiv:2401.14196, 2024.
[17] S. Suwarno and C. Aeni, “Pentingnya Rubrik Penilaian Dalam Pengukuran Kejujuran Peserta Didik,” Edukasi: Jurnal Pendidikan, vol. 19, no. 1, p. 161, 2021, doi: 10.31571/edukasi.v19i1.2364.
[18] H. N. Boone and D. A. Boone, “Analyzing Likert data,” J. Ext., vol. 50, no. 2, 2012, doi: 10.34068/joe.50.02.48.
[19] I. A. Kamal and A. B. Cahyono, “Pemanfaatan Chatbot Berbasis Dialogflow dan Google Sheet Api Untuk Penyimpanan Laporan Komplain Konsumen Toko Online,” Automata, vol. 3, no. 2, pp. 1–5, 2022.
[20] Snigdha Tadanki and Sai Kiran Reddy Malikireddy, “Context-aware chatbots with data engineering for multi-turn conversations,” World Journal of Advanced Engineering Technology and Sciences, vol. 4, no. 1, pp. 063–078, 2021, doi: 10.30574/wjaets.2021.4.1.0061.
[21] M. W. Shiferaw, T. Zheng, A. Winter, L. A. Mike, and L. N. Chan, “Assessing the accuracy and quality of artificial intelligence (AI) chatbot-generated responses in making patient-specific drug-therapy and healthcare-related decisions,” BMC Med. Inform. Decis. Mak., vol. 24, no. 1, 2024, doi: 10.1186/s12911-024-02824-5.
[22] M. C. B. Nicole M. Radziwill, “Evaluating Quality of Chatbots and Intelligent Conversational Agents,” CEUR Workshop Proc., vol. 1982, pp. 40–49, 2017, doi: doi.org/10.48550/arXiv.1704.04579.
[23] H. Liang and H. Li, “Towards Standard Criteria for human evaluation of Chatbots: A Survey,” 2021.
[24] Y. Heryanto, F. Farahdinna, and S. Wijanarko, “Evaluasi Responsivitas dan Akurasi: Perbandingan Kinerja ChatGPT dan Google BARD dalam Menjawab Pertanyaan seputar Python,” Jurnal Riset Sistem Informasi Dan Teknik Informatika (JURASIK, vol. 9, no. 1, pp. 248–256, 2024.
[25] A. R. Syah and A. Prihanto, “Auto Response Messages pada Telegram Bot untuk Pelayanan Sistem Informasi Praktek Industri dan Skripsi dengan Metode Webhook,” Journal of Informatics and Computer Science (JINACS), vol. 3, no. 04, pp. 547–556, 2022, doi: 10.26740/jinacs.v3n04.p547-556.
[26] R. Lumbantoruan, X. Zhou, Y. Ren, and L. Chen, “I-CARS: An Interactive Context-Aware Recommender System,” in IEEE International Conference on Data Mining (ICDM), 2019.
[27] S. Raimer and M. Vanhauer, “I don’t understand you – error-handling as a key aspect of conversational design for chatbots,” Human Interaction & Emerging Technologies (IHIET-AI 2022): Artificial Intelligence & Future Applications, vol. 23, no. April 2022, 2022, doi: 10.54941/ahfe100878.
[28] S. Vemuru, E. John, and S. Rao, “Handling Complex Queries Using Query Trees,” no. June, 2021, doi: 10.36227/techrxiv.14845212.v1.
[29] Z. Ma, Z. Dou, Y. Zhu, H. Zhong, and J. R. Wen, “One Chatbot per Person: Creating Personalized Chatbots based on Implicit User Profiles,” SIGIR 2021 - Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 555–564, 2021, doi: 10.1145/3404835.3462828.
[30] M. D. S. Monteiro, V. C. Pereira, and L. C. D. C. Salgado, “Investigating politeness strategies in chatbots through the lens of Conversation Analysis,” in Proceedings of the XXII Brazilian Symposium on Human Factors in Computing Systems, 2023, pp. 1–12.
[31] R. Ren, M. Zapata, J. W. Castro, O. Dieste, and S. T. Acuna, “Experimentation for Chatbot Usability Evaluation: A Secondary Study,” IEEE Access, vol. 10, pp. 12430–12464, 2022, doi: 10.1109/ACCESS.2022.3145323.
[32] ISO, “Ergonomic requirements for office work with visual display terminals. Part 11: Guidance on usability,” ISO No 924111, vol. 2008, no. February 9, p. 22, 1994.
[33] R. Sakaeda and D. Kawahara, “Generate, Evaluate, and Select: A Dialogue System with a Response Evaluator for Diversity-Aware Response Generation,” NAACL 2022 - 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Student Research Workshop, pp. 76–82, 2022, doi: 10.18653/v1/2022.naacl-srw.10.
[34] W. A. Pinheiro, At the intersections: Queer high school students’ experiences with the teaching of mathematics for social justice. Indiana University, 2022.
[35] S. Zitzmann and G. A. Orona, “Why we might still be concerned about low Cronbach’s alphas in domain-specific knowledge tests,” Educ. Psychol. Rev., vol. 37, no. 2, p. 37, 2025.
[31] D. Firmansyah and Dede, “Common Sampling Techniques in Research Methodology: Literature Review,“ (In Indonesia) “Teknik Pengambilan Sampel Umum dalam Metodologi Penelitian: Literature Review,” J. Ilm. Pendidik. Holistik, vol. 1, no. 2, pp. 85–114, 2022, doi: 10.55927/jiph.v1i2.937.
[32] J. Friedl-Knirsch, F. Pointecker, S. Pfistermüller, C. Stach, C. Anthes, and D. Roth, “A Systematic Literature Review of User Evaluation in Immersive Analytics,” Comput. Graph. Forum, vol. 43, no. 3, 2024, doi: 10.1111/cgf.15111.
[33] J. W. Sitopu, I. R. Purba, and T. Sipayung, “Statistical Data Processing Training Using SPSS Application,“ (In Indonesia) “Pelatihan Pengolahan Data Statistik Dengan Menggunakan Aplikasi SPSS,” Dedik. Sains dan Teknol., vol. 1, no. 2, pp. 82–87, 2021, doi: 10.47709/dst.v1i2.1068.
[34] V. L. Lechien JR, Maniaci A, Gengler I, Hans S, Chiesa-Estomba CM, “Validity and reliability of an instrument evaluating the performance of intelligent chatbot: the Artificial Intelligence Performance Instrument (AIPI).,” doi: 0.1007/s00405-023-08219-y.
[35] N. Janna, “VALIDITY AND RELIABILITY TESTING CONCEPT USING SPSS,“ (In Indonesia) “KONSEP UJI VALIDITAS DAN RELIABILITAS DENGAN MENGGUNAKAN SPSS,” Callaloo, vol. 23, no. 3, pp. 885–891, 2000, doi: 10.1353/cal.2000.0135.
Downloads
Published
How to Cite
Issue
Section
Citation Check
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
The article published in ELINVO became ELINVO's right in publication.
This work by ELINVO is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.



