Assessing Generative AI with Context-Augmented Zero-Shot Prompting for HOTS Question Generation Aligned with Bloom’s Taxonomy
DOI:
https://doi.org/10.15294/edukom.v12i2.34437Keywords:
Automatic Question Generation, Bloom’s Taxonomy, Context-Augmented Zero-Shot Prompting, Generative AI Higher Order Thinking Skills (HOTS)Abstract
This study investigates the use of HOTS-GenAI, a generative AI system employing context-augmented zero-shot prompting, to automatically generate multiple-choice questions aligned with higher-order thinking skills (HOTS) in Bloom’s taxonomy. A dataset of 200 items for vocational high schools was validated by three experts. The ground truth data demonstrated good quality with an inter-rater reliability of 0.75 (Gregory’s Index). System performance across analysis (C4), evaluation (C5), and creation (C6) levels was evaluated using Content Validity Index (CVI), gap analysis, and confusion-matrix-based metrics. The findings revealed that HOTS-GenAI performed relatively well at the analytical level, where 70% of items met the HOTS threshold, supported by higher expert consensus. However, only 10% of items achieved the threshold for evaluation, and none for creation. CVI results indicated moderate validity overall, with stronger agreement for C4 than for C5 or C6. Confusion matrix analysis further confirmed this imbalance: accuracy and F1-scores were highest for analysis items but dropped sharply for evaluation and creation, where recall and precision were near zero. These results suggest that while HOTS-GenAI has potential in generating analytical questions, its capacity to model evaluative and creative tasks remains underdeveloped. Future research should involve larger datasets, refined prompt design, and more operational rubrics to enhance both validity and reliability in AI-generated HOTS assessments.
References
Ahmed, A., Kerr, E., & O’Malley, A. (2025). Quality assurance and validity of AI-generated single best answer questions. BMC Medical Education, 25(1), 300. https://doi.org/10.1186/s12909-025-06881-w
Aithal, A., & Aithal, P. (2020). Development and validation of survey questionnaire & experimental data–a systematical review-based statistical approach. International Journal of Management, Technology, and Social Sciences (IJMTS), 5(2), 233–251.
Arfandi, A., SURYANI, H., Panennungi, P., & Sampebua, O. (2021). The Ability of Vocational High School Teachers to Developing HOTS Question. International Journal of Environment, Engineering, and Education, 3(3), 1–6.
Arslan Mancar, S., & Gülleroğlu, H. D. (2022). Comparison of Inter-Rater Reliability Techniques in Performance-Based Assessment. International Journal of Assessment Tools in Education, 9(2), 515–533. https://doi.org/10.21449/ijate.993805
Atayolu, Y., & Kutlu, Y. (2024). Similarity and Classification Analysis in Turkish Question Generation of AI tools According to Bloom’s Taxonomy. Tethys Env. Sci, 1(2), 87–98.
Bitew, S. K., Hadifar, A., Sterckx, L., Deleu, J., Develder, C., & Demeester, T. (2024). Learning to Reuse Distractors to Support Multiple-Choice Question Generation in Education. IEEE Transactions on Learning Technologies, 17, 375–390. https://doi.org/10.1109/TLT.2022.3226523
C. Shu, N. Yao, Y. Chen, V. Wijeratne, L. Ma, J. Loo, K. K. Chai, A. Alam, & A. Abuelmaatti. (2025). Ai-Assisted Multiple-Choice Questions Generation with Multimodal Large Language Models in Engineering Higher Education. 2025 IEEE Global Engineering Education Conference (EDUCON), 1–9. https://doi.org/10.1109/EDUCON62633.2025.11016449
Caufield, J. H., Hegde, H., Emonet, V., Harris, N. L., Joachimiak, M. P., Matentzoglu, N., Kim, H., Moxon, S., Reese, J. T., Haendel, M. A., Robinson, P. N., & Mungall, C. J. (2024). Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): A method for populating knowledge bases using zero-shot learning. Bioinformatics, 40(3), btae104. https://doi.org/10.1093/bioinformatics/btae104
Chan, Y.-H., & Fan, Y.-C. (2019). A Recurrent BERT-based Model for Question Generation. Proceedings of the 2nd Workshop on Machine Reading for Question Answering, 154–162. https://doi.org/10.18653/v1/D19-5821
Gupta, S., Grover, S., Menon, V., Indu, P., Vidhukumar, K., & Chacko, D. (2025). Item generation and establishing face and content validity of a rating scale: A primer. Indian Journal of Psychiatry, 67(8). https://journals.lww.com/indianjpsychiatry/fulltext/2025/08000/item_generation_and_establishing_face_and_content.12.aspx
Himawan, R., & Suyata, P. (2023). Analisis Sebaran Level Kognitif HOTS Berdasarkan Taksonomi Bloom pada Soal Penilaian Harian Materi Teks Pidato Persuasif di SMPN 1 Bambanglipuro Bantul. Stilistika: Jurnal Pendidikan Bahasa Dan Sastra, 16(1), 89–100.
Huber, T., & Niklaus, C. (2025). LLMs meet Bloom’s Taxonomy: A Cognitive View on Large Language Model Evaluations.
Isnaeni, W., Rudyatmi, E., Ridlo, S., Ingesti, S., & Adiani, L. R. (2021). Improving students’ communication skills and critical thinking ability with ICT-oriented problem-based learning and the assessment instruments with HOTS criteria on the immune system material. Journal of Physics: Conference Series, 1918(5), 052048. https://doi.org/10.1088/1742-6596/1918/5/052048
Koto, F. (2025). Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia (No. arXiv:2409.08564). arXiv. https://doi.org/10.48550/arXiv.2409.08564
Lam, Y. Y., Chu, S. K. W., Ong, E. L. C., Suen, W. W. L., Xu, L., Lam, L. C. L., & Wong, S. M. Y. (2024). Comparative Study of GenAI (ChatGPT) vs. Human in Generating Multiple Choice Questions Based on the PIRLS Reading Assessment Framework. Proceedings of the Association for Information Science and Technology, 61(1), 537–540.
Lee, U., Jung, H., Jeon, Y., Sohn, Y., Hwang, W., Moon, J., & Kim, H. (2024). Few-shot is enough: Exploring ChatGPT prompt engineering method for automatic question generation in english education. Education and Information Technologies, 29(9), 11483–11515.
Li, Y. (2023). A Practical Survey on Zero-shot Prompt Design for In-context Learning. Proceedings of the Conference Recent Advances in Natural Language Processing - Large Language Models for Natural Language Processings, 641–647. https://doi.org/10.26615/978-954-452-092-2_069
Mulyono, H., Istiyati, S., Atmojo, I. R. W., & Ardiansyah, R. (2019). Kompetensi Guru Dalam Penyusunan Soal Higher Order Thinking Skills (HOTS) Berbasis Critical Thinking Sesuai Kurikulum 2013 Guna Mengakselerasi Education 4.0. 7(2), 108–1111.
P. Y. Reddy, N. S. Satya Aarthi, S. Pooja, B. Veerendra, & S. D. (2025). A Deep Learning Approach for Identifying LLM-Generated Text. 2025 International Conference on Artificial Intelligence and Data Engineering (AIDE), 359–364. https://doi.org/10.1109/AIDE64228.2025.10987319
Pawar, P., Dube, R., Joshi, A., Gulhane, Z., & Patil, R. (2024). Automated Generation and Evaluation of MultipleChoice Quizzes using Langchain and Gemini LLM. 2024 International Conference on Electrical Electronics and Computing Technologies (ICEECT), 1, 1–7. https://doi.org/10.1109/ICEECT61758.2024.10739326
Retnawati, H. (2016). Proving content validity of self-regulated learning scale (The comparison of Aiken index and expanded Gregory index). REID (Research and Evaluation in Education), 2(2), 155–164. https://doi.org/10.21831/reid.v2i2.11029
Sallam, M., Snygg, J., Hamdan, A., Allam, D., Kassem, R., & Damani, M. (2025). Evaluating the Performance of Seven Large Language Models (GPT4.5, Gemini, Copilot, Claude, Perplexity, DeepSeek, and Manus) in Answering Healthcare Quality Management Inquiries. Research and Advances in Education, 4(4), 39–50. https://doi.org/10.63593/RAE.2788-7057.2025.05.005
Sudana, I. M., Oktarina, N., Apriyani, D., & Achmadi, T. A. (2020). An Implementation of HOTS Based Learning Strategy in Vocational High Schools. International Journal of Innovation, 13(12).



