Beyond translated benchmarks: a culturally grounded evaluation of large language models for Indonesian language understanding
DOI:
https://doi.org/10.67490/ijcl.v1i2.911Keywords:
cultural commonsense, Indonesian language understanding, large language models, multilingual evaluation, source-grounded benchmarkAbstract
Background: Translated benchmarks have expanded multilingual large language model evaluation, yet Indonesian remains a sociolinguistically dense setting where meaning is shaped by public institutions, regional languages, religious discourse, and culturally situated inference. Objective: This study aims to examine how Indonesian language understanding can be evaluated through source-grounded, culturally embedded corpus items rather than through translated English-centric tasks. Method: Using a qualitative-diagnostic benchmark design, this study constructs and codes an auditable corpus of public Indonesian documents, benchmark references, and contextual materials across semantic, pragmatic, cultural, regional, institutional, safety, and multimodal dimensions. Results: The analysis shows that semantic comprehension appears as a baseline requirement but does not sufficiently capture Indonesian understanding. Institutional register and cultural commonsense emerge as dominant dimensions, indicating that Indonesian prompts frequently require recognition of public authority, local values, collective memory, and socially constrained meaning. Implication: Task sensitivity is highest when documents combine policy language, cultural inference, safety cues, regional references, and pragmatic restraint, revealing why ordinary accuracy metrics may obscure fragile understanding. Novelty: This study contributes a culturally grounded evaluation perspective that reframes Indonesian LLM assessment as source-verifiable, context-sensitive interpretation rather than transferable language performance, and offers methodological resources for future multilingual benchmarking in low-resource and culturally plural settings within Southeast Asian digital ecologies
References
[1] Y.-C. Chang, X. Wang, J. Wang, Y. Wu, K. Zhu, H. Chen, et al., “A survey on evaluation of large language models,” ACM Transactions on Intelligent Systems and Technology, vol. 15, pp. 1–45, 2023, doi: 10.1145/3641289.
[2] V. D. Lai, N. T. Ngo, A. P. B. Veyseh, H. Man, F. Dernoncourt, T. Bui, and T. H. Nguyen, “ChatGPT beyond English: Towards a comprehensive evaluation of large language models in multilingual learning,” arXiv, 2023, doi: 10.48550/arxiv.2304.05613.
[3] H. Naveed, A. Khan, S. Qiu, M. Saqib, S. Anwar, M. Usman, N. Barnes, and A. S. Mian, “A comprehensive overview of large language models,” ACM Transactions on Intelligent Systems and Technology, vol. 16, pp. 1–72, 2023, doi: 10.1145/3744746.
[4] A. F. Aji and T. Cohn, “LoraxBench: A multitask, multilingual benchmark suite for 20 Indonesian languages,” arXiv, 2025, doi: 10.48550/arxiv.2508.12459.
[5] M. F. Azmi, M. D. A. Kautsar, A. Wicaksono, and F. Koto, “IndoSafety: Culturally grounded safety for LLMs in Indonesian languages,” arXiv, 2025, doi: 10.48550/arxiv.2506.02573.
[6] S. Cahyawijaya, G. I. Winata, B. Wilie, K. Vincentio, X. Li, A. Kuncoro, et al., “IndoNLG: Benchmark and resources for evaluating Indonesian natural language generation,” Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, doi: 10.18653/v1/2021.emnlp-main.699.
[7] F. Koto, N. Aisyah, H. Li, and T. Baldwin, “Large language models only pass primary school exams in Indonesia: A comprehensive test on IndoMMLU,” Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 12359–12374, 2023, doi: 10.48550/arxiv.2310.04928.
[8] F. Koto, R. Mahendra, N. Aisyah, and T. Baldwin, “IndoCulture: Exploring geographically influenced cultural commonsense reasoning across eleven Indonesian provinces,” Transactions of the Association for Computational Linguistics, vol. 12, pp. 1703–1719, 2024, doi: 10.1162/tacl_a_00726.
[9] L. Owen, V. Tripathi, A. Kumar, and B. Ahmed, “Komodo: A linguistic expedition into Indonesia’s regional languages,” arXiv, 2024, doi: 10.48550/arxiv.2403.09362.
[10] W. Q. Leong, J. G. Ngui, Y. Susanto, H. Rengarajan, K. Sarveswaran, and W.-C. Tjhi, “BHASA: A holistic Southeast Asian linguistic and cultural evaluation suite for large language models,” arXiv, 2023, doi: 10.48550/arxiv.2309.06085.
[11] Y. Susanto, A. V. Hulagadri, J. Montalan, J. G. Ngui, X. Yong, W. Leong, et al., “SEA-HELM: Southeast Asian holistic evaluation of language models,” arXiv, 2025, doi: 10.48550/arxiv.2502.14301.
[12] X. Huang, W. Zhu, H. Hu, C. He, L. Li, S. Huang, and F. Yuan, “BenchMAX: A comprehensive multilingual evaluation suite for large language models,” arXiv, 2025, doi: 10.48550/arxiv.2502.07346.
[13] T. Chang, C. Arnett, A. Eldesokey, A. Sadallah, A. Kashar, A. Daud, et al., “Global PIQA: Evaluating physical commonsense reasoning across 100+ languages and cultures,” arXiv, 2025, doi: 10.48550/arxiv.2510.24081.
[14] S. Singh, A. Romanou, C. Fourrier, D. I. Adelani, J. G. Ngui, D. Vila-Suero, et al., “Global MMLU: Understanding and addressing cultural and linguistic biases in multilingual evaluation,” arXiv, 2024, doi: 10.48550/arxiv.2412.03304.
[15] W. Xuan, R. Yang, H. Qi, Q. Zeng, Y. Xiao, Y. Xing, et al., “MMLU-ProX: A multilingual benchmark for advanced large language model evaluation,” arXiv, 2025, doi: 10.48550/arxiv.2503.10497.
[16] W. Zhang, S. M. Aljunied, C. Gao, Y. K. Chia, and L. Bing, “M3Exam: A multilingual, multimodal, multilevel benchmark for examining large language models,” arXiv, 2023, doi: 10.48550/arxiv.2306.05179.
[17] J. Myung, N. Lee, Y. Zhou, J. Jin, R. A. Putri, D. Antypas, et al., “BLEnD: A benchmark for LLMs on everyday knowledge in diverse cultures and languages,” arXiv, 2024, doi: 10.48550/arxiv.2406.09948.
[18] S. Shen, L. Logeswaran, M. Lee, H. Lee, S. Poria, and R. Mihalcea, “Understanding the capabilities and limitations of large language models for cultural commonsense,” arXiv, 2024, doi: 10.48550/arxiv.2405.04655.
[19] A. Mushtaq, I. Taj, M. Naeem, I. Ghaznavi, and J. Qadir, “WorldView-Bench: A benchmark for evaluating global cultural perspectives in large language models,” arXiv, 2025, doi: 10.48550/arxiv.2505.09595.
[20] S. Z. Ridoy, A. T. Wasi, and K. A. Tonmoy, “BengaliMoralBench: A benchmark for auditing moral reasoning in large language models within Bengali language and culture,” arXiv, 2025, doi: 10.48550/arxiv.2511.03180.
[21] T. C. Cheng, C. S. Cheng, C.-M. Lau, E. Lam, C. Y. Wong, H. O. Yu, and C. H. Chong, “HKCanto-Eval: A benchmark for evaluating Cantonese language understanding and cultural comprehension in LLMs,” arXiv, 2025, doi: 10.48550/arxiv.2503.12440.
[22] Y. A. Er, I. Kesen, G. Şahin, and A. Erdem, “Cetvel: A unified benchmark for evaluating language understanding, generation and cultural capacity of LLMs for Turkish,” arXiv, 2025, doi: 10.48550/arxiv.2508.16431.
[23] A. Maji, R. Kumar, A. Ghosh, A. Anushka, and S. Saha, “SANSKRITI: A comprehensive benchmark for evaluating language models’ knowledge of Indian culture,” arXiv, 2025, doi: 10.48550/arxiv.2506.15355.
[24] O. Nacar, S. Sibaee, S. Ahmed, S. B. Atitallah, A. Ammar, Y. AlHabashi, et al., “Towards inclusive Arabic LLMs: A culturally aligned benchmark in Arabic large language model evaluation,” Proceedings, 387–401, 2025.
[25] X. Wang, J. Yeo, J.-H. Lim, and H. Kim, “KULTURE Bench: A benchmark for assessing language model in Korean cultural context,” arXiv, 2024, doi: 10.48550/arxiv.2412.07251.
[26] H.-S. Lee, C.-C. Chang, C.-Y. Chen, and Y.-H. Hsu, “Evaluating cultural knowledge processing in large language models: A cognitive benchmarking framework integrating retrieval-augmented generation,” arXiv, 2025, doi: 10.48550/arxiv.2511.01649.
[27] T. Vo and O. Koyejo, “CURE: Cultural understanding and reasoning evaluation: A framework for “thick” culture alignment evaluation in LLMs,” arXiv, 2025, doi: 10.48550/arxiv.2511.12014.
[28] J. Zhang, S. Jiang, S. Guo, S. Chen, Y. Xiao, H. Feng, et al., “CultureScope: A dimensional lens for probing cultural understanding in LLMs,” arXiv, 2025, doi: 10.48550/arxiv.2509.16188.
[29] H. Ahmadian, T. F. Abidin, H. Riza, and K. Muchtar, “Hybrid models for emotion classification and sentiment analysis in Indonesian language,” Applied Computational Intelligence and Soft Computing, vol. 2024, Art. no. 2826773, 2024, doi: 10.1155/2024/2826773.
[30] X. Huang, T. K. Vangani, M. D. Pham, X. Zou, B. Wang, Z. Liu, and A. Aw, “MERaLiON-TextLLM: Cross-lingual understanding of large language models in Chinese, Indonesian, Malay, and Singlish,” arXiv, 2024, doi: 10.48550/arxiv.2501.08335.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Achraf El Bouazzaoui (Author)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.








Creative Commons Attribution 4.0 International License