Do large language models reason consistently across Indonesian languages? a multilingual evaluation of Indonesian, Javanese, Sundanese, and Buginese
DOI:
https://doi.org/10.67490/ijcl.v1i2.912Keywords:
Buginese, cross-linguistic consistency, Indonesian languages, large language models, multilingual reasoningAbstract
Background: Multilingual large language model evaluation increasingly recognises Indonesian as a significant benchmark language, yet Indonesia’s wider linguistic ecology requires assessment across local languages whose digital representation remains uneven. Objective: This study examines whether large language models reason consistently across Indonesian, Javanese, Sundanese, and Buginese when semantically aligned reasoning tasks are presented in each language. Method: Using a controlled multilingual evaluation design, this study compares model outputs through three analytic dimensions: final-answer consistency, explanation coherence and faithfulness, and language-linked error patterns. Results: Findings show that answer stability is not evenly preserved across language pairs, with stronger alignment in Indonesian–Javanese and Indonesian–Sundanese comparisons than in Indonesian–Buginese comparison. Explanation analysis further indicates that apparently correct answers may be accompanied by compressed, partially faithful, or weakly grounded reasoning. Implication: Error analysis reveals that inconsistency emerges through multiple pathways, including lexical-semantic drift, register mismatch, translation-related distortion, cultural inferential misreading, and language fallback. Novelty: The novelty of this study lies in shifting Indonesian multilingual LLM evaluation from isolated benchmark accuracy to cross-linguistic reasoning consistency, positioning local languages as central analytic sites for assessing reliability, equity, and epistemic accountability in multilingual artificial intelligence
References
[1] F. Koto, N. Aisyah, H. Li, and T. Baldwin, “Large language models only pass primary school exams in Indonesia: A comprehensive test on IndoMMLU,” Findings of the Association for Computational Linguistics: EMNLP 2023, 12359–12374, 2023, doi: 10.48550/arXiv.2310.04928.
[2] H. A. Wibowo, E. H. Fuadi, M. N. Nityasya, R. E. Prasojo, and A. F. Aji, “COPAL-ID: Indonesian language reasoning with local culture and nuances,” Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1404–1422, 2023, doi: 10.48550/arXiv.2311.01012.
[3] F. Koto, R. Mahendra, N. Aisyah, and T. Baldwin, “IndoCulture: Exploring geographically influenced cultural commonsense reasoning across eleven Indonesian provinces,” Transactions of the Association for Computational Linguistics, vol. 12, pp. 1703–1719, 2024, doi: 10.1162/tacl_a_00726.
[4] W. Q. Leong, J. G. Ngui, Y. Susanto, H. Rengarajan, K. Sarveswaran, and W.-C. Tjhi, “BHASA: A holistic Southeast Asian linguistic and cultural evaluation suite for large language models,” arXiv, 2023, doi: 10.48550/arXiv.2309.06085.
[5] Y. Susanto, A. V. Hulagadri, J. Montalan, J. G. Ngui, X. Yong, W. Leong, et al., “SEA-HELM: Southeast Asian holistic evaluation of language models,” arXiv, 2025, doi: 10.48550/arXiv.2502.14301.
[6] S. Cahyawijaya, H. Lovenia, F. Koto, R. A. Putri, E. Dave, J. Lee, et al., “Cendol: Open instruction-tuned generative large language models for Indonesian languages,” Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 14899–14914, 2024, doi: 10.48550/arXiv.2404.06138.
[7] X.-P. Nguyen, W. Zhang, X. Li, M. Aljunied, Q. Tan, L. Cheng, et al., “SeaLLMs—Large language models for Southeast Asia,” arXiv, 2023, doi: 10.48550/arXiv.2312.00738.
[8] A. F. Aji and T. Cohn, “LoraxBench: A multitask, multilingual benchmark suite for 20 Indonesian languages,” arXiv, 2025, doi: 10.48550/arXiv.2508.12459.
[9] H. Huang, T. Tang, D. Zhang, W. X. Zhao, T. Song, Y. Xia, and F. Wei, “Not all languages are created equal in LLMs: Improving multilingual capability by cross-lingual-thought prompting,” Findings of the Association for Computational Linguistics: EMNLP 2023, 12365–12394, 2023, doi: 10.48550/arXiv.2305.07004.
[10] X. Xing, Z. He, H. Xu, X. Wang, R. Wang, and Y. Hong, “Evaluating knowledge-based cross-lingual inconsistency in large language models,” arXiv, 2024, doi: 10.48550/arXiv.2407.01358.
[11] A. Gupta, M. Mehta, Z. Xu, and V. Srikumar, “Found in translation: Measuring multilingual LLM consistency as simple as translate then evaluate,” Proceedings of the 2025 Annual Conference of the North American Chapter of the Association for Computational Linguistics, 3477–3496, 2025, doi: 10.48550/arXiv.2505.21999.
[12] L. Qin, Q. Chen, Y. Zhou, Z. Chen, Y. Li, L. Liao, et al., “A survey of multilingual large language models,” Patterns, vol. 6, 2025, doi: 10.1016/j.patter.2024.101118.
[13] Y. Xu, L. Hu, J. Zhao, Z. Qiu, Y. Ye, and H. Gu, “A survey on multilingual large language models: Corpora, alignment, and bias,” Frontiers of Computer Science, vol. 19, 2024, doi: 10.1007/s11704-024-40579-4.
[14] W. Zheng, X. Huang, Z. Liu, T. K. Vangani, B. Zou, X. Tao, et al., “AdaMCoT: Rethinking cross-lingual factual reasoning through adaptive multilingual chain-of-thought,” Proceedings of the AAAI Conference on Artificial Intelligence, 33863–33871, 2025, doi: 10.1609/aaai.v40i40.40678.
[15] R. Zhao, Y. Liu, H. Schütze, and M. A. Hedderich, “A comprehensive evaluation of multilingual chain-of-thought reasoning: Performance, consistency, and faithfulness across languages,” arXiv, 2025, doi: 10.48550/arXiv.2510.09555.
[16] A. Agrawal, A. Dang, S. B. Nezhad, R. Pokharel, and R. Scheinberg, “Evaluating multilingual long-context models for retrieval and reasoning,” Proceedings of the First Workshop on Multilingual Retrieval-Augmented Generation, 3477–3496, 2024, doi: 10.18653/v1/2024.mrl-1.18.
[17] R. S. J. Xuan, J. Huseynov, and Y. Zhang, “Uncovering cross-linguistic disparities in LLMs using sparse autoencoders,” arXiv, 2025, doi: 10.48550/arXiv.2507.18918.
[18] W. Han, Y. Zhang, Z. Chen, B. Li, H. Lin, B. Zhang, et al., “MuBench: Assessment of multilingual capabilities of large language models across 61 languages,” arXiv, 2025, doi: 10.48550/arXiv.2506.19468.
[19] F. Faisal and A. Anastasopoulos, “An efficient approach for studying cross-lingual transfer in multilingual language models,” arXiv, 2024, doi: 10.48550/arXiv.2403.20088.
[20] J. Lee, S. Hong, H. Moon, and H.-J. Lim, “Cross-lingual optimization for language transfer in large language models,” arXiv, 2025, doi: 10.48550/arXiv.2505.14297.
[21] K. Ravisankar, H. Han, and M. Carpuat, “Can you map it to English? The role of cross-lingual alignment in the multilingual performance of LLMs,” Proceedings of the 2026 Conference of the European Chapter of the Association for Computational Linguistics, 4854–4872, 2025, doi: 10.18653/v1/2026.eacl-long.225.
[22] S. Rajaee and C. Monz, “Analyzing the evaluation of cross-lingual knowledge transfer in multilingual language models,” arXiv, 2024, doi: 10.48550/arXiv.2402.02099.
[23] W. Zhang, H. P. Chan, Y. Zhao, M. Aljunied, J. Wang, C. Liu, et al., “SeaLLMs 3: Open foundation and chat multilingual large language models for Southeast Asian languages,” Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 96–105, 2024, doi: 10.48550/arXiv.2407.19672.
[24] L. Dou, Q. Liu, G. Zeng, J. Guo, J. Zhou, W. Lu, and M. Lin, “Sailor: Open language models for South-East Asia,” Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 424–435, 2024, doi: 10.48550/arXiv.2404.03608.
[25] F. Adilazuarda, S. Cahyawijaya, G. I. Winata, P. Fung, and A. Purwarianti, “IndoRobusta: Towards robustness against diverse code-mixed Indonesian local languages,” arXiv, 2023, doi: 10.48550/arXiv.2311.12405.
[26] W. Wongso, D. S. Setiawan, S. Limcorn, and A. Joyoadikusumo, “NusaBERT: Teaching IndoBERT to be multilingual and multicultural,” arXiv, 2024, doi: 10.48550/arXiv.2403.01817.
[27] W. Tan and K. Zhu, “NusaMT-7B: Machine translation for low-resource Indonesian languages with large language models,” arXiv, 2024, doi: 10.48550/arXiv.2410.07830.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Fitriya Dessi Wulandari (Author)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.








Creative Commons Attribution 4.0 International License