When safety fails in local languages: evaluating culturally sensitive responses of large language models in multilingual Indonesia

Authors

  • Nadiyah Universitas Nurul Jadid Author

DOI:

https://doi.org/10.67490/ijcl.v1i2.914

Keywords:

AI safety, Buginese, culturally grounded evaluation, Indonesian languages, multilingual LLMs

Abstract

Background: In Indonesia’s digitally dense and linguistically diverse environment, LLM safety cannot be evaluated only through English or standardized Indonesian because harm is often expressed through local registers, indirectness, stigma, and culturally situated forms of advice-seeking. Objective: This study examines whether large language models produce safe, culturally sensitive, and verifiably helpful responses to safety-relevant prompts in Indonesian, Javanese, Sundanese, and Buginese. Method: Using a public-source corpus of 24 verified documents and 160 coded model responses, this study combines safety response auditing, cultural-pragmatic analysis, and grounded helpfulness assessment across domains including bullying, online gender-based violence, mental health, misinformation, and digital harm. Results: The analysis shows that safe refusal and mitigated completion dominate the response distribution, but unsafe compliance, over-refusal, evasion, and culturally inadequate redirection remain visible. Cultural-pragmatic adequacy declines in local-language prompts, where register mismatch, missed indirect distress, stigma reproduction, and unsupported cultural assumptions become more prominent. Implication: Grounded helpfulness is also uneven, as local-language responses more often contain incomplete referral pathways or unverifiable local claims in public-risk contexts. Novelty: This study contributes a multilingual Indonesian safety-evaluation framework that treats local languages not as translated inputs but as pragmatic environments where AI safety may shift, weaken, or become socially inadequate

References

[1] Z. Zhang, L. Lei, L. Wu, R. Sun, Y. Huang, C. Long, et al., “SafetyBench: Evaluating the safety of large language models with multiple choice questions,” arXiv, 2023, doi: 10.48550/arXiv.2309.07045.

[2] P. Röttger, H. R. Kirk, B. Vidgen, G. Attanasio, F. Bianchi, and D. Hovy, “XSTest: A test suite for identifying exaggerated safety behaviours in large language models,” arXiv, 2023, doi: 10.48550/arXiv.2308.01263.

[3] L. Li, B. Dong, R. Wang, X. Hu, W. Zuo, D. Lin, Y. Qiao, and J. Shao, “SALAD-Bench: A hierarchical and comprehensive safety benchmark for large language models,” arXiv, 2024, doi: 10.48550/arXiv.2402.05044.

[4] T. Xie, X. Qi, Y. Zeng, Y. Huang, U. M. Sehwag, K. Huang, et al., “SORRY-Bench: Systematically evaluating large language model safety refusal behaviors,” arXiv, 2024, doi: 10.48550/arXiv.2406.14598.

[5] Y. Chang, X. Wang, J. Wang, Y. Wu, K. Zhu, H. Chen, et al., “A survey on evaluation of large language models,” ACM Transactions on Intelligent Systems and Technology, vol. 15, no. 3, pp. 1–45, 2024, doi: 10.1145/3641289.

[6] X. Huang, W. Ruan, W. Huang, G. Jin, Y. Dong, C. Wu, et al., “A survey of safety and trustworthiness of large language models through the lens of verification and validation,” Artificial Intelligence Review, vol. 57, Art. no. 175, 2024, doi: 10.1007/s10462-024-10824-0.

[7] W. Wang, Z. Tu, C. Chen, Y. Yuan, J.-T. Huang, W. Jiao, and M. R. Lyu, “All languages matter: On the multilingual safety of large language models,” arXiv, 2023, doi: 10.48550/arXiv.2310.00905.

[8] G. I. Meadows, N. W. L. Lau, E. A. Susanto, C. Yu, and A. Paul, “LocalValueBench: A collaboratively built and extensible benchmark for evaluating localized value alignment and ethical safety in large language models,” arXiv, 2024, doi: 10.48550/arXiv.2408.01460.

[9] M. Bhutani, K. Robinson, V. Prabhakaran, S. Dave, and S. Dev, “SeeGULL Multilingual: A dataset of geo-culturally situated stereotypes,” arXiv, 2024, doi: 10.48550/arXiv.2403.05696.

[10] Y. Ashraf, Y. Wang, B. Gu, P. Nakov, and T. Baldwin, “Arabic dataset for LLM safeguard evaluation,” arXiv, 2024, doi: 10.48550/arXiv.2410.17040.

[11] J. Song, Y. Huang, Z. Zhou, and L. Li, “Multilingual blending: Large language model safety alignment evaluation with language mixture,” Findings of the Association for Computational Linguistics: NAACL 2025, pp. 3433–3449, 2025, doi: 10.18653/v1/2025.findings-naacl.191.

[12] S. Cahyawijaya, H. Lovenia, F. Koto, R. A. Putri, E. Dave, J. Lee, et al., “Cendol: Open instruction-tuned generative large language models for Indonesian languages,” arXiv, 2024, doi: 10.48550/arXiv.2404.06138.

[13] M. F. Azmi, M. D. A. Kautsar, A. Wicaksono, and F. Koto, “IndoSafety: Culturally grounded safety for LLMs in Indonesian languages,” arXiv, 2025, doi: 10.48550/arXiv.2506.02573.

[14] R. Joshi, R. Paul, K. Singla, A. Kamath, M. Evans, K. Luna, et al., “CultureGuard: Towards culturally-aware dataset and guard model for multilingual safety applications,” arXiv, 2025, doi: 10.48550/arXiv.2508.01710.

[15] P. Tasawong, J. G. Ngui, A. F. Aji, T. Cohn, and P. Limkonchotiwat, “SEA-SafeguardBench: Evaluating AI safety in SEA languages and cultures,” arXiv, 2025, doi: 10.48550/arXiv.2512.05501.

[16] L. Weidinger, J. F. J. Mellor, M. Rauh, C. Griffin, J. Uesato, P.-S. Huang, et al., “Ethical and social risks of harm from language models,” arXiv, 2021.

[17] U. Anwar, A. Saparov, J. Rando, D. Paleka, M. Turpin, P. Hase, et al., “Foundational challenges in assuring alignment and safety of large language models,” arXiv, 2024, doi: 10.48550/arXiv.2404.09932.

[18] Z. Guo, R. Jin, C. Liu, Y. Huang, D. Shi, L. Yu, et al., “Evaluating large language models: A comprehensive survey,” arXiv, 2023, doi: 10.48550/arXiv.2310.19736.

[19] X. Qi, Y. Zeng, T. Xie, P.-Y. Chen, R. Jia, P. Mittal, and P. Henderson, “Fine-tuning aligned language models compromises safety, even when users do not intend to!,” arXiv, 2023, doi: 10.48550/arXiv.2310.03693.

[20] H. Sun, Z. Zhang, J. Deng, J. Cheng, and M. Huang, “Safety assessment of Chinese large language models,” arXiv, 2023, doi: 10.48550/arXiv.2304.10436.

[21] T. Abdullahi, M. Mgonzo, M. Oduwole, P. Okewunmi, A. Owodunni, R. Singh, and C. Eickhoff, “UbuntuGuard: A culturally-grounded policy benchmark for equitable AI safety in African languages,” arXiv, 2026, doi: 10.48550/arXiv.2601.12696.

[22] D. Choi, E. Kim, J.-W. Noh, S. Seo, E. Kim, M. Oh, et al., “XL-SafetyBench: A country-grounded cross-cultural benchmark for LLM safety and cultural sensitivity,” , 2026.

[23] H. Kang, D. Lin, Z. Liao, P. Bai, X. Zeng, J. Jiang, Y. Zhu, and T. Qian, “AdaCultureSafe: Adaptive cultural safety grounded by cultural knowledge in large language models,” , 2026.

[24] H. Qiu, K.-H. Huang, R. Zheng, J. Sun, and N. Peng, “Multimodal cultural safety: Evaluation frameworks and alignment strategies,” arXiv, 2025, doi: 10.48550/arXiv.2505.14972.

[25] S. Banerjee, R. Hazra, and A. Mukherjee, “Bridging the multilingual safety divide: Efficient, culturally-aware alignment for Global South languages,” arXiv, 2026, doi: 10.48550/arXiv.2602.13867.

[26] P. Pattnayak and S. Chowdhuri, “IndicSafe: A benchmark for evaluating multilingual LLM safety in South Asia,” , 2026.

[27] T. Ukarapol, N. Chukamphaeng, K. Pipatanakul, and P. Sarapat, “ThaiSafetyBench: Assessing language model safety in Thai cultural contexts,” , 2026.

[28] J. Song, Y. Huang, Z. Zhou, and L. Li, “Multilingual blending: LLM safety alignment evaluation with language mixture,” arXiv, 2024, doi: 10.48550/arXiv.2407.07342.

[29] Z. Ning, T. Gu, J. Song, S. Hong, L. Li, H. Liu, et al., “LinguaSafe: A comprehensive multilingual safety benchmark for large language models,” arXiv, 2025, doi: 10.48550/arXiv.2508.12733.

Downloads

Published

30-06-2026

How to Cite

Nadiyah. (2026). When safety fails in local languages: evaluating culturally sensitive responses of large language models in multilingual Indonesia. Indonesian Journal of Computational Language Studies, 1(2), 93-105. https://doi.org/10.67490/ijcl.v1i2.914