Hallucinating the archipelago: detecting factual and cultural errors in retrieval-augmented generation for Indonesian knowledge
DOI:
https://doi.org/10.67490/ijcl.v1i2.915Keywords:
cultural specificity, factuality, Indonesian knowledge, retrieval-augmented generation, source provenanceAbstract
Background: Retrieval-augmented generation has been promoted as a practical response to hallucination in large language models, yet Indonesian knowledge presents a harder evidential problem because factual claims are often inseparable from regional attribution, cultural terminology, legal classification, and administrative scope. Objective: This study examines how factual and cultural errors appear in RAG-generated answers about Indonesian knowledge and how such errors can be traced across claim alignment, retrieval provenance, and cultural specificity. Method: Using a qualitative-computational design, this study evaluates 186 claim units derived from a public-source corpus of Indonesian cultural-heritage records, legal documents, statistical portals, lexical references, government explainers, and local institutional sources. Results: The findings indicate that RAG performs most reliably when generated claims reproduce explicit institutional facts, such as official entities, legal instruments, statistical categories, and cultural-heritage designations. However, unsupported and partially supported claims emerge when retrieved evidence is general, incomplete, or transformed into causal, nationalising, or culturally overextended explanations. Implication: Cultural errors are especially visible in regional scope reduction, terminological loss, ritual simplification, and temporal or administrative decontextualisation. Novelty: The novelty of this study lies in integrating claim–evidence alignment, retrieval-provenance diagnosis, and cultural-specificity coding into one framework for evaluating Indonesian RAG hallucination
References
[1] L. Huang, W. Yu, W. Zhong, Z. Feng, H. Wang, Q. Chen, et al., “A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,” ACM Transactions on Information Systems, vol. 43, pp. 1–55, 2023, doi: 10.1145/3703155.
[2] S. Tonmoy, S. Zaman, V. Jain, A. Rani, V. Rawte, A. Chadha, and A. Das, “A comprehensive survey of hallucination mitigation techniques in large language models,” arXiv, 2024, doi: 10.48550/arXiv.2401.01313.
[3] W. Zhang and J. Zhang, “Hallucination mitigation for retrieval-augmented large language models: A review,” Mathematics, vol. 13, no. 5, Art. no. 856, 2025, doi: 10.3390/math13050856.
[4] N. Thakur, L. Bonifacio, X. C. Zhang, O. Ogundepo, E. Kamalloo, D. Alfonso-Hermelo, et al., ““Knowing when you don’t know”: A multilingual relevance assessment dataset for robust retrieval-augmented generation,” , 2023.
[5] H. Tohir, N. Merlina, and M. Haris, “Utilizing retrieval-augmented generation in large language models to enhance Indonesian language NLP,” JITK (Jurnal Ilmu Pengetahuan dan Teknologi Komputer), 2024, doi: 10.33480/jitk.v10i2.5916.
[6] W. Christian, D. Adamlu, A. Yu, and D. Suhartono, “Bridging language gaps with adaptive RAG: Improving Indonesian language question answering,” arXiv, 2025, doi: 10.48550/arXiv.2510.21068.
[7] F. Muhammad, S. Y. Yakub, M. Overbeek, and F. A. T. Tobing, “RAG-fusion based framework for Indonesian hoax news detection. In 2025 IEEE 2nd International Conference on Cryptography, Informatics, and Cybersecurity (ICoCICs) (pp. 1–6),” IEEE, 2025, doi: 10.1109/ICoCICs68032.2025.11384087.
[8] I. K. R. Arthana, N. Gunantara, M. Sudarma, and M. Sukarsa, “LoRA-based fine-tuning of local LLMs for hallucination detection in Indonesian RAG systems,” International Journal of Advanced Computer Science and Applications, 2026, doi: 10.14569/ijacsa.2026.0170389.
[9] T. A. Chang, K. Tomanek, J. Hoffmann, N. Thain, E. M. Van Liemt, K. Meier-Hellstern, and L. Dixon, “Detecting hallucination and coverage errors in retrieval augmented generation for controversial topics,” arXiv, 2024, doi: 10.48550/arXiv.2403.08904.
[10] K. Wu, E. Wu, and J. Zou, “ClashEval: Quantifying the tug-of-war between an LLM’s internal prior and external evidence,” Advances in Neural Information Processing Systems, vol. 37, 2024, doi: 10.52202/079017-1053.
[11] S. S. Rahman, M. A. Islam, M. M. Alam, M. Zeba, M. A. Rahman, S. S. Chowa, M. A. K. Raiaan, and S. Azam, “Hallucination to truth: A review of fact-checking and factuality evaluation in large language models,” Artificial Intelligence Review, vol. 59, 2025, doi: 10.1007/s10462-025-11454-w.
[12] H. Orgad, M. Toker, Z. Gekhman, R. Reichart, I. Szpektor, H. Kotek, and Y. Belinkov, “LLMs know more than they show: On the intrinsic representation of LLM hallucinations,” arXiv, 2024, doi: 10.48550/arXiv.2410.02707.
[13] Z. Sun, X. Zang, K. Zheng, Y. Song, J. Xu, X. Zhang, W. Yu, and H. Li, “ReDeEP: Detecting hallucination in retrieval-augmented generation via mechanistic interpretability,” arXiv, 2024, doi: 10.48550/arXiv.2410.11414.
[14] L. Wang, “SEReDeEP: Hallucination detection in retrieval-augmented models via semantic entropy and context-parameter fusion,” arXiv, 2025, doi: 10.48550/arXiv.2505.07528.
[15] W. Su, Y. Tang, Q. Ai, C. Wang, Z. Wu, and Y. Liu, “Mitigating entity-level hallucination in large language models. In Proceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region,” ACM, 2024, doi: 10.1145/3673791.3698403.
[16] V. Rawte, S. M. Towhidul, I. Tonmoy, K. Rajbangshi, S. Nag, A. Chadha, et al., “FACTOID: FACtual enTailment fOr hallucInation Detection,” arXiv, 2024, doi: 10.48550/arXiv.2403.19113.
[17] M. Li, Z. Zhan, H. Yang, Y. Xiao, H. Zhou, J. Huang, and R. Zhang, “Benchmarking retrieval-augmented large language models in biomedical NLP: Application, robustness, and self-awareness,” Science Advances, vol. 11, 2025, doi: 10.1126/sciadv.adr1443.
[18] A. Mishra, A. Asai, V. Balachandran, Y. Wang, G. Neubig, Y. Tsvetkov, and H. Hajishirzi, “Fine-grained hallucination detection and editing for language models,” arXiv, 2024, doi: 10.48550/arXiv.2401.06855.
[19] B. Wang, S. Chern, E. Chern, and P. Liu, “Halu-J: Critique-based hallucination judge,” arXiv, 2024, doi: 10.48550/arXiv.2407.12943.
[20] I. Zimmerman, J. Tredup, E. Selfridge, and J. Bradley, “Two-tiered encoder-based hallucination detection for retrieval-augmented generation in the wild. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track (pp. 8–22),” Association for Computational Linguistics, 2024, doi: 10.18653/v1/2024.emnlp-industry.2.
[21] L. M. Amugongo, P. Mascheroni, S. Brooks, S. Doering, and J. Seidel, “Retrieval augmented generation for large language models in healthcare: A systematic review,” PLOS Digital Health, vol. 4, 2025, doi: 10.1371/journal.pdig.0000877.
[22] E. Fadeeva, A. Rubashevskii, R. Vashurin, S. Dhuliawala, A. Shelmanov, T. Baldwin, et al., “Faithfulness-aware uncertainty quantification for fact-checking the output of retrieval augmented generation,” arXiv, 2025, doi: 10.48550/arXiv.2505.21072.
[23] M. Niu, H. Li, J. Shi, H. Haddadi, and F. Mo, “Mitigating hallucinations in large language models via self-refinement-enhanced knowledge retrieval,” arXiv, 2024, doi: 10.48550/arXiv.2405.06545.
[24] H. Kang, J. Ni, and H. Yao, “Ever: Mitigating hallucination in large language models through real-time verification and rectification,” arXiv, 2023, doi: 10.48550/arXiv.2311.09114.
[25] G. Kim, J. Park, J. Kim, S. Han, K. Son, and I. Jang, “SERC: LDPC-inspired semantic error correction for retrieval-augmented generation,” , 2026.
[26] R. J. Wowor, S. Y. Yakub, and S. Indarjani, “AI-driven misinformation mitigation: Implementation of RAG system for hoax detection and cybersecurity in Indonesia. In 2025 IEEE 2nd International Conference on Cryptography, Informatics, and Cybersecurity (ICoCICs) (pp. 458–463),” IEEE, 2025, doi: 10.1109/ICoCICs68032.2025.11383967.
[27] H.-S. Lee, C.-C. Chang, C.-Y. Chen, and Y.-H. Hsu, “Evaluating cultural knowledge processing in large language models: A cognitive benchmarking framework integrating retrieval-augmented generation,” arXiv, 2025, doi: 10.48550/arXiv.2511.01649.
[28] M. Mortaheb, M. A. Khojastepour, S. Chakradhar, and S. Ulukus, “RAG-Check: Evaluating multimodal retrieval augmented generation performance,” arXiv, 2025, doi: 10.48550/arXiv.2501.03995.
[29] S. Hildebrand, C. Taylor, S. Oesch, J. Ghawaly, A. Sadovnik, R. Shivers, B. Schreiber, and K. Kurian, “FATHOMS-RAG: A framework for the assessment of thinking and observation in multimodal systems that use retrieval augmented generation,” arXiv, 2025, doi: 10.11578/dc.20260227.1.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Radna Tulus Wibisono (Author)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.








Creative Commons Attribution 4.0 International License