Small models, many languages: parameter-efficient adaptation of multilingual language models for low-resource Indonesian languages

Authors

  • Ainul Hadziqi Universitas Sarjanawiyata Tamansiswa Author

DOI:

https://doi.org/10.67490/ijcl.v1i2.916

Keywords:

corpus adequacy, Indonesian local languages, low-resource NLP, multilingual language models, parameter-efficient fine-tuning

Abstract

Background: Multilingual language models have expanded computational access to many languages, yet Indonesian local languages remain unevenly represented in corpora, benchmarks, and adaptation pipelines despite Indonesia’s exceptional linguistic diversity. Objective: This study examines how parameter-efficient adaptation can support small multilingual language models for low-resource Indonesian languages by treating corpus adequacy, source provenance, language coverage, and evaluation readiness as central methodological conditions. Method: Using a corpus-level research design, this study maps 20 public source-level items consisting of task datasets, benchmark resources, documentation, registry infrastructure, supplementary corpus portals, and model-context collections for Indonesian and selected local languages. Results: The findings show that task-ready datasets form the strongest part of the corpus frame, while benchmarking, reproducibility, and auxiliary corpus layers remain less evenly distributed. Language coverage is stratified, with Javanese and Sundanese occupying stronger cross-task positions, while Acehnese, Ngaju, Bima, Toba Batak, and Ambonese Malay are better treated as focused transfer-stress cases. Implication: These patterns indicate that parameter-efficient fine-tuning should not be interpreted through model efficiency alone, because adaptation gains remain bounded by corpus distribution and language-specific evidence. Novelty: The novelty of this study lies in repositioning PEFT for Indonesian local languages as a corpus-dependent multilingual adaptation problem rather than a purely architectural optimization problem

References

[1] S. Cahyawijaya, G. I. Winata, B. Wilie, K. Vincentio, X. Li, A. Kuncoro, et al., “IndoNLG: Benchmark and resources for evaluating Indonesian natural language generation,” Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 8875–8898, 2021, doi: 10.18653/v1/2021.emnlp-main.699.

[2] J. A. Lopo and R. Tanone, “Constructing and expanding low-resource and underrepresented parallel datasets for Indonesian local languages,” ArXiv, abs/2404.01009, 2024, doi: 10.48550/arxiv.2404.01009.

[3] W. Wongso, A. Joyoadikusumo, B. S. Buana, and D. Suhartono, “Many-to-many multilingual translation model for languages of Indonesia,” IEEE Access, vol. 11, pp. 91385–91397, 2023, doi: 10.1109/access.2023.3308818.

[4] W. Tan and K. Zhu, “NusaMT-7B: Machine translation for low-resource Indonesian languages with large language models,” ArXiv, abs/2410.07830, 2024, doi: 10.48550/arxiv.2410.07830.

[5] S. Cahyawijaya, H. Lovenia, F. Koto, R. A. Putri, E. Dave, J. Lee, et al., “Cendol: Open instruction-tuned generative large language models for Indonesian languages,” ArXiv, abs/2404.06138, 14899–14914, 2024, doi: 10.48550/arxiv.2404.06138.

[6] “Compacter: Efficient low-rank hypercomplex adapter layers,” , pp. 1022–1035, 2021.

[7] L. Xu, H. Xie, S. Qin, X. Tao, and F. Wang, “Parameter-efficient fine-tuning methods for pretrained language models: A critical review and assessment,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, pp. 6107–6126, 2023, doi: 10.1109/tpami.2026.3657354.

[8] D. Aggarwal, A. Sathe, I. Watts, and S. Sitaram, “MAPLE: Multilingual evaluation of parameter efficient finetuning of large language models,” ArXiv, abs/2401.07598, 2024, doi: 10.48550/arxiv.2401.07598.

[9] T. Su, X. Peng, S. Thillainathan, D. Guzmán, S. Ranathunga, and E.-S. A. Lee, “Unlocking parameter-efficient fine-tuning for low-resource language translation,” ArXiv, abs/2404.04212, 4217–4225, 2024, doi: 10.48550/arxiv.2404.04212.

[10] P.-H. Chen and Y.-N. Chen, “Efficient unseen language adaptation for multilingual pre-trained language models,” Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 18983–18994, 2024, doi: 10.18653/v1/2024.emnlp-main.1057.

[11] Z. Csaki, P. Pawakapan, U. Thakker, and Q. Xu, “Efficiently adapting pretrained language models to new languages,” ArXiv, abs/2311.05741, 2023, doi: 10.48550/arxiv.2311.05741.

[12] C. M. Downey, T. Blevins, D. Serai, D. Parikh, and S. Steinert-Threlkeld, “Targeted multilingual adaptation for low-resource language families,” ArXiv, abs/2405.12413, 15647–15663, 2024, doi: 10.48550/arxiv.2405.12413.

[13] D. Gurgurov, M. Hartmann, and S. Ostermann, “Adapting multilingual LLMs to low-resource languages with knowledge graphs via adapters,” Proceedings of the Workshop on Knowledge Augmented Methods for Natural Language Processing, 1–7, 2024, doi: 10.18653/v1/2024.kallm-1.7.

[14] D. Gurgurov, I. Vykopal, J. Van Genabith, and S. Ostermann, “Small models, big impact: Efficient corpus and graph-based adaptation of small multilingual language models for low-resource languages,” ArXiv, abs/2502.10140, 355–395, 2025, doi: 10.48550/arxiv.2502.10140.

[15] J. O. Alabi, D. I. Adelani, M. Mosbach, and D. Klakow, “Adapting pre-trained language models to African languages via multilingual adaptive fine-tuning,” Proceedings of the 29th International Conference on Computational Linguistics, 4336–4349, 2022.

[16] G. Pu, A. Jain, J. Yin, and R. Kaplan, “Empirical analysis of the strengths and weaknesses of PEFT techniques for LLMs,” ArXiv, abs/2304.14999, 2023, doi: 10.48550/arxiv.2304.14999.

[17] N. J. Prottasha, U. R. Chowdhury, S. Mohanto, T. Nuzhat, A. A. Sami, M. S. Ali, et al., “PEFT A2Z: Parameter-efficient fine-tuning survey for large language and vision models,” ArXiv, abs/2504.14117, 2025, doi: 10.48550/arxiv.2504.14117.

[18] O. Khade, S. Jagdale, A. Phaltankar, G. Takalikar, and R. Joshi, “Challenges in adapting multilingual LLMs to low-resource languages using LoRA PEFT tuning,” ArXiv, abs/2411.18571, 2024, doi: 10.48550/arxiv.2411.18571.

[19] R. Joshi, K. Singla, A. Kamath, R. Kalani, R. Paul, U. Vaidya, et al., “Adapting multilingual LLMs to low-resource languages using continued pre-training and synthetic corpus,” ArXiv, abs/2410.14815, 50–57, 2024, doi: 10.48550/arxiv.2410.14815.

[20] S. Purkayastha, S. Ruder, J. Pfeiffer, I. Gurevych, and I. Vulić, “Romanization-based large-scale adaptation of multilingual language models,” Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7996–8005, 2023, doi: 10.48550/arxiv.2304.08865.

[21] F. Remy, P. Delobelle, H. Avetisyan, A. Khabibullina, M. De Lhoneux, and T. Demeester, “Trans-tokenization and cross-lingual vocabulary transfers: Language adaptation of LLMs for low-resource NLP,” ArXiv, abs/2408.04303, 2024, doi: 10.48550/arxiv.2408.04303.

[22] D. Isnaeni, N. Afra, I. Budi, and M. T. Uliniansyah, “Fine-tuning pretrained language models for extremely low-resource machine translation: The case of Makassarese-Indonesian,” 2025 International Conference on Computer, Control, Informatics and Its Applications (IC3INA), 208–213, 2025, doi: 10.1109/ic3ina68387.2025.11325367.

[23] J. Sha, M. Zhu, C. Feng, and Y. Shang, “VEEF-Multi-LLM: Effective vocabulary expansion and parameter efficient finetuning towards multilingual large language models,” 7963–7981, 2025.

[24] H. Kang, J. Ni, and H. Yao, “Ever: Mitigating hallucination in large language models through real-time verification and rectification,” arXiv, 2023, doi: 10.48550/arXiv.2311.09114.

[25] G. Kim, J. Park, J. Kim, S. Han, K. Son, and I. Jang, “SERC: LDPC-inspired semantic error correction for retrieval-augmented generation,” , 2026.

Downloads

Published

30-06-2026

How to Cite

Ainul Hadziqi. (2026). Small models, many languages: parameter-efficient adaptation of multilingual language models for low-resource Indonesian languages. Indonesian Journal of Computational Language Studies, 1(2), 121-135. https://doi.org/10.67490/ijcl.v1i2.916

Similar Articles

1-10 of 11

You may also start an advanced similarity search for this article.