Indonesian named entity recognition using BiLSTM-CRF with domain-specific word embeddings
DOI:
https://doi.org/10.67490/ijcl.v1i1.548Keywords:
BiLSTM-CRF, domain-specific embeddings, Indonesian language, named entity recognition, sequence labelingAbstract
Background: The rapid growth of Indonesian digital news content increases the demand for accurate Named Entity Recognition (NER), yet linguistic complexity and limited annotated data remain key challenges. Objective: This study evaluates the effectiveness of a BiLSTM-CRF model enhanced with domain-specific word embeddings for Indonesian NER across person, location, and organization entities. Method: A supervised sequence-labeling approach was applied using annotated news data, with embeddings trained on large-scale political and economic corpora and evaluated via precision, recall, and F1-score. Results: The model shows stable performance for person and location entities, while domain-specific embeddings improve all categories, especially organization entities with high variation; remaining errors relate to boundary detection and semantic ambiguity. Implication: These findings highlight the importance of domain-adaptive representations for improving NER systems in low-resource languages and support more reliable information extraction in Indonesian digital media. Novelty: This study demonstrates that embedding-level domain adaptation significantly enhances Indonesian NER without increasing model complexity, while clarifying distinctions between corpus resources and annotated data for better reproducibility.
References
[1] Y. Qiu, L. Dong, W. Zhang, H. Xing, and J. Huang, “A diffusion enhanced CRF and BiLSTM framework for accurate entity recognition,” Scientific Reports, vol. 15, Jun. 2025, doi: 10.1038/s41598-025-04036-x.
[2] E. Yulianti, N. Bhary, J. Abdurrohman, F. W. Dwitilas, E. Q. Nuranti, and H. S. Husin, “Named entity recognition on Indonesian legal documents: a dataset and study using transformer-based models,” International Journal of Electrical and Computer Engineering (IJECE), Oct. 2024, doi: 10.11591/ijece.v14i5.pp5489-5501.
[3] E. Subowo, I. Bukhori, and Warto, “Corpus Development and NER Model for Identification of Legal Entities (Articles, Laws, and Sanctions) in Corruption Court Decisions in Indonesia,” Transactions on Informatics and Data Science, Jun. 2025, doi: 10.24090/tids.v2i1.13592.
[4] I. M. K. Karo, S. Dewi, and A. S. Nasution, “Named Entity Recognition on Indonesian Online News Based on Bidirectional LSTM-CRF,” 2025 4th International Conference on Electronics Representation and Algorithm (ICERA), pp. 251–256, Jun. 2025, doi: 10.1109/icera66156.2025.11087367.
[5] K. L. Thant, K. Nongpong, Y. K. Thu, T. Aung, K. Wai, and T. Oo, “myNER: Contextualized Burmese Named Entity Recognition with Bidirectional LSTM and fastText Embeddings via Joint Training with POS Tagging,” 2025 IEEE International Conference on Cybernetics and Innovations (ICCI), pp. 1–6, Apr. 2025, doi: 10.1109/icci64209.2025.10987237.
[6] P. Chen, M. Zhang, X. Yu, and S. Li, “Named entity recognition of Chinese electronic medical records based on a hybrid neural network and medical MC-BERT,” BMC Medical Informatics and Decision Making, vol. 22, Dec. 2022, doi: 10.1186/s12911-022-02059-2.
[7] S. Arslan, “Application of BiLSTM-CRF model with different embeddings for product name extraction in unstructured Turkish text,” Neural Computing and Applications, vol. 36, pp. 8371–8382, Feb. 2024, doi: 10.1007/s00521-024-09532-1.
[8] R. Anam et al., “A deep learning approach for Named Entity Recognition in Urdu language,” PLOS ONE, vol. 19, Mar. 2024, doi: 10.1371/journal.pone.0300725.
[9] Z. Zainuddin, -. Mudassir, and Z. Tahir, “Entity Extraction in Indonesian Online News Using Named Entity Recognition (NER) with Hybrid Method Transformer, Word2Vec, Attention and Bi-LSTM,” JOIV : International Journal on Informatics Visualization, May 2025, doi: 10.62527/joiv.9.3.2902.
[10] A. Trewartha et al., “Quantifying the advantage of domain-specific pre-training on named entity recognition tasks in materials science,” Patterns, vol. 3, Apr. 2022, doi: 10.1016/j.patter.2022.100488.
[11] G. Shidik et al., “Indonesian Disaster Named Entity Recognition from Multi Source Information using Bidirectional LSTM (BiLSTM),” Journal of Open Innovation: Technology, Market, and Complexity, Aug. 2024, doi: 10.1016/j.joitmc.2024.100358.
[12] E. Dave and A. Chowanda, “IPerFEX-2023: Indonesian personal financial entity extraction using indoBERT-BiGRU-CRF model,” Journal of Big Data, vol. 11, Sep. 2024, doi: 10.1186/s40537-024-00987-6.
[13] R. Yan, X. Jiang, and D. Dang, “Named Entity Recognition by Using XLNet-BiLSTM-CRF,” Neural Processing Letters, vol. 53, pp. 3339–3356, Jun. 2021, doi: 10.1007/s11063-021-10547-1.
[14] A. Fawaid, H. Baharun, M. Hamzah, Rohimah, I. Munawwaroh, and D. F. Putri, “AI-based career management to improve the quality of decision making in higher education,” in 2025 IEEE Integrated STEM Education Conference (ISEC), IEEE, Mar. 2025, pp. 1–8. doi: 10.1109/isec64801.2025.11147274.
[15] N. Hussain, A. Qasim, G. Mehak, O. Kolesnikova, A. Gelbukh, and G. Sidorov, “ORUD-Detect: A Comprehensive Approach to Offensive Language Detection in Roman Urdu Using Hybrid Machine Learning-Deep Learning Models with Embedding Techniques,” Inf., vol. 16, p. 139, Feb. 2025, doi: 10.3390/info16020139.
[16] K. Li, H. Zha, Y. Su, and X. Yan, “Concept Mining via Embedding,” presented at the Proceedings - IEEE International Conference on Data Mining, ICDM, 2018, pp. 267–276. doi: 10.1109/ICDM.2018.00042.
[17] J. Neumann, R. Lange, Y. Susanti, and M. Färber, “Optimizing Small Transformer-Based Language Models for Multi-Label Sentiment Analysis in Short Texts,” ArXiv, vol. abs/2509.04982, Sep. 2025, doi: 10.48550/arxiv.2509.04982.
[18] N. Yogeesh et al., “Modeling Lexical Ambiguity in English Literature Using Fuzzy Logic and Equations,” Applied Mathematics and Information Sciences, vol. 19, no. 4, pp. 873–889, 2025, doi: 10.18576/amis/190413.
[19] S. Zad, M. Heidari, P. Hajibabaee, and M. Malekzadeh, “A Survey of Deep Learning Methods on Semantic Similarity and Sentence Modeling,” presented at the 2021 IEEE 12th Annual Information Technology, Electronics and Mobile Communication Conference, IEMCON 2021, 2021, pp. 466–472. doi: 10.1109/IEMCON53756.2021.9623078.
[20] R. Adyanthaya and R. S. P, “YenLP_CS@DravidianLangTech 2025: Sentiment Analysis on Code-Mixed Tamil-Tulu Data Using Machine Learning and Deep Learning Models,” Proceedings of the Fifth Workshop on Speech, Vision, and Language Technologies for Dravidian Languages, Jan. 2025, doi: 10.18653/v1/2025.dravidianlangtech-1.50.
[21] C. Anderson, M. Nguyen, and R. Coto-Solano, “Unsupervised, Semi-Supervised and LLM-Based Morphological Segmentation for Bribri,” Proceedings of the Fifth Workshop on NLP for Indigenous Languages of the Americas (AmericasNLP), Jan. 2025, doi: 10.18653/v1/2025.americasnlp-1.7.
[22] A. Petropoulos and V. Siakoulis, “Can central bank speeches predict financial market turbulence? Evidence from an adaptive NLP sentiment index analysis using XGBoost machine learning technique,” Central Bank Review, Dec. 2021, doi: 10.1016/j.cbrev.2021.12.002.
[23] X. Fang and J. Yan, “Named entity recognition for U.S. export control based on the ERNIE-BiLSTM-Attention-CRF model,” vol. 13646, pp. 136460–136460, May 2025, doi: 10.1117/12.3056127.
[24] A. Fawaid, R. Assyabani, I. Abdullah, C. Muali, M. S. Itqan, and S. Islam, “Human Intelligence and Algorithmic Precision: An Experimental Study of Indonesian Translation Pedagogy in Higher Education,” Asian Journal of University Education, vol. 21, no. 3, pp. 779–792, 2025, doi: 10.24191/ajue.v21i3.53.
[25] A. Ustun and B. Can, “Incorporating word embeddings in unsupervised morphological segmentation,” Natural Language Engineering, vol. 27, pp. 609–629, Jul. 2020, doi: 10.1017/s1351324920000406.
[26] R. Chauhan, A. Mishra, M. Jena, E. Yafi, and M. F. Zuhairi, “Mental Health Diagnostics Using Machine Learning: A Clustering and Predictive Modeling Approach,” presented at the Proceedings of the 2025 19th International Conference on Ubiquitous Information Management and Communication, IMCOM 2025, 2025. doi: 10.1109/IMCOM64595.2025.10857500.
[27] V. Mayik, N. Lotoshynska, L. Mayik, K. Bazylyuk, and M. Khariv, “Mathematical Modeling of the Braille Font Formation for the Information and Communication Environment Creation for Blind People,” presented at the CEUR Workshop Proceedings, 2023, pp. 248–254. [Online]. Available: https://ceur-ws.org/Vol-3641/short5.pdf
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Joko Santoso (Author)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.








Creative Commons Attribution 4.0 International License