Evolution and Adaptation of Large Language Models for Bahasa Indonesia
DOI:
https://doi.org/10.38035/dit.v4i1.3623Keywords:
Language Adaptation, Bahasa Indonesia, Large Language Models, Vocabulary Expansion, Continual Pre-Training, Cultural AlignmentAbstract
This study evaluates the systematic evolution and computational adaptation of pre-trained language models and Large Language Models (LLMs) for Bahasa Indonesia and its low-resource regional dialects. Initially centered on bidirectional encoder-based representations like IndoBERT, the regional natural language processing (NLP) field has transitioned toward generative sequence-to-sequence structures and massive decoder-only architectures. This paper investigates the engineering methodologies of cross-lingual vocabulary adaptation, parameter initialization heuristics, and language-adaptive pre-training strategies designed to address text overfragmentation, representational misalignment, and tokenization cost inefficiencies. Through extensive structural benchmarks, this analysis compares discriminative and generative performances across tasks including sentiment classification, extractive question answering, text style normalization, domain-specific retrieval-augmented pipelines, and entity linking. While localized generative models such as Komodo, Sailor, and the SEA-LION suite improve contextual reasoning, colloquial style transfers, and regional dialect preservation, they remain susceptible to architectural anomalies like template leakage and entity hallucination. This study provides foundational benchmarks and methodological frameworks for adapting massive language models to morphologically rich, culturally diverse, and low-resource linguistic environments.
References
Cahyawijaya, S., Winata, G. I., Wilie, B., Vincentio, K., Li, X., Kuncoro, A., Ruder, S., Lim, Z. Y., Bahar, S., Khodra, M. L., Purwarianti, A., & Fung, P. (2021). IndoNLG: Benchmark and Resources for Evaluating Indonesian Natural Language Generation. EMNLP 2021 - 2021 Conference on Empirical Methods in Natural Language Processing, Proceedings, 8875–8898. https://doi.org/10.18653/v1/2021.emnlp-main.699
Chakrabarty, A., Chaturvedi, A., & Garain, U. (2020). NeuMorph. ACM Transactions on Asian and Low-Resource Language Information Processing, 19(1), 1–19. https://doi.org/10.1145/3342354
Chizhov, P., Arnett, C., Korotkova, E., & Yamshchikov, I. P. (2024). BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training. EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference, 16587–16604. https://doi.org/10.18653/v1/2024.emnlp-main.925
Di Marco, M., & Fraser, A. (2024). Subword Segmentation in LLMs: Looking at Inflection and Consistency. EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference, 12050–12060. https://doi.org/10.18653/v1/2024.emnlp-main.672
Ding, W., Wang, W., Kwok, S. H. D., Liu, M., Fang, T., Bai, J., Liu, X., Yu, C., Li, Z., Luo, C., Yin, Q., Yin, B., He, J., & Song, Y. (2024). IntentionQA: A Benchmark for Evaluating Purchase Intention Comprehension Abilities of Language Models in E-commerce. Findings of the Association for Computational Linguistics: EMNLP 2024, 2247–2266. https://doi.org/10.18653/v1/2024.findings-emnlp.123
Fridkin, S., & Bendersky, M. (2026). Interpretable Machine Learning: A Comprehensive Review of Foundations, Methods, and the Path Forward. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 16(1). https://doi.org/10.1002/widm.70075
Fujii, K., Nakamura, T., Loem, M., Iida, H., Ohi, M., Hattori, K., Shota, H., Mizuki, S., Yokota, R., & Okazaki, N. (2024). Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities. http://arxiv.org/abs/2404.17790
Hasmawati, & Ade Romadhony. (2023). Similar Questions Identification on Indonesian Language Subject Using Machine Learning. Jurnal Nasional Pendidikan Teknik Informatika (JANAPATI), 12(2), 196–202. https://doi.org/10.23887/janapati.v12i2.62582
Jap, B. A. J., & Arumsari, C. (2017). Adaptation of the Token Test in Standard Indonesian. Makara Human Behavior Studies in Asia, 21(1), 44. https://doi.org/10.7454/mssh.v21i1.3499
Jin, B. (2024). Multilingual Neural Machine Translation with Integrated Language Adapters. Proceedings of the 2024 International Conference on Generative Artificial Intelligence and Information Security, 29–35. https://doi.org/10.1145/3665348.3665355
Jo, E., Cho, E., Lee, Y., Song, S., & Joo, H. J. (2025). Domain and Language adaptive pre-training of BERT models for Korean-English bilingual clinical text analysis. BMC Medical Informatics and Decision Making, 25(1). https://doi.org/10.1186/s12911-025-03262-7
Kartika, B. V., Alfredo, M. J., & Kusuma, G. P. (2023). Fine-Tuned IndoBERT based model and data augmentation for indonesian language paraphrase identification. Revue d’Intelligence Artificielle, 37(3), 733–743. https://doi.org/10.18280/ria.370322
Lee, K. S. J. W., Qorib, M. R., Soegeng, A. I., & Ng, H. T. (2026). Equity with Efficiency: An Empirical Study of Tokenizers for Multilingual Large Language Models. http://arxiv.org/abs/2606.15044
Linder, M. R. (2026). Vocabulary Expansion of Large Language Models via Kullback-Leibler-Based Self-Distillation. http://arxiv.org/abs/2508.15807
McGiff, J., & Nikolov, N. S. (2025). Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review. http://arxiv.org/abs/2505.04531
Mu, L., Wang, X., Ni, L., Li, Y., Wu, Z., Jin, P., & Zhang, Y. (2025). DenseLoRA: Dense Low-Rank Adaptation of Large Language Models. Proceedings of the Annual Meeting of the Association for Computational Linguistics, 1, 10198–10211. https://doi.org/10.18653/v1/2025.acl-long.503
Owen, L., Tripathi, V., Kumar, A., & Ahmed, B. (2024). Komodo: A Linguistic Expedition into Indonesia’s Regional Languages. http://arxiv.org/abs/2403.09362
Petrov, A., La Malfa, E., Torr, P. H. S., & Bibi, A. (2023). Language Model Tokenizers Introduce Unfairness Between Languages. http://arxiv.org/abs/2305.15425
Provilkov, I., Emelianenko, D., & Voita, E. (2020). BPE-dropout: Simple and effective subword regularization. Proceedings of the Annual Meeting of the Association for Computational Linguistics, 1, 1882–1892. https://doi.org/10.18653/v1/2020.acl-main.170
Ridwan, M. (2018). National and Official Language: The Long Journey of Indonesian Language. Budapest International Research and Critics Institute (BIRCI-Journal) : Humanities and Social Sciences, 1(2), 72–78. https://doi.org/10.33258/birci.v1i2.14
Subhan Mahendrasyah, M., & Hariguna, T. (2024). Analisis Sentimen Pengguna Aplikasi Bukalapak di Platform Playstore Menggunakan Metode Naïve Bayes. Building of Informatics, Technology and Science (BITS), 6(2), 733–745. https://doi.org/10.47065/bits.v6i2.5528
Xu, Y., & Che, W. (2026). Evolving LLMs from Next-Token Prediction to Multi-Token Prediction via Self-Distillation. Electronics (Switzerland), 15(7). https://doi.org/10.3390/electronics15071533
Yamaguchi, A., Villavicencio, A., & Aletras, N. (2026). How Can We Effectively Expand the Vocabulary of LLMs with 0.01GB of Target Language Text? Computational Linguistics, 52(1), 295–330. https://doi.org/10.1162/COLI.a.581
Zhang, L., Lou, Z., Ying, Y., Yang, C., & Zhou, H. (2025). Efficient Fine-Tuning of Large Language Models via a Low-Rank Gradient Estimator. Applied Sciences (Switzerland), 15(1), 1–16. https://doi.org/10.3390/app15010082
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Allan Desi Alexander, Siti Setiawati

This work is licensed under a Creative Commons Attribution 4.0 International License.
Copyright :
Authors who publish their manuscripts in this journal agree to the following conditions:
- Copyright in each article belongs to the author.
- The author acknowledges that the DIT has the right to be the first to publish under a Creative Commons Attribution 4.0 International license (Attribution 4.0 International CC BY 4.0).
- Authors can submit articles separately, arrange the non-exclusive distribution of manuscripts that have been published in this journal to other versions (for example, sent to the author's institutional repository, publication in a book, etc.), by acknowledging that the manuscript has been published for the first time at DIT.




















