Three-Category News Text Classification Using LSTM on a Local Government Portal
DOI:
https://doi.org/10.69533/informatech.volume3number1.582Keywords:
Text Classification, Long Short-Term Memory, Government PortalAbstract
This research develops an automated three-category news classification system for the official news portal of the Regional Revenue Agency (Bapenda) of Surabaya City, using a Long Short-Term Memory (LSTM) deep learning model to reduce reliance on manual, inconsistent categorization. The training corpus combined original institutional data, web-scraped articles from other regional revenue agencies, and GPT-generated synthetic text, all validated by domain experts, then processed through a structured preprocessing pipeline. On an unseen test set of 513 authentic institutional articles, the model achieved an overall accuracy of 77.78% and a weighted F1-score of 0.78, driven mainly by strong performance on the majority "Important Information" class (F1-score 0.87). However, the macro-averaged F1-score was only 0.4434, reflecting substantially weaker performance on minority classes, particularly "Information Technology" (F1-score 0.069). The model was deployed as a human-in-the-loop recommendation feature within the agency's existing dashboard.
Downloads
References
I. Mutambik et al., “Usability of the G7 Open Government Data Portals and Lessons Learned,” Sustainability, vol. 13, no. 24, p. 13740, Dec. 2021, doi: 10.3390/su132413740.
S. Sheoran, S. Mohanasundaram, R. Kasilingam, and S. Vij, “Usability and Accessibility of Open Government Data Portals of Countries Worldwide: An Application of TOPSIS and Entropy Weight Method,” International Journal of Electronic Government Research, vol. 19, no. 1, pp. 1–25, Apr. 2023, doi: 10.4018/IJEGR.322307.
K. Taha, P. D. Yoo, C. Yeun, and A. Taha, “A Comprehensive Survey of Text Classification Techniques and Their Research Applications: Observational and Experimental Insights,” Computer Science Review, vol. 54, p. 100664, Nov. 2024, doi: 10.1016/j.cosrev.2024.100664.
S. Minaee, N. Kalchbrenner, E. Cambria, N. Nikzad, M. Chenaghlu, and J. Gao, “Deep Learning--based Text Classification: A Comprehensive Review,” ACM Comput. Surv., vol. 54, no. 3, p. 62:1-62:40, Apr. 2021, doi: 10.1145/3439726.
C. Liu, “Long short-term memory (LSTM)-based news classification model,” PLoS ONE, vol. 19, no. 5, p. e0301835, May 2024, doi: 10.1371/journal.pone.0301835.
Y. Rafat, P. Narayana, R. M. Mohana, and K. Srilatha, “LSTM-Based News Article Category Classification,” Computer Sciences & Mathematics Forum, vol. 12, no. 1, p. 8, 2025, doi: 10.3390/cmsf2025012008.
A. Ezen-Can, “A Comparison of LSTM and BERT for Small Corpus,” 2020, arXiv. doi: 10.48550/ARXIV.2009.05451.
F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” in Proceedings of the 28th International Conference on Computational Linguistics, Barcelona, Spain (Online): International Committee on Computational Linguistics, 2020, pp. 757–770. doi: 10.18653/v1/2020.coling-main.66.
Z. R. K. Rostam and G. Kertész, “Advances in Pre-trained Language Models for Domain-Specific Text Classification: A Systematic Review,” ACM Trans. Intell. Syst. Technol., vol. 16, no. 6, pp. 1–41, Dec. 2025, doi: 10.1145/3763002.
Y. HaCohen-Kerner, D. Miller, and Y. Yigal, “The influence of preprocessing on text classification using a bag-of-words representation,” PLoS ONE, vol. 15, no. 5, p. e0232525, May 2020, doi: 10.1371/journal.pone.0232525.
M. Siino, I. Tinnirello, and M. La Cascia, “Is text preprocessing still worth the time? A comparative survey on the influence of popular preprocessing methods on Transformers and traditional classifiers,” Information Systems, vol. 121, p. 102342, Mar. 2024, doi: 10.1016/j.is.2023.102342.
D. Y. Yefferson, V. Lawijaya, and A. S. Girsang, “Hybrid model: IndoBERT and long short-term memory for detecting Indonesian hoax news,” IJ-AI, vol. 13, no. 2, p. 1913, Jun. 2024, doi: 10.11591/ijai.v13.i2.pp1913-1924.
E. Sirait and J. Ismail, “Comparative Evaluation of IndoBERT-Based Architectures for Imbalanced Indonesian News Title Classification,” SinkrOn, vol. 10, no. 3, pp. 1896–1907, Jul. 2026, doi: 10.33395/sinkron.v10i3.16209.
P. F. Jacobs, G. M. de B. Wenniger, M. Wiering, and L. Schomaker, “Active learning for reducing labeling effort in text classification tasks,” Nov. 03, 2021, arXiv: arXiv:2109.04847. doi: 10.48550/arXiv.2109.04847.
H. Dai et al., “AugGPT: Leveraging ChatGPT for Text Data Augmentation,” 2023.
X.-S. Hong, J.-J. Lee, S.-H. Wu, and M. T.-J. Jiang, “CYUT at the NTCIR-16 FinNum-3 Task: Data Resampling and Data Augmentation by Generation,” 2022.
Z. Li, H. Zhu, Z. Lu, and M. Yin, “Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali, Eds., Singapore: Association for Computational Linguistics, Dec. 2023, pp. 10443–10461. doi: 10.18653/v1/2023.emnlp-main.647.
E. Herrewijnen, D. Nguyen, F. Bex, and K. van Deemter, “Human-annotated rationales and explainable text classification: a survey,” Front. Artif. Intell., vol. 7, May 2024, doi: 10.3389/frai.2024.1260952.
S. Choi, J. Sim, and G. Choi, “Synthetic Text as Data: On Usefulness and Limitations,” Applied Sciences, vol. 15, no. 10, p. 5460, Jan. 2025, doi: 10.3390/app15105460.
M. A. Brown, A. Gruen, G. Maldoff, S. Messing, Z. Sanderson, and M. Zimmer, “Web scraping for research: Legal, ethical, institutional, and scientific considerations,” Big Data & Society, vol. 12, no. 4, p. 20539517251381686, Dec. 2025, doi: 10.1177/20539517251381686.
A. Luscombe, K. Dick, and K. Walby, “Algorithmic thinking in the public interest: navigating technical, legal, and ethical hurdles to web scraping in the social sciences,” Qual Quant, vol. 56, no. 3, pp. 1023–1044, Jun. 2022, doi: 10.1007/s11135-021-01164-0.
“Bapenda Jatim.” Accessed: May 10, 2026. [Online]. Available: https://bapenda.jatimprov.go.id/?utm_source=chatgpt.com
“Website Resmi Bapenda Madiun.” Accessed: May 10, 2026. [Online]. Available: https://bapenda.madiunkota.go.id/
“Teknologi Informasi – Badan Pendapatan Daerah Kota Padang.” Accessed: May 10, 2026. [Online]. Available: https://bapenda.padang.go.id/?cat=7
“SIMANTAP - Sistem Informasi Manajemen Pendapatan Daerah Kota Pasuruan.” Accessed: May 10, 2026. [Online]. Available: https://bapenda.pasuruankota.go.id/?utm_source=chatgpt.com
“BAPENDA KOTA MALANG.” Accessed: May 10, 2026. [Online]. Available: https://bapenda.malangkota.go.id/?utm_source=chatgpt.com
“Bapenda Kabupaten Blitar – Pemerintah Kabupaten Blitar.” Accessed: May 10, 2026. [Online]. Available: https://bapenda.blitarkab.go.id/
“Bapenda Kab. Mojokerto Kab Mojokerto.” Accessed: May 10, 2026. [Online]. Available: https://bapenda.mojokertokab.go.id/?utm_source=chatgpt.com
O. Rainio, J. Teuho, and R. Klén, “Evaluation metrics and statistical tests for machine learning,” Sci Rep, vol. 14, p. 6086, Mar. 2024, doi: 10.1038/s41598-024-56706-x.
H. Heymann, H. Mende, M. Frye, and R. H. Schmitt, “Assessment Framework for Deployability of Machine Learning Models in Production,” Procedia CIRP, vol. 118, pp. 32–37, Jan. 2023, doi: 10.1016/j.procir.2023.06.007.
S. Amershi et al., “Guidelines for Human-AI Interaction,” in Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, in CHI ’19. New York, NY, USA: Association for Computing Machinery, Mei 2019, pp. 1–13. doi: 10.1145/3290605.3300233.
S. F. Taskiran, B. Turkoglu, E. Kaya, and T. Asuroglu, “A comprehensive evaluation of oversampling techniques for enhancing text classification performance,” Sci Rep, vol. 15, no. 1, p. 21631, Jul. 2025, doi: 10.1038/s41598-025-05791-7.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Febrisari Amalia, Andhy Permadi, Ilham Ilham (Author)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.









