Indonesian Online News Topic Classification Using Naive Bayes and Support Vector Machine
DOI:
https://doi.org/10.47709/brilliance.v6i3.8905Keywords:
Indonesian online news, Naive Bayes, Support Vector Machine, Text Classification, Web ScrapingAbstract
Indonesian online news portals publish large volumes of articles every day, making manual topic grouping time-consuming, inconsistent, and difficult to maintain when articles contain overlapping contexts. Objective: This study develops and compares Naive Bayes and support vector machine models for Indonesian online news topic classification and identifies the model that can be used as the basis for an automatic newsroom support tool. Methods: A dataset of 1,631 Kompas.com articles was collected through web scraping from six channels: economy, automotive, technology, lifestyle, politics, and sport. The article body was used as input, while the source channel became the class label. The text was prepared through duplicate checking, text normalization, case folding, tokenizing, filtering with Indonesian and custom stopwords, and stemming with Sastrawi. The processed texts were transformed with TF-IDF using a maximum of 10,000 unigram and bigram features, min_df=2, and max_df=0.9. The data were split into 80% training data and 20% testing data with stratification. GridSearchCV with five-fold cross validation was applied to tune Multinomial Naive Bayes and support vector machine parameters. Results: Naive Bayes achieved 97.25% testing accuracy, while support vector machine with a linear kernel, C=10, and gamma=scale achieved 98.17%. SVM also produced higher macro precision, recall, and F1-score. Remaining errors mainly appeared in technology articles because their vocabulary overlapped with lifestyle, automotive, and politics topics. Conclusion: TF-IDF with linear SVM effectively classifies Indonesian online news topics and supports automated content organization workflows.
References
Afdal, M., & Elita, L. R. (2022). PENERAPAN TEXT MINING PADA APLIKASI TOKOPEDIA MENGGUNAKAN ALGORITMA K-NEAREST NEIGHBOR. Jurnal Ilmiah Rekayasa dan Manajemen Sistem Informasi, 8(1), 78. https://doi.org/10.24014/rmsi.v8i1.16595
Arifin, N., Enri, U., & Sulistiyowati, N. (2021). Penerapan Algoritma Support Vector Machine (SVM) dengan TF-IDF N-Gram untuk Text Classification. STRING (Satuan Tulisan Riset dan Inovasi Teknologi), 6(2), 129. https://doi.org/10.30998/string.v6i2.10133
Astuti, Y. P., Wibowo, A. R., Kartikadarma, E., Subhiyakto, E. R., Sri Winarsih, N. A., & Rohman, M. S. (2024). Penerapan Metode Naïve Bayes Classifier Untuk Klasifikasi Sentimen Pada Judul Berita. LogicLink, 1–12. https://doi.org/10.28918/logiclink.v1i1.7684
Asyhar, E. S., Wijoyo, S. H., & Setiawan, N. Y. (2024). Analisis Sentimen dan Pemodelan Topik Terhadap Ulasan Aplikasi Jenius Menggunakan Metode Support Vector Machine dan Latent Dirichlet Allocation.
Djiwadikusumah, F., & Irawan, G. H. (2021). WEB SCRAPING SITUS E-COMMERCE MENGGUNAKAN TEKNIK PARSING DOM.
Habib, S. M., Haerani, E., Gusti, S. K., & Ramadhani, S. (2022). Klasifikasi Berita Menggunakan Metode Naïve Bayes Classifier. Jurnal Nasional Komputasi dan Teknologi Informasi (JNKTI), 5(2), 248–258. https://doi.org/10.32672/jnkti.v5i2.4191
Handoko, M. R. (2021). SISTEM PAKAR DIAGNOSA PENYAKIT SELAMA KEHAMILAN MENGGUNAKAN METODE NAIVE BAYES BERBASIS WEB. Jurnal Teknologi dan Sistem Informasi, 2(1).
Hartati, D., Khaira, U., & Bintana, R. R. (2025). Penerapan Support Vector Machine dan Latent Dirichlet Allocation dalam Analisis Sentimen Terhadap Pengalaman Pengguna Aplikasi Alfagift. 8(5).
Insani, D. F., & Zamzamy, A. (2023). Analisis Framing Pemberitaan Media Online CNBC Indonesia.com dan Kompas.com Mengenai Dampak Lingkungan Pemindahan Ibu Kota Negara.
Joergensen E Munthe, C., Astuti Hasibuan, N., & Hutabarat, H. (2022). Penerapan Algoritma Text Mining Dan TF-RF Dalam Menentukan Promo Produk Pada Marketplace. Resolusi?: Rekayasa Teknik Informatika dan Informasi, 2(3), 110–115. https://doi.org/10.30865/resolusi.v2i3.309
Mayasari, N., Chairul Rizal, & Fitri Wulandari. (2024). Penerapan Metode Klasifikasi Karakteristik Kepribadian Berbasis Website. JOURNAL ZETROEM, 6(2), 27–30. https://doi.org/10.36526/ztr.v6i2.3755
Nanda, R., Haerani, E., Gusti, S. K., & Ramadhani, S. (2022). Klasifikasi Berita Menggunakan Metode Support Vector Machine. Jurnal Nasional Komputasi dan Teknologi Informasi (JNKTI), 5(2), 269–278. https://doi.org/10.32672/jnkti.v5i2.4193
Nur’Faradila, D. A., Magdalena, L., & Febima, M. (2024). Penerapan Naïve Bayes Classifier Untuk Klasifikasi Postingan Berita Hoaks Di Instagram Cirebon Saber Hoaks. Media Jurnal Informatika, 16(2), 243. https://doi.org/10.35194/mji.v16i2.4616
Puspitasari, A. D., & Wulandari, E. R. (2024). Proses Produksi Berita dalam Konvergensi Jurnalistik di Website Suarasurabaya.net. DIGICOM?: Jurnal Komunikasi dan Media, 4(2), 146–151. https://doi.org/10.37826/digicom.v4i2.815
Rianto, Mutiara, A. B., Wibowo, E. P., & Santosa, P. I. (2021). Improving the accuracy of text classification using stemming method, a case of non-formal Indonesian conversation. Journal of Big Data, 8(1), 26. https://doi.org/10.1186/s40537-021-00413-1
Solahuddin, M., Purnamasari, A. I., & Dikananda, A. R. (2023). Klasifikasi Kualitas Berita Pada Majalah Menggunakan Metode Decision Tree. 1(2).
Utami, N. W., & Eka Putra, I. G. J. (2022). TEXT MINIG CLUSTERING UNTUK PENGELOMPOKAN TOPIK DOKUMEN PENELITIAN MENGGUNAKAN ALGORITMA K-MEANS DENGAN COSINE SIMILARITY. Jurnal Informatika Teknologi dan Sains, 4(3), 255–259. https://doi.org/10.51401/jinteks.v4i3.1907
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Winando Parbo Kusuma, Pradita Eko Prasetyo Utomo, Mutia Fadhila Putri

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.















