Indonesian Online News Topic Classification Using Naive Bayes and Support Vector Machine

Authors

  • Winando Parbo Kusuma Universitas Jambi
  • Pradita Eko Prasetyo Utomo Universitas Jambi
  • Mutia Fadhila Putri Universitas Jambi

DOI:

https://doi.org/10.47709/brilliance.v6i3.8905

Keywords:

Indonesian online news, Naive Bayes, Support Vector Machine, Text Classification, Web Scraping

Abstract

Indonesian online news portals publish large volumes of articles every day, making manual topic grouping time-consuming, inconsistent, and difficult to maintain when articles contain overlapping contexts. Objective: This study develops and compares Naive Bayes and support vector machine models for Indonesian online news topic classification and identifies the model that can be used as the basis for an automatic newsroom support tool. Methods: A dataset of 1,631 Kompas.com articles was collected through web scraping from six channels: economy, automotive, technology, lifestyle, politics, and sport. The article body was used as input, while the source channel became the class label. The text was prepared through duplicate checking, text normalization, case folding, tokenizing, filtering with Indonesian and custom stopwords, and stemming with Sastrawi. The processed texts were transformed with TF-IDF using a maximum of 10,000 unigram and bigram features, min_df=2, and max_df=0.9. The data were split into 80% training data and 20% testing data with stratification. GridSearchCV with five-fold cross validation was applied to tune Multinomial Naive Bayes and support vector machine parameters. Results: Naive Bayes achieved 97.25% testing accuracy, while support vector machine with a linear kernel, C=10, and gamma=scale achieved 98.17%. SVM also produced higher macro precision, recall, and F1-score. Remaining errors mainly appeared in technology articles because their vocabulary overlapped with lifestyle, automotive, and politics topics. Conclusion: TF-IDF with linear SVM effectively classifies Indonesian online news topics and supports automated content organization workflows.

References

Afdal, M., & Elita, L. R. (2022). PENERAPAN TEXT MINING PADA APLIKASI TOKOPEDIA MENGGUNAKAN ALGORITMA K-NEAREST NEIGHBOR. Jurnal Ilmiah Rekayasa dan Manajemen Sistem Informasi, 8(1), 78. https://doi.org/10.24014/rmsi.v8i1.16595

Arifin, N., Enri, U., & Sulistiyowati, N. (2021). Penerapan Algoritma Support Vector Machine (SVM) dengan TF-IDF N-Gram untuk Text Classification. STRING (Satuan Tulisan Riset dan Inovasi Teknologi), 6(2), 129. https://doi.org/10.30998/string.v6i2.10133

Astuti, Y. P., Wibowo, A. R., Kartikadarma, E., Subhiyakto, E. R., Sri Winarsih, N. A., & Rohman, M. S. (2024). Penerapan Metode Naïve Bayes Classifier Untuk Klasifikasi Sentimen Pada Judul Berita. LogicLink, 1–12. https://doi.org/10.28918/logiclink.v1i1.7684

Asyhar, E. S., Wijoyo, S. H., & Setiawan, N. Y. (2024). Analisis Sentimen dan Pemodelan Topik Terhadap Ulasan Aplikasi Jenius Menggunakan Metode Support Vector Machine dan Latent Dirichlet Allocation.

Djiwadikusumah, F., & Irawan, G. H. (2021). WEB SCRAPING SITUS E-COMMERCE MENGGUNAKAN TEKNIK PARSING DOM.

Habib, S. M., Haerani, E., Gusti, S. K., & Ramadhani, S. (2022). Klasifikasi Berita Menggunakan Metode Naïve Bayes Classifier. Jurnal Nasional Komputasi dan Teknologi Informasi (JNKTI), 5(2), 248–258. https://doi.org/10.32672/jnkti.v5i2.4191

Handoko, M. R. (2021). SISTEM PAKAR DIAGNOSA PENYAKIT SELAMA KEHAMILAN MENGGUNAKAN METODE NAIVE BAYES BERBASIS WEB. Jurnal Teknologi dan Sistem Informasi, 2(1).

Hartati, D., Khaira, U., & Bintana, R. R. (2025). Penerapan Support Vector Machine dan Latent Dirichlet Allocation dalam Analisis Sentimen Terhadap Pengalaman Pengguna Aplikasi Alfagift. 8(5).

Insani, D. F., & Zamzamy, A. (2023). Analisis Framing Pemberitaan Media Online CNBC Indonesia.com dan Kompas.com Mengenai Dampak Lingkungan Pemindahan Ibu Kota Negara.

Joergensen E Munthe, C., Astuti Hasibuan, N., & Hutabarat, H. (2022). Penerapan Algoritma Text Mining Dan TF-RF Dalam Menentukan Promo Produk Pada Marketplace. Resolusi?: Rekayasa Teknik Informatika dan Informasi, 2(3), 110–115. https://doi.org/10.30865/resolusi.v2i3.309

Mayasari, N., Chairul Rizal, & Fitri Wulandari. (2024). Penerapan Metode Klasifikasi Karakteristik Kepribadian Berbasis Website. JOURNAL ZETROEM, 6(2), 27–30. https://doi.org/10.36526/ztr.v6i2.3755

Nanda, R., Haerani, E., Gusti, S. K., & Ramadhani, S. (2022). Klasifikasi Berita Menggunakan Metode Support Vector Machine. Jurnal Nasional Komputasi dan Teknologi Informasi (JNKTI), 5(2), 269–278. https://doi.org/10.32672/jnkti.v5i2.4193

Nur’Faradila, D. A., Magdalena, L., & Febima, M. (2024). Penerapan Naïve Bayes Classifier Untuk Klasifikasi Postingan Berita Hoaks Di Instagram Cirebon Saber Hoaks. Media Jurnal Informatika, 16(2), 243. https://doi.org/10.35194/mji.v16i2.4616

Puspitasari, A. D., & Wulandari, E. R. (2024). Proses Produksi Berita dalam Konvergensi Jurnalistik di Website Suarasurabaya.net. DIGICOM?: Jurnal Komunikasi dan Media, 4(2), 146–151. https://doi.org/10.37826/digicom.v4i2.815

Rianto, Mutiara, A. B., Wibowo, E. P., & Santosa, P. I. (2021). Improving the accuracy of text classification using stemming method, a case of non-formal Indonesian conversation. Journal of Big Data, 8(1), 26. https://doi.org/10.1186/s40537-021-00413-1

Solahuddin, M., Purnamasari, A. I., & Dikananda, A. R. (2023). Klasifikasi Kualitas Berita Pada Majalah Menggunakan Metode Decision Tree. 1(2).

Utami, N. W., & Eka Putra, I. G. J. (2022). TEXT MINIG CLUSTERING UNTUK PENGELOMPOKAN TOPIK DOKUMEN PENELITIAN MENGGUNAKAN ALGORITMA K-MEANS DENGAN COSINE SIMILARITY. Jurnal Informatika Teknologi dan Sains, 4(3), 255–259. https://doi.org/10.51401/jinteks.v4i3.1907

Downloads

Published

2026-07-05

How to Cite

Kusuma, W. P., Prasetyo Utomo, P. E., & Fadhila Putri, M. (2026). Indonesian Online News Topic Classification Using Naive Bayes and Support Vector Machine. Brilliance: Research of Artificial Intelligence, 6(3), 343–353. https://doi.org/10.47709/brilliance.v6i3.8905

Similar Articles

<< < 1 2 3 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.