Evaluasi Kinerja Pendekatan Latent Semantic Indexing dan Kombinasi Latent Semantic Indexing - K Nearest Neighbor pada Klasifikasi Dokumen Beban Kerja Dosen


Authors

  • Khairun Nadiah Universitas Negeri Medan, Medan, Indonesia
  • Mansur As Universitas Negeri Medan, Medan, Indonesia
  • Said Iskandar Al Idrus Universitas Negeri Medan, Medan, Indonesia
  • Mulyono Mulyono Universitas Negeri Medan, Medan, Indonesia
  • Zulfahmi Indra Universitas Negeri Medan, Medan, Indonesia

DOI:

https://doi.org/10.47065/bulletincsr.v6i5.1274

Keywords:

Lecturer Workload; Document Classification; Latent Semantic Indexing; Cosine Similarity; K-Nearest Neighbor

Abstract

The Lecturer Workload (BKD) document is an academic administrative document categorized based on higher education's Tridharma (triple dharma) activities, making it highly valuable as a dataset for document classification research. However, the characteristics of BKD documents which feature diverse academic terms and an imbalanced class distribution pose a distinct challenge in the classification process. This study aims to evaluate and compare the performance of two classification approaches based on Latent Semantic Indexing (LSI): LSI + Cosine Similarity and LSI + K-Nearest Neighbor (KNN). As a benchmark, TF-IDF + KNN is utilized as a baseline method to analyze the impact of LSI on classification performance. The dataset consists of 352 Postgraduate BKD documents from Universitas Negeri Medan, which underwent text extraction, preprocessing (case folding, cleaning, tokenization, stopword removal, stemming, normalization, and BKD keyword enrichment), TF-IDF weighting, and dimensionality reduction using LSI. Evaluation was conducted using Stratified 3-Fold Cross Validation with Accuracy, Precision, Recall, F1-Score, ROC-AUC, and Average Precision (AP) as metrics. The results indicate that the LSI + Cosine Similarity approach delivers the best performance, achieving an Accuracy of 86.3%, F1-Weighted of 0.865, AUC-Macro of 0.965, and AP-Macro of 0.887. Meanwhile, LSI + KNN achieved 84.6% Accuracy, and the TF-IDF + KNN baseline reached 82.6% Accuracy. These findings demonstrate that semantic representation using LSI successfully enhances the quality of BKD document classification, while the Cosine Similarity-based approach proves more stable than KNN on datasets characterized by varied academic terminology and imbalanced class distributions.

Downloads

Download data is not yet available.

References

M. Solekhah and W. Lasniah, “Analisis Proses Bisnis Sistem Informasi Manajemen Dokumen Pendukung Beban Kerja Dosen Dengan Metode Protoyping Model,” Sebatik, vol. 25, no. 2, pp. 356–365, 2021, doi: 10.46984/sebatik.v25i2.1656.

A. H. Ardiansyah, K. P. Kartika, and S. N. Budiman, “Penerapan Latent Semantic Indexing Pada Sistem Temu Balik Informasi Pada Undang-Undang Pemilu Berdasarkan Kasus,” J. Mnemon., vol. 4, no. 2, pp. 64–70, 2021, doi: 10.36040/mnemonic.v4i2.4165.

F. Istighfarizky, N. A. Sanjaya ER, I. M. Widiartha, L. G. Astuti, I. G. N. A. C. Putra, and I. K. G. Suhartana, “Klasifikasi Jurnal menggunakan Metode KNN dengan Mengimplementasikan Perbandingan Seleksi Fitur,” Jeliku (Jurnal Elektron. Ilmu Komput. Udayana), vol. 11, no. 1, p. 167, 2022, doi: 10.24843/jlk.2022.v11.i01.p18.

R. A. Murniati, “Analisis Sentimen Ulasan Aplikasi Mobile Banking M-Smile Menggunakan Metode Latent Semantic Indexing (LSI),” Universitas Paramadina, 2021. [Online]. Available: https://repository.paramadina.ac.id/1483/1/Jurnal_Retno Asih M_120103007.pdf

R. C. Rivaldi and T. D. Wismarini, “Analisis Sentimen Pada Ulasan Produk Dengan Metode Natural Language Processing (NLP) (Studi Kasus Zalika Store 88 Shopee),” Elkom J. Elektron. dan Komput., vol. 17, no. 1, pp. 120–128, 2024, doi: 10.51903/elkom.v17i1.1680.

N. Ajijah, A. Kurniawan, and Susilawati, “Klasifikasi Teks Mining Terhadap Analisa Isu Kegiatan Tenaga Lapangan Menggunakan Algoritma K-Nearest Neighbor (KNN),” J-Sakti (Jurnal Sains Komput. Inform., vol. 7, no. 1, pp. 254–262, 2023, doi: 10.30645/j-sakti.v7i1.589.

R. Kosasih and A. Alberto, “Analisis Sentimen Produk Permainan Menggunakan Metode TF-IDF Dan Algoritma K-Nearest Neighbor,” InfoTekJar J. Nas. Inform. dan Teknol. Jar., vol. 6, no. 1, pp. 134–139, 2021, doi: 10.30743/infotekjar.v6i1.3893.

P. Mehta, S. Aggarwal, and A. Tandon, “The Effect of Topic Modelling on Prediction of Criticality Levels of Software Vulnerabilities,” Inform., vol. 47, no. 6, pp. 145–158, 2023, doi: 10.31449/inf.v47i6.3712.

T. Ridwansyah, “Implementasi Text Mining Terhadap Analisis Sentimen Masyarakat Dunia Di Twitter Terhadap Kota Medan Menggunakan K-Fold Cross Validation Dan Naïve Bayes Classifier,” Klik Kaji. Ilm. Inform. Dan Komput., vol. 2, no. 5, pp. 178–185, 2022, doi: 10.30865/klik.v2i5.362.

O. W. Yuda, D. Tuti, L. S. Yee, and Susanti, “Penerapan Penerapan Data Mining Untuk Klasifikasi Kelulusan Mahasiswa Tepat Waktu Menggunakan Metode Random Forest,” Satin - Sains dan Teknol. Inf., vol. 8, no. 2, pp. 122–131, 2022, doi: 10.33372/stn.v8i2.885.

A. Dahari, D. Herwanto, and J. Arifin, “Analisa Pengendalian Persediaan Bahan Baku Bumbu Racik Makanan dari Raw Material Hingga Barang Jadi (Finish Good) di PT. Ariake Europe Indonesia,” J. Ilm. Wahana Pendidik., vol. 7, no. 1, pp. 391–402, 2021, doi: 10.5281/zenodo.5535550.

J. Jefriyanto, N. Ainun, and M. A. Al Ardha, “Application of Naïve Bayes Classification to Analyze Performance Using Stopwords,” Jiste (Journal Inf. Syst. Technol. Eng., vol. 1, no. 1, pp. 49–53, 2023, doi: 10.61487/jiste.v1i2.15.

A. Sinaga and S. P. Nainggolan, “Analisis Perbandingan Akurasi Dan Waktu Proses Algoritma Stemming Arifin-Setiono Dan Nazief-Adriani Pada Dokumen Teks Bahasa Indonesia,” Sebatik, vol. 27, no. 1, pp. 63–69, 2023, doi: 10.46984/sebatik.v27i1.2072.

A. R. Lubis and M. K. M. Nasution, “Twitter Data Analysis and Text Normalization in Collecting Standard Word,” J. Appl. Eng. Technol. Sci., vol. 4, no. 2, pp. 855–863, 2023, doi: 10.37385/jaets.v4i2.1991.

Supiyanto and Sriyono, “Metode Cosine Similarity Untuk Mendeteksi Kemiripan Pada Dokumen Teks,” J. Mipa dan Pengajarannya, vol. 1, no. 1, pp. 1–7, 2023, doi: 10.31957/sains.v23i1.3661.

R. F. Putra et al., Algoritma Pembelajaran Mesin ( Dasar , Teknik , dan Aplikasi ), Pertama., no. April. Bekasi: PT. Sonpedia Publishing Indonesia, 2024. [Online]. Available: https://www.researchgate.net/publication/379479664_ALGORITMA_PEMBELAJARAN_MESIN_Dasar_Teknik_dan_Aplikasi

A. Muzakir, A. Desiani, and A. Amran, “Klasifikasi Penyakit Kanker Prostat Menggunakan Algoritma Naïve Bayes Classification of Prostate Cancer Using Naïve Bayes and K-Nearest Neighbor Algorithms,” Komputika J. Sist. Komput., vol. 12, no. 148, pp. 73–79, 2023, doi: 10.34010/komputika.v12i1.9629.

W. P. Anggraini, M. S. Utami, J. M. Berlianty, E. S. Hutagalung, Y. Juniarto, and R. Nooraeni, “Klasifikasi Sentimen Masyarakat Terhadap Kebijakan Kartu Pekerja Di Indonesia,” Fakt. Exacta, vol. 13, no. 4, p. 255, 2021, doi: 10.30998/faktorexacta.v13i4.7964.

R. R. Adhitya, W. Witanti, and R. Yuniarti, “Perbandingan Metode Cart Dan Naïve Bayes Untuk Klasifikasi Customer Churn,” Infotech J., vol. 9, no. 2, pp. 307–318, 2023, doi: 10.31949/infotech.v9i2.5641.

M. C. Rani, F. D. Azkia, R. A. Dewi, M. Wahyudi, Sumanto, and A. S. Budiman, “Perbandingan Algoritma Random Forest, Naive Bayes, dan Neural Network dalam Klasifikasi Penyakit Jantung,” J. Sains Inform. Terap. ( Jsit ), vol. 4, no. 2, pp. 187–201, 2025, doi: 10.62357/jsit.v4i2.609.

Angelina, N. Filosofia, and R. Arga Wijaya, “Optimizing Predictions For Thyroid Disease Sufferers Using Correlation Matrix And Random Forest With Hyperparameter Tuning,” Media J. Gen. Comput. Sci., vol. 2, no. 1, pp. 1–11, 2025, doi: 10.62205/mjgcs.v2i1.26.

W. A. Kurniawan and A. Salam, “Penggunaan Feature Space Smote Untuk Mengurangi Overfitting Akibat Imbalance Dataset,” Techno.Com, vol. 23, no. 2, pp. 328–337, 2024, doi: 10.62411/tc.v23i2.10215.

I. W. Amitaba, A. Pramudiya, H. Ahyadi, and L. S. C. Ranesa, “Optimasi Pemodelan Hujan Limpasan dan Analisis Genangan Banjir untuk Evaluasi Dampak Infrastruktur DAS Rea,” J. Sumber Daya Air, vol. 21, no. 2, pp. 101–114, 2025, doi: 10.32679/jsda.v21i2.970.


Bila bermanfaat silahkan share artikel ini

Berikan Komentar Anda terhadap artikel Evaluasi Kinerja Pendekatan Latent Semantic Indexing dan Kombinasi Latent Semantic Indexing - K Nearest Neighbor pada Klasifikasi Dokumen Beban Kerja Dosen

Dimensions Badge

ARTICLE HISTORY

Published: 2026-08-10

Abstract View: 0 times
PDF Download: 0 times

How to Cite

Nadiah, K., As, M., Idrus, S. I. A., Mulyono, M., & Indra, Z. (2026). Evaluasi Kinerja Pendekatan Latent Semantic Indexing dan Kombinasi Latent Semantic Indexing - K Nearest Neighbor pada Klasifikasi Dokumen Beban Kerja Dosen. Bulletin of Computer Science Research, 6(5), 1988-2000. https://doi.org/10.47065/bulletincsr.v6i5.1274

Issue

Section

Articles

Most read articles by the same author(s)