Deteksi Potensi Karir Mahasiswa Universitas Airlangga Berbasis Pemodelan Machine Learning

Mohammad Nathiq Ulman, Faisal Fahmi

Abstract

Abstract. Many students and recent graduates face uncertainty when mapping post-graduation pathways—whether to enter employment, continue formal education, or pursue entrepreneurship—creating practical constraints for university career services and tracer-study reporting. Using tracer-study records as an institutional evidence base, this study develops and evaluates a multiclass classification model for predicting alumni employment status from tabular data and positions the exercise as an institutional analytics problem rather than a purely technical benchmark. The target variable is grouped into three classes: Employed (full-time/part-time) / previously employed, Pursuing further education, and Entrepreneur / previously entrepreneur. The modeling workflow uses XGBoost with an integrated preprocessing pipeline consisting of data cleaning, numerical missing-value imputation, numerical feature standardization, and one-hot encoding for categorical variables. The data are split using an 80:20 stratified train-test scheme. Performance is evaluated with class-wise precision, recall, F1-score, and a confusion matrix; model interpretability is examined using SHAP. The model reaches an overall accuracy of 0.88, but class performance is uneven: Employed is classified very well (precision 0.87; recall 1.00; F1-score 0.93), Pursuing further education is predicted with high precision but moderate recall (precision 0.96; recall 0.50; F1-score 0.66), and Entrepreneur is detected very poorly (precision 1.00; recall 0.04; F1-score 0.08). SHAP shows different combinations of academic, institutional, and financing-related features across classes, while the Entrepreneur class lacks sufficiently strong and consistent separability. These findings indicate that tracer study modeling should prioritize class-wise evaluation, class imbalance analysis, and interpretability, because aggregate accuracy alone can overstate practical readiness for institutional use.

Keywords

alumni tracer study; multiclass classification; XGBoost; class imbalance; SHAP; model interpretability; alumni employment status

Full Text:

PDF

References

Referensi

Abdulloh, F. F., Rahardi, M., Aminuddin, A., Anggita, S. D., & Nugraha, A. Y. A. (2022). Observation of imbalance tracer study data for graduates employability prediction in Indonesia. International Journal of Advanced Computer Science and Applications (IJACSA), 13(8), 169–174. https://doi.org/10.14569/IJACSA.2022.0130820

Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). https://doi.org/10.1145/2939672.2939785

Haque, R., Quek, A., Ting, C.-Y., Goh, H.-N., & Hasan, M. R. (2024). Classification techniques using machine learning for graduate student employability predictions. International Journal on Advanced Science, Engineering and Information Technology, 14(1), 45–56.

He, H., & Garcia, E. A. (2009). Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering, 21(9), 1263–1284. https://doi.org/10.1109/TKDE.2008.239

Krawczyk, B. (2016). Learning from imbalanced data: Open challenges and future directions. Progress in Artificial Intelligence, 5(4), 221–232. https://doi.org/10.1007/s13748-016-0094-0

Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (Vol. 30). (Also available as arXiv:1705.07874). https://arxiv.org/abs/1705.07874

Refbacks

  • There are currently no refbacks.