UNDERSTANDING DATA PRIVACY BEHAVIOR AMONG COLLEGE STUDENTS USING EXPLAINABLE MACHINE LEARNING: THE ROLE OF PRIVACY LITERACY, AUTHENTICATION PRACTICES, AND RISK PERCEPTION

Authors

  • Nupur Pandey Author
  • Angrish K. Agarwal Author
  • Arun Kumar Tripathi Author

DOI:

https://doi.org/10.4238/81x40a46

Keywords:

Privacy literacy; privacy protective behavior; explainable artificial intelligence; machine learning; two-factor authentication; risk perception; college students

Abstract

A central question in academic institutions cybersecurity and privacy research is why only some students engage in online privacy-protective behavior while others do not. The previous research relies on linear models that may overlook nonlinear and interactive effects. To address this gap, we applied explainable machine learning to survey data from 745 undergraduate and postgraduate students across seven universities in India, using twelve theory-driven predictors spanning privacy literacy, risk perception, institutional trust, privacy fatigue, and contextual constraints. Five classifiers were evaluated after grid-search tuning and 10-fold cross-validation, and XGBoost emerged as the best-performing model, achieving 68.5% accuracy and an AUC of 0.704 on a held-out test set. SHAP analysis identified privacy literacy as the strongest predictor of privacy-protective behavior, followed by privacy fatigue and reactive behaviors, while the remaining variables contributed relatively little. These findings were corroborated by logistic regression and t-tests, and exploratory clustering suggested weak segment separation, indicating that students vary more continuously than as distinct behavioral types. Overall, the results highlight privacy literacy as a key target for university interventions and demonstrate the value of explainable machine learning for transparent and interpretable student privacy research. Self-reported privacy-policy reading frequency and two-factor authentication use were not predictive of broader protective behavior. These findings suggest that Institutions interventions should prioritize literacy-building over policy-transparency messaging, and that explainable AI offers a transparent, audit-ready alternative to black-box prediction in student-facing cybersecurity research.

Downloads

Published

2026-06-02