Abstract
This study evaluates four clustering methodologies—Hierarchical Agglomerative Clustering (HAC), K-Prototypes, KAMILA, and Spectral Clustering—for detecting organized test collusion using mixed-type data. A two-phase design employed simulated data across 12 conditions and real certification exam data from a documented security breach. Results showed clustering efficacy increased substantially with higher exact response match rates and larger collusion group proportions. HAC emerged as the most robust method in simulations, while all four methods demonstrated exceptional convergence with empirically identified collusion groups in real data. These findings advance test security practice by providing empirically validated guidelines for selecting clustering methodologies and developing strategic implementation approaches, establishing a methodological foundation that equips testing organizations with enhanced collusion detection capabilities for evidence-based enforcement decisions.