Speaker
Description
Machine learning–based anomaly detection can search for new physics in high-dimensional data with minimal theory bias. However, because these methods scan many possibilities at once, they suffer from a look-elsewhere effect that weakens statistical significance. We study this in weakly supervised settings and find a key trade-off: training and testing on the same data gives high sensitivity but badly miscalibrated p-values due to overfitting, while splitting the data and using a fully independent test set yields correct calibration but reduced sensitivity. Methods like early stopping help, but also cost sensitivity. We find that k-folding provides an effective balance, maintaining good calibration while preserving much of the sensitivity. Our findings are supported by numerical studies with Gaussian random variables as well as from collider physics using the LHC Olympics benchmark anomaly detection dataset.