Great reflections from SPHERE summer intern Tania-Amanda Fredrick Eneye on what it really takes to make cybersecurity research artifacts reproducible and reusable. We greatly appreciate Tania’s work and contributions to SPHERE this summer, and we’re pleased to see her share some of the lessons she learned along the way. Thank you, Tania!
I just wrapped up my summer internship at USC’s Information Sciences Institute, where I worked on SPHERE Research Infrastructure, a national testbed for reproducible cybersecurity experimentation. My role involved taking published research artifacts from venues including IEEE S&P, USENIX Security, ACSAC, NDSS, and PETS and getting them running on SPHERE so that other researchers could redeploy and verify them. It sounds straightforward. It is not. What I did not expect was that the hardest failures would be the quiet ones: scripts that report success while producing empty output files, a locale setting that silently corrupts results, or a configuration path with the wrong quotation marks that fills a disk without warning until the build fails 40 minutes later. Almost nothing announced itself as broken. Another major lesson was that testing on the machine you built on proves very little. A development node accumulates state that your installation script never had to earn. The real test is deploying on a fresh node from scratch, and that exposed gaps every single time. Reproducibility is often discussed as a research virtue. Up close, it is mostly unglamorous debugging, careful documentation, and being honest when something does not reproduce, then clearly documenting why. Thank you to Jelena Mirkovic, David Balenson, and the SPHERE Research Infrastructure team for their guidance. And thank you to my fellow interns, Allison Lu, Yizhu Wen, Prathyush Turaga, and Sonali Singh, for making a summer full of stack traces genuinely fun.