Presented at DEF CON 34
Improving AI Red-Teaming by Systematizing Red-Teaming Reports
Abstract
AI red-teaming faces serious challenges: lack of scientific grounding, inconsistent testing practices, and misalignment with external stakeholders. Because the practice is highly context-dependent and qualitative, reporting remains inconsistent, creating a gap between the information produced and what is needed to improve the field. This fragmentation hinders interpretability, limits stakeholder trust, and complicates efforts to mitigate AI risks. These problems are especially pertinent as the practice of AI red-teaming adapts to evaluate agentic systems.
In our work, we conducted a qualitative interview study with 17 AI red-teaming practitioners to explore how testing practices influence reporting and identify the challenges preventing standardization. Our thematic analysis reveals that organizational context and threat models are the primary frameworks shaping red-teaming exercises. We identify four key dimensions that practitioners agree are essential to report for improved transparency and utility: the threat model, methodological details, harms elicited during testing, and actionable information for mitigation.
We will present on the challenges to the AI red-teaming ecosystem, what AI red-teamers want from reports, and how red-teamers can systematize reporting practices to help the AI community move toward more consistent, useful, and interpretable evaluations of AI systems. Solving the challenges that face AI red-teaming is vital for the emerging era of agentic AI, where increased autonomy and interaction complexity require more rigorous, interpretable evaluation methods.