- What: OWASP project focuses on agent security testing
- Impact: Helps organizations evaluate AI system security
Subscribe Share Full episode and show notes Application security , AI benefits/risks , Attack surface management Inside the OWASP Agent Security Regression Harness Project – Mert Satilmaz – ASW #393 Orgs need to be able to use agents, MCPs, and LLMs in ways that don’t lead to unexpected actions and undesirable outcomes. The OWASP Agent Security Regression Harness project is an approach for defining customizable scenarios and testing whether those systems fail against known security threats. Mert Saltimaz talks about the background of the project, how orgs can use it as they bring more LLMs into their environment, and how the project intends to grow. Importantly, we also talk about the security controls and designs that orgs can build around the systems and data that models interact with in addition to evaluating the security of the agents and agent harnesses themselves. Segment Resource... July 28, 2026 Full Segment Notes Orgs need to be able to use agents, MCPs, and LLMs in ways that don't lead to unexpected actions and undesirable outcomes. The OWASP Agent Security Regression Harness project is an approach for defining customizable scenarios and testing whether those systems fail against known security threats. Mert Saltimaz talks about the background of the project, how orgs can use it as they bring more LLMs into their environment, and how the project intends to grow. Importantly, we also talk about the security controls and designs that orgs can build around the systems and data that models interact with in addition to evaluating the security of the agents and agent harnesses themselves. Segment Resources: https://github.com/OWASP/Agent-Security-Regression-Harness https://youtu.be/6DWs5EwbFQ0?si=r0IJ_F0SZnkPzzYg -- "What Trading Systems Taught Me About Breaking (And Defending) Infrastructure" Guest Mert Satilmaz SecPortal-Founder at SecPortal Founder of SecPortal, an AI-native cybersecurity platform designed to manage the full security lifecycle, from vulnerability discovery to compliance. OWASP Project Lead for the Agent Security Regression Harness, focused on advancing security testing for AI-driven systems. Security researcher with publicly disclosed CVEs, identifying and reporting real-world vulnerabilities across production systems. I also publish technical research and security insights, with articles featured on HackerNoon, and speak at industry events including SteelCon. Background in application security, cloud environments, and security automation, with over a decade of hands-on engineering experience. Certified Cyber Essentials Plus Lead Assessor, supporting organisations in improving their security posture. MSc in Information Security from Royal Holloway, University of London. Building SecPortal for security teams that need faster scanning, triage, remediation tracking, reporting, and compliance. Hosts Mike Shema https://dangerouserrors.com Tyler Shields https://www.90degree.vc/ Announcements AppSec teams, your backlog is growing faster than you can fix it. SAST and DAST tools are flooding you with findings, developers are pushing back, and prioritizing what actually matters in code is getting harder. So how do you reduce risk without slowing releases? Join the Vulnerability Management Virtual Cybersecurity Summit to learn how teams are prioritizing real exploitable issues, reducing noise, and integrating remediation into modern development workflows. Security Weekly listeners can register for free at https://securityweekly.com/vulnmanagement using the promo code: CSS26-SW InfoSec World brings cybersecurity professionals together across industries, from healthcare and financial services to government and the Fortune 500. Join the community in Orlando, October 12–14, for practical education, new perspectives, and cybersecurity research unveiled live. Listeners save 30% on their pass with code ISW26-SWSAVINGS at securityweekly.com/infosecworld2026. List of Articles Mike Shema OpenAI and Hugging Face partner to address security incident during model evaluation This story has been framed in so many different ways that it's almost indistinguishable from telling the future by reading tea leaves. See also https://huggingface.co/blog/security-incident-july-2026 https://www.forrester.com/blogs/an-ai-security-facepalm-openais-evaluation-became-hugging-faces-incident/ https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations https://news.risky.biz/srsly-risky-biz-knives-are-out-for-open-weight-ai-models/ And, of course, try your own hand at ExploitGym Releases now reject new files after 14 days – The Python Package Index Blog Why not be proactive about a secure design? Not only does this change remove an attack vector, but it reduces ambiguity in the state of a compromised release. And it was done in partnership with the community and evaluation of what side-effects it might cause. Next chapter: Restructuring GitHub’s bug bounty program Another, perhaps unsurprising, side-effect of the proliferation of LLMs is that guardrails to defend against LLM-generated cruft increasingly rely on human experts. RefluXFS: A Linux Kernel Local Privilege Escalation to Root in XFS (CVE-2026-64600) | Qualys Security at Machine Speed Is the Wrong Race Show More Stay in the Know, No Smoke and Mirrors – Join Our Newsletter Get expert insights and technical breakdowns straight to your inbox. Join Now Related Segments Application security MacOS Security Design Features, Flaws, And Futures – Patrick Wardle – ASW #392 Vulnerability Management AI Security at Scale, CMMC phase II paused, and the Weekly Enterprise News – Keith Hollender – ESW #468 Application security Discovering & Securing Your AI Agent Attack Surface – Jeremy Snyder – ASW #391 Related Content Application security GitHub and PyPI implement new security measures against supply-chain attacks Application security Malvertising campaign assembles malware in browser AI/ML AI plans are being approved. Recovery plans are not You can skip this ad in 5 seconds