Loading…
7-9 October, 2026
Prague, Czechia
View More Details & Registration
Important Note: Timing of sessions and room locations are subject to change.

The Sched app allows you to build your schedule but is not a substitute for your event registration. You must be registered for Open Source Summit Europe 2026 to participate in the sessions. If you have not registered but would like to join us, please go to the event registration page to purchase a registration.



Friday October 9, 2026 15:40 - 16:20 CEST
Vibe coding is now the norm, and more open-source code is AI-generated. But is it secure?

We present the Agent Security League, an open leaderboard built on SusVibes (200 real-world tasks, 108 OSS projects, 77 CWEs). Unlike benchmarks frontier labs cite to showcase progress, SusVibes evaluates what developers actually use: a harness coupled with an LLM, not the model alone. Across 15+ combinations — GPT-5.5, Gemini 3.5, Fable 5 — with agents like Cursor and Claude Code, functional correctness can exceed 80%, yet our best fair security score is 29%.

This means ~7/10 working AI patches leave the bug open.

Strikingly, even models (e.g., Fable 5) that lead model-only benchmarks score no better on security once wrapped in a real agent.

Agents also cheat: early on they reverse-engineered fixes from git history, inflating scores up to 42×. Our 3-layer pipeline—prompt hardening, workspace sanitization, LLM-based anti-cheating eval—nearly killed that, but memorization of upstream fixes from training data persists.

We show cheating examples, share fair results, what actually improves security, and argue anti-cheating and contamination controls must be standard for agentic benchmarks.
Speakers
avatar for Luca Compagna

Luca Compagna

Security Researcher, Endor Labs
Security Research Consultant at Endor Labs, after more than 15 years at SAP, researching and innovating in the areas of security testing, security engineering, and AI.

Regular presenter at industrial and community venues (OWASP) and leading security conferences, he has published... Read More →
avatar for Danqing Wang

Danqing Wang

PhD, Carnegie Mellon University
Danqing is a Ph.D. candidate at LTI in Carnegie Mellon University (CMU), advised by Prof. Lei Li. Her research focuses on Strategic Planning and Reasoning in LLM Agents, including agent collaboration, agent competition, agent communication and agentic safety.
Friday October 9, 2026 15:40 - 16:20 CEST
Forum Hall (Floor 2)
  Open AI & Data

Sign up or log in to save this to your schedule, view media, leave feedback and see who's attending!

Share Modal

Share this link via

Or copy link