The important change: AI has to prove the bug
Security reports can sound convincing without demonstrating a real exploit. MobileCybench approaches that problem differently. Researchers created executable probes that check whether a security property was actually violated after an agent's attempted exploit was replayed.
Five coding-agent configurations were evaluated across four settings: a malicious application on the victim's device or a remote attacker with a low-privilege account, with either an obfuscated APK or application source code available.
Top-agent results with only the obfuscated APK
A triggered probe demonstrates a checked security property failed. It is not the same as saying the agent compromised 53.8% of arbitrary Android apps.
Source: Zhang et al., MobileCybench, submitted September 21, 2026. Results are pass@2 benchmark trigger rates for the top agent in these specific settings.
What we know
Current coding agents can autonomously perform meaningful parts of vulnerability discovery. In MobileCybench, the strongest tested agent triggered executable security probes in multiple real applications even when given only an obfuscated APK. Building and running the benchmark surfaced 23 previously unreported vulnerabilities, with the paper reporting that a majority had been confirmed by maintainers.
What we don't know
This benchmark does not establish that AI can reliably compromise arbitrary real-world software. It covers 13 Android applications, specific attacker privileges, five coding-agent configurations, and a finite inventory of security properties. A probe that does not trigger does not prove an application is secure, and a probe trigger does not by itself tell us the severity or real-world impact of a vulnerability.
What is still speculation
The larger societal effect depends on deployment. If vulnerability discovery becomes dramatically cheaper, defenders could continuously scan and patch software. Attackers could use similar capability to search many more targets. It is too early to know which side gains more.
We used to ask whether AI could find real vulnerabilities. Increasingly, the more useful question is whether defenders can discover, verify, prioritize, and patch them faster than attackers can exploit them.
How to prepare
Developers
Treat AI security testing as an additional layer, not a replacement for review, threat modeling, testing, and responsible disclosure.
Security teams
Prepare for higher report volume. Reproducible evidence and automated validation may become as important as finding the vulnerability itself.
Everyone else
Don't translate benchmark percentages into claims that AI can hack half of all apps. The capability is meaningful; the scope is specific.