The source reports a private benchmark that evaluated models with DeepSec, an open-source harness for finding vulnerabilities in large codebases. The open-core application used in the tests was not disclosed.
The comparison considers recall, precision, and cost. These measures can help teams assess performance and expense, but the reported details identify no winner and do not establish that results will transfer to other projects.
To apply the approach, decide in advance which vulnerability types to assess and use a codebase you are authorized to test. Record model versions and test settings so the comparison can be reproduced.
Measure recall, precision, and cost using the same criteria in each run. Review findings before acting: benchmark results are evidence for investigation, not proof that a vulnerability exists or that code is secure.
If you use AI to study or adapt the method, do not submit proprietary code, secrets, or personal data without authorization. Prefer synthetic examples or reviewed excerpts, and follow your organization’s security rules.