Skip to content
Rota Nacional

Guides ·

Benchmark compares models for vulnerability discovery

A private benchmark uses the open-source DeepSec harness to compare model recall, precision, and cost on an undisclosed open-core application, offering reference points for software security evaluations.

The source reports a private benchmark that evaluated models with DeepSec, an open-source harness for finding vulnerabilities in large codebases. The open-core application used in the tests was not disclosed.

The comparison considers recall, precision, and cost. These measures can help teams assess performance and expense, but the reported details identify no winner and do not establish that results will transfer to other projects.

To apply the approach, decide in advance which vulnerability types to assess and use a codebase you are authorized to test. Record model versions and test settings so the comparison can be reproduced.

Measure recall, precision, and cost using the same criteria in each run. Review findings before acting: benchmark results are evidence for investigation, not proof that a vulnerability exists or that code is secure.

If you use AI to study or adapt the method, do not submit proprietary code, secrets, or personal data without authorization. Prefer synthetic examples or reviewed excerpts, and follow your organization’s security rules.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free