We reproduced Anthropic's Mythos findings with public models. See the results >>

Jev is cheap. But it's not good for security

Cheap, fast models look like an easy win for the small decisions inside a vulnerability detection pipeline. I tested that idea by having Jev and seven LLMs score real CVEs using CVSS. The savings are real and too small to matter, and the accuracy you give up is not.

Sep 30, 2026 | 13 min read

Since last week, every second post in my feed has been about Jev - the model that doesn't return open text but closed answers instead. How good it is, how quick it is, how cheap it is... Thus, since I'm in the security industry and use LLMs a lot, the question naturally popped into my mind - can we use it to make decisions within the flow of vulnerability detection? Be it a static or dynamic vulnerability hunt. Cause if it's good, quick and cheap - why wouldn't we use it in our product? So I made a test to see how its understanding of security compares to 7 different LLM models.

Here's a sneak peek into the results👀

Metric accuracy, bare prompt
sonnet-589.11%
gemini-3.8-flash88.84%
kimi-k388.31%
glm-5.387.77%
gemini-3.1-pro87.10%
gpt-5.6-sol86.96%
gpt-5.6-luna84.81%
jev80.38%
70%100%
Each bar: 744 metric decisions, 31 advisories x 8 metrics x 3 runs.

What Jev is

Jev is not an LLM - it's a classifier. It won't give you text, or the reasoning behind an answer. It can give you a yes/no answer, a score or an answer selected from a closed set of propositions that you give it. You can ask an LLM to reply in a fixed format and it mostly will. Here the format is guaranteed.

Another great feature of Jev is that it returns a probability per option. So if you give it a very simple question and three options, it will maybe tell you that there's a 0% probability it's option A, 97% it's option B and 3% it's option C. If it's unsure of the answer, maybe the split would be 30%, 40%, 30%. In both examples option B won but in the latter example you might consider running the question again through a smarter model.

Jev's designed to make quick, calibrated decisions. It's advertised to be 193.6x quicker and 444.6x cheaper than LLMs on system one tasks. That's why I want to see if it has its use in security.

The test

My idea was to test Jev's understanding of security by giving it a vulnerability report and asking it about CVSS metrics of that particular report. You can ask 8 questions about the 8 metrics in parallel and you end up with the whole CVSS vector. If you have the vector, you can calculate the score with a few lines of code.

And as much as CVSS is prone to generate disagreements, it's the best idea I've had to verify such a model's understanding of security in a quantifiable way.

The dataset

To gather the dataset of bugs, I first downloaded a few dozen reports from the GitHub Security Advisory database (GHSA). I filtered out those that had CVSS scores clearly incorrectly set - and it was a lot. I was left with 11 criticals, 9 highs and 11 mediums. You can see all the reports in the companion repository to this blogpost.

Importantly, the report, stripped of its CVSS vector, base score and severity rating, was all the models were given - they couldn't read the source and get the whole context around the bug. If they were allowed to, their results would be better. But this is not a blogpost to answer whether models are good at rating CVSS metrics.

The models

These are the models I tested.

modelserved as
jevTypeSafe /v1/systemone, jev-latest (answered as jev-1.13.0)
sonnet-5anthropic/claude-sonnet-5
gpt-5.6-solopenai/gpt-5.6-sol
gpt-5.6-lunaopenai/gpt-5.6-luna
gemini-3.1-progoogle/gemini-3.1-pro-preview
gemini-3.8-flashgoogle/gemini-3.8-flash
glm-5.3z-ai/glm-5.3
kimi-k3moonshotai/kimi-k3

Every model that reasons ran at low reasoning effort. Temperature was set to 0 wherever the model accepts it - Sonnet 5 and both GPT-5.6 variants reject the parameter, so they ran at their own default.

Everything but Jev ran through OpenRouter, with max_tokens 16000.

The prompts

For each bug, I had two different prompts. One that would rely on each model's pre-training and another that also included the CVSS specification and its worked examples, straight from FIRST.

An example bare prompt would look like this:

1You are scoring a software vulnerability report with CVSS v3.1 base metrics. 2Read the report and output the CVSS v3.1 base vector string. 3 4Output rules: 5- Output ONLY the vector string, nothing else. No explanation, no markdown, no 6 code fence. 7- Exact format: CVSS:3.1/AV:?/AC:?/PR:?/UI:?/S:?/C:?/I:?/A:? 8 9Vulnerability report: 10 11# MLflow: Unauthenticated full-read SSRF in webhook delivery: 12_validate_webhook_url bypassed via unvalidated HTTP redirects (and DNS 13rebinding) 14 15- CWE: CWE-918 16- Packages: pip:mlflow 17 18### Summary 19The default MLflow Tracking Server (`mlflow server`, no authentication, 20default SQLite 21backend) exposes the model-registry webhooks API unauthenticated, including a 22synchronous 23`POST /api/2.0/mlflow/webhooks/{id}/test` endpoint that returns the upstream 24response 25status and body to the caller. ... 26 27CVSS v3.1 base vector:

And one with the reference would look like this:

1You are scoring a software vulnerability report with CVSS v3.1 base metrics. 2Read the report and output the CVSS v3.1 base vector string. 3 4Output rules: 5- Output ONLY the vector string, nothing else. No explanation, no markdown, no 6 code fence. 7- Exact format: CVSS:3.1/AV:?/AC:?/PR:?/UI:?/S:?/C:?/I:?/A:? 8 9## AV - Attack Vector 10 11This metric reflects the context by which vulnerability exploitation is 12possible. This metric value (and consequently the Base Score) will be larger 13the more remote (logically, and physically) an attacker can be in order to 14exploit the vulnerable component. ... 15 16- `AV:N` Network - The vulnerable component is bound to the network stack and 17 the set of possible attackers extends beyond the other options listed below, 18 up to and including the entire Internet. ... 19- `AV:A` Adjacent - ... 20- `AV:L` Local - ... 21- `AV:P` Physical - ... 22 23## Scoring guidance 24 25- When deciding between Network and Adjacent, if an attack can be launched 26 over a wide area network or from outside the logically adjacent 27 administrative network domain, use Network. ... 28 29## Examples 30 31- MySQL Stored SQL Injection (CVE-2013-0375), `AV:N` - The attacker connects 32 to the exploitable MySQL database over a network. 33- ... 34 35## AC - Attack Complexity 36... 37 38Vulnerability report: 39 40# MLflow: Unauthenticated full-read SSRF in webhook delivery: 41_validate_webhook_url bypassed via unvalidated HTTP redirects (and DNS 42rebinding) 43 44- CWE: CWE-918 45- Packages: pip:mlflow 46 47### Summary 48The default MLflow Tracking Server (`mlflow server`, no authentication, 49default SQLite 50backend) exposes the model-registry webhooks API unauthenticated, including a 51synchronous 52`POST /api/2.0/mlflow/webhooks/{id}/test` endpoint that returns the upstream 53response 54status and body to the caller. ... 55 56CVSS v3.1 base vector:

For Jev, the bare prompt would look like this:

1{ 2 "model": "jev-latest", 3 "state": "Vulnerability report:\n\n# MLflow: Unauthenticated full-read ...", 4 "questions": { 5 "AV": { 6 "type": "choice", 7 "instructions": "Attack Vector (AV)", 8 "criteria": { 9 "Attack Vector: Network": "Attack Vector: Network", 10 "Attack Vector: Adjacent": "Attack Vector: Adjacent", 11 "Attack Vector: Local": "Attack Vector: Local", 12 "Attack Vector: Physical": "Attack Vector: Physical" 13 } 14 }, 15 "AC": { ... }, "PR": { ... }, "UI": { ... }, 16 "S": { ... }, "C": { ... }, "I": { ... }, "A": { ... } 17 } 18}

And the one with the spec like this:

1{ 2 "model": "jev-latest", 3 "state": { 4 "advisory": "Vulnerability report:\n\n# MLflow: Unauthenticated full ...", 5 "cvss_specification": "## AV - Attack Vector\n\nThis metric reflects ..." 6 }, 7 "questions": { 8 "AV": { 9 "type": "choice", 10 "instructions": "Attack Vector (AV)", 11 "criteria": { 12 "Attack Vector: Network": { 13 "covers": "The vulnerable component is bound to the network st ...", 14 "belongs_to_another_option": [ 15 "Adjacent covers: The vulnerable component is bound to the n ...", 16 "Local covers: The vulnerable component is not bound to the ...", 17 "Physical covers: The attack requires the attacker to physic ..." 18 ], 19 "examples": [ 20 "MySQL Stored SQL Injection (CVE-2013-0375), `AV:N` - The ...", 21 "..." 22 ] 23 }, 24 "Attack Vector: Adjacent": { ... }, 25 "Attack Vector: Local": { ... }, 26 "Attack Vector: Physical": { ... } 27 } 28 }, 29 "AC": { ... }, "PR": { ... }, "UI": { ... }, 30 "S": { ... }, "C": { ... }, "I": { ... }, "A": { ... } 31 } 32}

All eight metrics were sent in a single call. After testing a few versions of where to put the CVSS reference, the best option turned out to be in both the state and the response criteria. It is redundant, and it still scored best.

From now on, the versions without the CVSS reference will be called bare and the ones with will be called spec.

Results

In total, I had 31 reports, 8 models, 2 variants and 3 runs per combination. On each model, each report was run three times with the CVSS reference and three times without.

Here are the results with bare prompts:

Metric accuracy, bare prompt
sonnet-589.11%
gemini-3.8-flash88.84%
kimi-k388.31%
glm-5.387.77%
gemini-3.1-pro87.10%
gpt-5.6-sol86.96%
gpt-5.6-luna84.81%
jev80.38%
70%100%
Each bar: 744 metric decisions, 31 advisories x 8 metrics x 3 runs.
modelmetric accuracyexact 8/8total costvs Jevmedian latency
jev80.38%20.43%$0.01161x0.36s
sonnet-589.11%44.09%$0.648055.6x1.99s
gpt-5.6-sol86.96%27.96%$0.847672.8x3.97s
gpt-5.6-luna84.81%26.88%$0.05354.6x5.42s
gemini-3.1-pro87.10%27.96%$1.2261105.3x4.88s
gemini-3.8-flash88.84%33.33%$0.237220.4x9.23s
glm-5.387.77%43.01%$0.164814.1x1.24s
kimi-k388.31%29.03%$0.558247.9x2.30s

The accuracy is low enough that I wouldn't use Jev this way, so let's move on to the results with the CVSS specification added to the prompt.

Results with the CVSS specification in the prompt

Metric accuracy, spec prompt
sonnet-589.65%+0.54
gpt-5.6-sol88.58%+1.61
glm-5.388.17%+0.40
gemini-3.8-flash87.90%-0.94
kimi-k387.77%-0.54
gpt-5.6-luna87.37%+2.55
gemini-3.1-pro86.69%-0.40
jev85.08%+4.70
70%100%
Each bar: 744 metric decisions, 31 advisories x 8 metrics x 3 runs. Second figure: change in points against the same model without the spec.
modelmetric accuracyexact 8/8total costvs Jevmedian latency
jev85.08%22.58%$0.05731x0.44s
sonnet-589.65%51.61%$2.539744.3x2.08s
gpt-5.6-sol88.58%30.11%$1.887832.9x3.79s
gpt-5.6-luna87.37%33.33%$0.10781.9x4.35s
gemini-3.1-pro86.69%34.41%$2.141537.4x5.16s
gemini-3.8-flash87.90%26.88%$0.865015.1x12.49s
glm-5.388.17%36.56%$0.20673.6x1.70s
kimi-k387.77%29.03%$0.878815.3x3.23s

We can see that the two cheapest models - Jev and gpt-5.6-luna - have significant increases in accuracy. On this chart, Jev does actually get close to those other models. We see some improvement on gpt-5.6-sol and then Sonnet, both Gemini models, GLM and Kimi are probably within the error margin.

Price of the spec run, as a multiple of Jev
jev1x
gpt-5.6-luna1.9x
glm-5.33.6x
gemini-3.8-flash15.1x
kimi-k315.3x
gpt-5.6-sol32.9x
gemini-3.1-pro37.4x
sonnet-544.3x
1x50x
Billed cost of the same 93 calls.

As much as Jev is cheaper than all of the models, with some of them being multiple times as expensive, in comparison to gpt-5.6-luna or glm-5.3 the difference is not that big. If you take a look at glm-5.3 and compare its performance with a bare prompt versus Jev with the specification, GLM is 2.87x as expensive while delivering 2.69pp better accuracy.

While 2.87x or 1.9x still sound like a lot, the tradeoff is also significant. The context window especially. Jev takes 64k tokens per request, of which the state plus the longest question can use 32k, with Luna at 1M. That's crucial for security, especially as the answer often does not lie in the prompt itself but in files that the LLM can read while Jev can't.

In Jev's defense, the confidence it returns with every answer was a good predictor of whether the case was easy to judge. On average, the confidence for its good answers was 87% versus 50% when it was wrong. So you could find a threshold somewhere there and use Jev's answer if it has high enough confidence and use an LLM if Jev's response does not have satisfying confidence.

Which metrics are the most problematic?

Here's the heatmap of how often the models got each metric right.

Accuracy per CVSS base metric, both prompts pooled
60%100%
jevsonnet-5gpt-5.6-solgpt-5.6-lunagemini-3.1-progemini-3.8-flashglm-5.3kimi-k3
AVAttack Vector97%100%100%99%100%98%99%100%
ACAttack Complexity97%95%97%96%94%96%97%96%
PRPrivileges Required92%90%95%94%94%92%93%94%
UIUser Interaction90%96%97%98%98%97%98%97%
SScope62%78%73%73%81%74%74%72%
CConfidentiality81%86%87%85%84%85%85%89%
IIntegrity79%84%84%78%80%86%81%82%
AAvailability63%85%69%66%66%79%77%75%
One cell: one model on one metric, 186 answers. Colours follow the 60% to 100% key above.

You can see that they definitely had more problems with the Impact metrics plus Scope than with Exploitability. It definitely makes sense. Exploitability metrics are directly derived from what the model sees in the report. The impact often has more abstraction to it. For example, it's obvious to me and you that reading another user's access token has Integrity impact because you can take over someone's account and make changes as them. But at the very first glance, it sounds like a read-only bug, so a simple model might easily get this wrong.

Here's the exact list of bugs, starting with the ones that were the most confusing for models.

The hardest bugs to rate

The metric column is the one the models missed most on that bug: C Confidentiality 100% means every model missed Confidentiality, every time.

advisorymost mistaken metricaccuracy
Sorted hardest first.

Or maybe the models were right and the authors were wrong?👀 Again, this was not the blogpost about models' accuracy in rating CVSS.

The verdict

Is Jev a good way to save money and time in a vulnerability detection flow?

No.

Its place is in flows that make huge numbers of quick decisions, where everything needed is already in the state. Vulnerability detection too often requires more reasoning and perhaps reading more than one file. While saving 2x or 4x sounds like a lot, these are small decisions so you'd need to be making millions of them to see measurable savings expressed as dollars and not as a ratio. I'd rather dedicate my time to shaving off 10% from a $10,000 bill than to reducing a $10 bill by 100x.

Share:

Ready to secure your application?

Vidoc finds and fixes vulnerabilities in real-time.
Ship secure applications faster.

Get VIDOC

More articles

Explore insights, trends, and tips to stay ahead in cybersecurity.