Since last week, every second post in my feed has been about Jev - the model that doesn't return open text but closed answers instead. How good it is, how quick it is, how cheap it is... Thus, since I'm in the security industry and use LLMs a lot, the question naturally popped into my mind - can we use it to make decisions within the flow of vulnerability detection? Be it a static or dynamic vulnerability hunt. Cause if it's good, quick and cheap - why wouldn't we use it in our product? So I made a test to see how its understanding of security compares to 7 different LLM models.
Here's a sneak peek into the results👀
What Jev is
Jev is not an LLM - it's a classifier. It won't give you text, or the reasoning behind an answer. It can give you a yes/no answer, a score or an answer selected from a closed set of propositions that you give it. You can ask an LLM to reply in a fixed format and it mostly will. Here the format is guaranteed.
Another great feature of Jev is that it returns a probability per option. So if you give it a very simple question and three options, it will maybe tell you that there's a 0% probability it's option A, 97% it's option B and 3% it's option C. If it's unsure of the answer, maybe the split would be 30%, 40%, 30%. In both examples option B won but in the latter example you might consider running the question again through a smarter model.
Jev's designed to make quick, calibrated decisions. It's advertised to be 193.6x quicker and 444.6x cheaper than LLMs on system one tasks. That's why I want to see if it has its use in security.
The test
My idea was to test Jev's understanding of security by giving it a vulnerability report and asking it about CVSS metrics of that particular report. You can ask 8 questions about the 8 metrics in parallel and you end up with the whole CVSS vector. If you have the vector, you can calculate the score with a few lines of code.
And as much as CVSS is prone to generate disagreements, it's the best idea I've had to verify such a model's understanding of security in a quantifiable way.
The dataset
To gather the dataset of bugs, I first downloaded a few dozen reports from the GitHub Security Advisory database (GHSA). I filtered out those that had CVSS scores clearly incorrectly set - and it was a lot. I was left with 11 criticals, 9 highs and 11 mediums. You can see all the reports in the companion repository to this blogpost.
Importantly, the report, stripped of its CVSS vector, base score and severity rating, was all the models were given - they couldn't read the source and get the whole context around the bug. If they were allowed to, their results would be better. But this is not a blogpost to answer whether models are good at rating CVSS metrics.
The models
These are the models I tested.
| model | served as |
|---|---|
| jev | TypeSafe /v1/systemone, jev-latest (answered as jev-1.13.0) |
| sonnet-5 | anthropic/claude-sonnet-5 |
| gpt-5.6-sol | openai/gpt-5.6-sol |
| gpt-5.6-luna | openai/gpt-5.6-luna |
| gemini-3.1-pro | google/gemini-3.1-pro-preview |
| gemini-3.8-flash | google/gemini-3.8-flash |
| glm-5.3 | z-ai/glm-5.3 |
| kimi-k3 | moonshotai/kimi-k3 |
Every model that reasons ran at low reasoning effort. Temperature was set to 0 wherever the model accepts it - Sonnet 5 and both GPT-5.6 variants reject the parameter, so they ran at their own default.
Everything but Jev ran through OpenRouter, with max_tokens 16000.
The prompts
For each bug, I had two different prompts. One that would rely on each model's pre-training and another that also included the CVSS specification and its worked examples, straight from FIRST.
An example bare prompt would look like this:
1You are scoring a software vulnerability report with CVSS v3.1 base metrics.
2Read the report and output the CVSS v3.1 base vector string.
3
4Output rules:
5- Output ONLY the vector string, nothing else. No explanation, no markdown, no
6 code fence.
7- Exact format: CVSS:3.1/AV:?/AC:?/PR:?/UI:?/S:?/C:?/I:?/A:?
8
9Vulnerability report:
10
11# MLflow: Unauthenticated full-read SSRF in webhook delivery:
12_validate_webhook_url bypassed via unvalidated HTTP redirects (and DNS
13rebinding)
14
15- CWE: CWE-918
16- Packages: pip:mlflow
17
18### Summary
19The default MLflow Tracking Server (`mlflow server`, no authentication,
20default SQLite
21backend) exposes the model-registry webhooks API unauthenticated, including a
22synchronous
23`POST /api/2.0/mlflow/webhooks/{id}/test` endpoint that returns the upstream
24response
25status and body to the caller. ...
26
27CVSS v3.1 base vector:And one with the reference would look like this:
1You are scoring a software vulnerability report with CVSS v3.1 base metrics.
2Read the report and output the CVSS v3.1 base vector string.
3
4Output rules:
5- Output ONLY the vector string, nothing else. No explanation, no markdown, no
6 code fence.
7- Exact format: CVSS:3.1/AV:?/AC:?/PR:?/UI:?/S:?/C:?/I:?/A:?
8
9## AV - Attack Vector
10
11This metric reflects the context by which vulnerability exploitation is
12possible. This metric value (and consequently the Base Score) will be larger
13the more remote (logically, and physically) an attacker can be in order to
14exploit the vulnerable component. ...
15
16- `AV:N` Network - The vulnerable component is bound to the network stack and
17 the set of possible attackers extends beyond the other options listed below,
18 up to and including the entire Internet. ...
19- `AV:A` Adjacent - ...
20- `AV:L` Local - ...
21- `AV:P` Physical - ...
22
23## Scoring guidance
24
25- When deciding between Network and Adjacent, if an attack can be launched
26 over a wide area network or from outside the logically adjacent
27 administrative network domain, use Network. ...
28
29## Examples
30
31- MySQL Stored SQL Injection (CVE-2013-0375), `AV:N` - The attacker connects
32 to the exploitable MySQL database over a network.
33- ...
34
35## AC - Attack Complexity
36...
37
38Vulnerability report:
39
40# MLflow: Unauthenticated full-read SSRF in webhook delivery:
41_validate_webhook_url bypassed via unvalidated HTTP redirects (and DNS
42rebinding)
43
44- CWE: CWE-918
45- Packages: pip:mlflow
46
47### Summary
48The default MLflow Tracking Server (`mlflow server`, no authentication,
49default SQLite
50backend) exposes the model-registry webhooks API unauthenticated, including a
51synchronous
52`POST /api/2.0/mlflow/webhooks/{id}/test` endpoint that returns the upstream
53response
54status and body to the caller. ...
55
56CVSS v3.1 base vector:For Jev, the bare prompt would look like this:
1{
2 "model": "jev-latest",
3 "state": "Vulnerability report:\n\n# MLflow: Unauthenticated full-read ...",
4 "questions": {
5 "AV": {
6 "type": "choice",
7 "instructions": "Attack Vector (AV)",
8 "criteria": {
9 "Attack Vector: Network": "Attack Vector: Network",
10 "Attack Vector: Adjacent": "Attack Vector: Adjacent",
11 "Attack Vector: Local": "Attack Vector: Local",
12 "Attack Vector: Physical": "Attack Vector: Physical"
13 }
14 },
15 "AC": { ... }, "PR": { ... }, "UI": { ... },
16 "S": { ... }, "C": { ... }, "I": { ... }, "A": { ... }
17 }
18}And the one with the spec like this:
1{
2 "model": "jev-latest",
3 "state": {
4 "advisory": "Vulnerability report:\n\n# MLflow: Unauthenticated full ...",
5 "cvss_specification": "## AV - Attack Vector\n\nThis metric reflects ..."
6 },
7 "questions": {
8 "AV": {
9 "type": "choice",
10 "instructions": "Attack Vector (AV)",
11 "criteria": {
12 "Attack Vector: Network": {
13 "covers": "The vulnerable component is bound to the network st ...",
14 "belongs_to_another_option": [
15 "Adjacent covers: The vulnerable component is bound to the n ...",
16 "Local covers: The vulnerable component is not bound to the ...",
17 "Physical covers: The attack requires the attacker to physic ..."
18 ],
19 "examples": [
20 "MySQL Stored SQL Injection (CVE-2013-0375), `AV:N` - The ...",
21 "..."
22 ]
23 },
24 "Attack Vector: Adjacent": { ... },
25 "Attack Vector: Local": { ... },
26 "Attack Vector: Physical": { ... }
27 }
28 },
29 "AC": { ... }, "PR": { ... }, "UI": { ... },
30 "S": { ... }, "C": { ... }, "I": { ... }, "A": { ... }
31 }
32}All eight metrics were sent in a single call. After testing a few versions of where to put the CVSS reference, the best option turned out to be in both the state and the response criteria. It is redundant, and it still scored best.
From now on, the versions without the CVSS reference will be called bare and the ones with will be called spec.
Results
In total, I had 31 reports, 8 models, 2 variants and 3 runs per combination. On each model, each report was run three times with the CVSS reference and three times without.
Here are the results with bare prompts:
| model | metric accuracy | exact 8/8 | total cost | vs Jev | median latency |
|---|---|---|---|---|---|
| jev | 80.38% | 20.43% | $0.0116 | 1x | 0.36s |
| sonnet-5 | 89.11% | 44.09% | $0.6480 | 55.6x | 1.99s |
| gpt-5.6-sol | 86.96% | 27.96% | $0.8476 | 72.8x | 3.97s |
| gpt-5.6-luna | 84.81% | 26.88% | $0.0535 | 4.6x | 5.42s |
| gemini-3.1-pro | 87.10% | 27.96% | $1.2261 | 105.3x | 4.88s |
| gemini-3.8-flash | 88.84% | 33.33% | $0.2372 | 20.4x | 9.23s |
| glm-5.3 | 87.77% | 43.01% | $0.1648 | 14.1x | 1.24s |
| kimi-k3 | 88.31% | 29.03% | $0.5582 | 47.9x | 2.30s |
The accuracy is low enough that I wouldn't use Jev this way, so let's move on to the results with the CVSS specification added to the prompt.
Results with the CVSS specification in the prompt
| model | metric accuracy | exact 8/8 | total cost | vs Jev | median latency |
|---|---|---|---|---|---|
| jev | 85.08% | 22.58% | $0.0573 | 1x | 0.44s |
| sonnet-5 | 89.65% | 51.61% | $2.5397 | 44.3x | 2.08s |
| gpt-5.6-sol | 88.58% | 30.11% | $1.8878 | 32.9x | 3.79s |
| gpt-5.6-luna | 87.37% | 33.33% | $0.1078 | 1.9x | 4.35s |
| gemini-3.1-pro | 86.69% | 34.41% | $2.1415 | 37.4x | 5.16s |
| gemini-3.8-flash | 87.90% | 26.88% | $0.8650 | 15.1x | 12.49s |
| glm-5.3 | 88.17% | 36.56% | $0.2067 | 3.6x | 1.70s |
| kimi-k3 | 87.77% | 29.03% | $0.8788 | 15.3x | 3.23s |
We can see that the two cheapest models - Jev and gpt-5.6-luna - have significant increases in accuracy. On this chart, Jev does actually get close to those other models. We see some improvement on gpt-5.6-sol and then Sonnet, both Gemini models, GLM and Kimi are probably within the error margin.
As much as Jev is cheaper than all of the models, with some of them being multiple times as expensive, in comparison to gpt-5.6-luna or glm-5.3 the difference is not that big. If you take a look at glm-5.3 and compare its performance with a bare prompt versus Jev with the specification, GLM is 2.87x as expensive while delivering 2.69pp better accuracy.
While 2.87x or 1.9x still sound like a lot, the tradeoff is also significant. The context window especially. Jev takes 64k tokens per request, of which the state plus the longest question can use 32k, with Luna at 1M. That's crucial for security, especially as the answer often does not lie in the prompt itself but in files that the LLM can read while Jev can't.
In Jev's defense, the confidence it returns with every answer was a good predictor of whether the case was easy to judge. On average, the confidence for its good answers was 87% versus 50% when it was wrong. So you could find a threshold somewhere there and use Jev's answer if it has high enough confidence and use an LLM if Jev's response does not have satisfying confidence.
Which metrics are the most problematic?
Here's the heatmap of how often the models got each metric right.
You can see that they definitely had more problems with the Impact metrics plus Scope than with Exploitability. It definitely makes sense. Exploitability metrics are directly derived from what the model sees in the report. The impact often has more abstraction to it. For example, it's obvious to me and you that reading another user's access token has Integrity impact because you can take over someone's account and make changes as them. But at the very first glance, it sounds like a read-only bug, so a simple model might easily get this wrong.
Here's the exact list of bugs, starting with the ones that were the most confusing for models.
The metric column is the one the models missed most on that bug: C Confidentiality 100% means every model missed Confidentiality, every time.
Or maybe the models were right and the authors were wrong?👀 Again, this was not the blogpost about models' accuracy in rating CVSS.
The verdict
Is Jev a good way to save money and time in a vulnerability detection flow?
No.
Its place is in flows that make huge numbers of quick decisions, where everything needed is already in the state. Vulnerability detection too often requires more reasoning and perhaps reading more than one file. While saving 2x or 4x sounds like a lot, these are small decisions so you'd need to be making millions of them to see measurable savings expressed as dollars and not as a ratio. I'd rather dedicate my time to shaving off 10% from a $10,000 bill than to reducing a $10 bill by 100x.