Jev vs. Claude vs. GPT: Your “Critical” Findings Are Stealing Your Sprint
It isn't revelatory to note that coverage of existing AI chatbots targets criticism towards their sycophancy. The redundancies of chatbot speak are so well documented to have earned the title "Claudespeak"—"It's not X, it's Y."
Of course, while these complaints are generally directed at the concept of artificial intelligence broadly, they, in fact, identify issues with chatbots reliant upon large language models, in particular.
So, when TypeSafe announced Jev on September 15, its pitch was compelling: a model built for structured decisions inside software stripped of the inefficiency and redundancy intrinsic to interacting with a human user. Rather than output expensive, redundant, and potentially hallucinated strings, Jev outputs type-safe structured values accompanied by calibrated probabilities and confidence intervals.
Jev won so much hype, so rapidly, that within forty-eight hours, a meme caricaturing the model's partisans entered the X canon.
Exhibit A.
Exhibit B.
Need I scroll more?
But just how "insane" is Jev, actually? The Casco team decided to put Jev to the test against seven other popular models. Security severity scoring looked to us like a perfect test. CVSS turns a judgment call into eight fixed choices that feed a published, deterministic formula, so every model's answer can be checked against a reference score, and every confidence claim can be tested against how often the model was actually right.
So, we sampled 2,449 findings and passed them through to eight models, including Jev, Claude, and GPT. We compared their CVSS 3.1 scores with Casco’s published scores. Every model received the same finding evidence, without Casco’s scoring policy or internal application context.
Every model scored higher than the reference on average. Jev assigned 550 critical scores to a dataset containing 55 published critical findings. Of those 550, 500 were below critical in Casco’s reference. Even the best overall performer, GPT-6 Astra, gave a positive score to 652 of 1,002 findings Casco had published as informational.
Overdiagnosing creates a busywork problem: an AI can generate a plausible severity rating much faster than a person can establish whether its implied impact is real. As a result, security engineers are overwhelmed by "security slop", distracting them from identifying and resolving the vulnerabilities that actually matter.
Explore the model results, inspect the evidence below, or jump to the time-savings calculator.
Why so serious? Upward bias in severity scores
We asked each model to select the eight CVSS 3.1 base metrics. The same calculator converted every vector into a score. This isolates the judgment from the arithmetic.
The primary ranking measures mean absolute error: the average distance between a model’s score and Casco’s published reference. Astra ranked first at 1.92 points. Jev ranked seventh at 3.40, ahead of Haiku at 3.94. Jev’s structured output did not make its severity judgments the most accurate against this reference.
Importantly, the direction of the disparity remained consistent across models: Jev’s mean signed error was +3.30 points, Haiku’s +3.89, and Astra’s +1.18. All shifted upward on a ten-point scale. In other words, every single model overestimates the severity of vulnerabilities in software.
01 / The noise scoreboard
Overestimation across the board
2,449 findings per model. Lower is better for the selected measure.
Average distance from Casco’s published score, in CVSS points. This is agreement with a reference, not an independent verdict on correctness.
Show agreement, upward bias, and the risk of under-scoring
| Model | Mean signed error | Severity agreement | Critical recall |
|---|---|---|---|
| GPT-6 Astra | +1.18 | 53.4% | 60.0% |
| Claude Opus 4.6 | +1.68 | 47.0% | 74.5% |
| GPT-5.6 Terra | +1.71 | 49.0% | 67.3% |
| GPT-5.6 Sol | +1.85 | 45.0% | 61.8% |
| Claude Sonnet 4.6 | +2.08 | 42.8% | 74.5% |
| GPT-5.6 Luna | +2.50 | 40.5% | 87.3% |
| Jev 1.13.0 | +3.30 | 30.8% | 90.9% |
| Claude Haiku 4.5 | +3.89 | 25.6% | 94.5% |
Try switching from average error to informational escalation. 65.1% to 98.4% of the informational references received positive scores, depending on the model. For Jev, 35.7% of informational findings became high or critical; for Haiku, 52.2% did.
Among 2,394 references below critical, Jev promoted 20.9% to critical; Astra promoted 1.5%. Note the denominator: these rates cover only references that were below critical, so every promotion counted here is a false one. A model that rarely assigns "critical" at all would score near zero, so Astra's 1.5% doesn't by itself show that it distinguishes critical from noncritical findings better than Jev.
In fact, Astra is worse at correctly identifying true vulnerabilities. Astra kept only 60.0% of the 55 critical references in the critical band, versus 90.9% for Jev and 94.5% for Haiku. Jev is more likely to sound the alarm on insignificant vulnerabilities whereas Astra is more likely to ignore real concerns, placing security specialists asked to rely on either in a difficult bind. In cybersecurity, we need both justified urgency and preserved coverage.
The biggest gap is understanding impact
A model can recognize that an endpoint is reachable over the network, but may fail to gauge why an attacker is incentivized to reach that endpoint in the first place.
Jev agreed with the published attack vector 98.5% of the time but agreed on confidentiality impact only 50.2% of the time, and integrity impact 59.8%. Haiku’s confidentiality agreement was 35.0%. Even Astra and Opus scored around 64–65% on that component.
What protected information becomes accessible? What can an attacker improperly change? What can an attacker disrupt or take offline? The answers depend on context: the attacker's identity and privileges, who owns the affected data, what the system is designed to permit, and whether a viable exploit path exists. A description that merely resembles a security flaw establishes none of these.
The category breakdown makes the pattern more concrete:
- Security misconfiguration: 879 findings. This is the largest group, and 655 were informational in the reference. Jev gave positive scores to 537 of those 655; Haiku did so for 643. Missing or permissive controls are particularly easy to turn into hypothetical downstream damage.
- Access control and authorization: 656 findings. Jev’s average error fell to 1.96 points here, compared with 4.53 for misconfiguration. The severity mixes differ. Astra even had a small negative bias here; upward bias is not universal across every slice.
- Authentication: 271 findings. Jev and Haiku assigned positive scores to all 49 informational references. A weak control and a demonstrated account takeover are different claims, with different evidence requirements.
- Smaller categories need more caution. Resource consumption contains 35 findings, and AI/LLM security contains 27. Treat their model rankings cautiously.
02 / Where the score goes wrong
The impact gap.
How often does Jev 1.13.0 choose the same CVSS component as the published reference? Each component has 2,449 comparisons.
Break it down by finding category
Positive bias means a higher score than Casco’s reference. Category labels are grouped by meaning across OWASP versions; these are descriptive slices, not equally sized test sets.
| Category / sample | Mean error | Mean bias | 0.0 → positive |
|---|---|---|---|
| Security misconfigurationn = 879 | 4.53 | +4.51 | 537 / 655 |
| Access control / authorizationn = 656 | 1.96 | +1.87 | 47 / 52 |
| Other / unclassifiedn = 300 | 3.87 | +3.76 | 121 / 150 |
| Authenticationn = 271 | 3.27 | +3.21 | 49 / 49 |
| Injectionn = 92 | 2.17 | +1.98 | 14 / 17 |
| Design / business logicn = 87 | 2.50 | +2.20 | 20 / 24 |
| Cryptographyn = 69 | 4.15 | +3.79 | 23 / 23 |
| Resource consumptionn = 35 · small sample | 3.23 | +2.97 | 12 / 13 |
| Server-side request forgeryn = 33 · small sample | 2.49 | +1.69 | 3 / 5 |
| AI / LLM securityn = 27 · small sample | 3.65 | +3.51 | 12 / 14 |
The last column counts positive model scores among informational references in that category. Different category mixes can change the overall ranking.
These groups consolidate existing category labels across OWASP versions; they are not independent reclassifications.
The context failure: turning “could” into “did”
We inspected disagreements in the raw evidence. A recurring pattern was a conditional chain: a control is weak → another attack might become possible → assume the downstream impact.
The missing step is often where the security decision lives. Can an attacker actually inject a script? Obtain someone else’s token? Affect another tenant? Exhaust a service within reachable limits? Or are they using a feature as intended?
Select a case below, compare models, and reveal the context check. The evidence is paraphrased; the scores and vectors are measured.
03 / Read between the scores
The word “could” is doing a lot of work.
Four selected cases, paraphrased to remove customer details. Scores and vectors are from the frozen run. Our annotations explain the evidence gap; the models were not asked to provide reasoning.
A permissive policy becomes a successful exploit.
What the evidence says
A content security policy allows WebSocket connections to arbitrary destinations. No XSS injection point was demonstrated.
The conditional impact
If an attacker can execute JavaScript, that policy could let it send data elsewhere.
Jev 1.13.0
7.5 / 10
CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:NCasco published reference
0.0
CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:N/I:N/A:NThe missing step is attacker-controlled JavaScript execution. A weakened containment control does not, by itself, establish access to protected data.
Ask before escalating: Where is the injection point, and what protected data can the attacker actually read?
These are illustrative disagreements, not proof that the published reference is always correct or that every overstatement has the same cause.
The permissive-CSP case is especially instructive. Its evidence explicitly said no XSS vector had been demonstrated. Jev scored it 7.5; Casco’s published reference was 0.0. Astra, Terra, and Luna also returned zero. The relevant limitation was already in the supplied text, yet models applied it differently.
The WebSocket and AI-instructions cases ask different versions of the same question: what establishes the claimed harm? Open connections are not an observed outage. A user changing their own AI output is not automatically crossing a security boundary.
These are our interpretations, not the models’ internal reasoning: the run collected choices, not explanations. Without a rerun adding context, we cannot attribute every disagreement to context failures. Scoring conventions, ambiguous evidence, and reference errors can also contribute.
TypeSafe’s own Jev limitations describe difficulties with multi-step indirection and distracting context. Typed output does not remove the work of establishing the right premises.
So is CVSS wrong?
CVSS base severity describes technical vulnerability characteristics and assumes reasonable worst-case impact. It is not a complete business-risk or sprint-priority score. FIRST explicitly distinguishes base scoring from environmental adjustments and other organizational prioritization factors. CVSS 3.1 specification.
Two mistakes follow: assigning impact without sufficient evidence, and treating base severity as an instruction to drop everything.
A hypothetical chain still needs defensible premises. But the opposite shortcut is also wrong: “we did not execute a destructive exploit” is not sufficient to assign zero when other evidence establishes impact. And a local mitigation may belong in environmental scoring rather than silently changing the base score.
This experiment measures agreement with Casco’s published decisions, not universal CVSS correctness. Models were not taught Casco’s conventions. The repeated upward bias is measurable; a claim that their training made them alarmist is not established here.
How much engineering time does that cost?
When notified of an inflated security ticket, an engineer is tasked to reconstruct the prerequisites, check intended behavior, and explain why the ticket does not deserve its severity. Only then can they recover their place in the feature they were building.
The benchmark did not track those hours. The calculator below turns the observed informational-escalation counts into a planning scenario, with your investigation time and an editable assumption for Casco’s remaining rechecks.
At the starting settings—100 findings per month, 30 minutes per investigation, Jev as the comparator, and an assumed 5% Casco recheck rate—the estimate is 85% less time on those rechecks: about 14.6 hours a month back for building features. That is a modeled outcome, not a measured customer result or an 85% reduction in all security work.
04 / Put a price on the interruption
How much of your sprint is noise?
Estimate time spent rechecking findings that a model scores above zero but Casco published as informational. The model’s rate is measured here. Casco’s future recheck rate and your investigation time are assumptions you control.
Assumes the same severity mix as the benchmark, including 40.9% informational findings.
Here, false positives means unnecessary positive-severity escalations, not nonexistent vulnerabilities.
Share of all findings requiring the same kind of avoidable recheck. The 5% starting value is a planning assumption, not a measured result of this benchmark.
85%
less time on avoidable rechecks
In this scenario, using Casco frees up 14.6 hours / month for your team to build features.
34.2 rechecks × 30 minutes
5.0 rechecks × 30 minutes
Scenario estimate, not observed customer savings. Excludes genuine finding investigation, remediation, onboarding, and other security work.
Show the math and assumptions
Jev 1.13.0 escalated 838 of the 1,002 informational references. Across all 2,449 findings, that is a 34.2% workload rate. We use this all-findings denominator because your monthly input includes every severity.
Model hours = 100 × (838 / 2,449) × 30 / 60 = 17.1
Casco hours = 100 × 5% × 30 / 60 = 2.5
Reduction = (model hours − Casco hours) / model hours
Each escalation is assumed to trigger one investigation of the same duration in both workflows. Informational findings can still merit attention. A different finding mix, triage threshold, or Casco recheck rate changes the result. This does not measure detection recall or guarantee time savings.
The workload denominator is all findings: Jev escalated 838 / 2,449, or 34.2%, from informational to positive severity. The estimate assumes each causes an investigation; a different triage threshold changes that assumption.
Use Casco to resolve the context before it consumes your sprint
Casco builds an understanding of your application, validates exploit claims, and retains what your team teaches it about intended behavior.
Our context refinement loops turn false-positive feedback into application context that your engineers can inspect and edit. That gives future investigations the information a standalone scoring call lacks: which behavior is intentional, what a role can legitimately do, and where the real security boundaries sit.
Use Casco to reduce avoidable investigations, so your engineers can get back to building features. Put your own assumptions into the calculator, then book a demo to see the workflow against your application.
Methodology and data
This run relied upon 2,449 records composed of 1,002 informational, 130 low, 932 medium, 330 high, and 55 critical references. Remediated findings retained historical severity.
Jev 1.13.0 answered eight Choice questions. Claude Haiku 4.5, Sonnet 4.6, and Opus 4.6 used forced structured output at temperature zero without extended thinking. GPT-5.6 Sol, Terra, and Luna disabled reasoning; GPT-6 Astra used low reasoning and a larger output budget. This is not an equal-reasoning-budget comparison. Existing Jev/Claude outcomes were reused for the OpenAI extension.
The shared calculator was checked against all 2,592 possible base vectors. Downloaded mean-error intervals bootstrap applications, not independent findings. One prediction per case measures neither repeated-run variability nor discovery performance. There was no independent adjudication, context ablation, or time-savings study.