
A security-advisory triage agent built from a retrieval engine and a decision model. One finds what's true right now. The other decides what to do about it.
Security teams are drowning, and it isn’t because they’re bad at their jobs.
In 2025 the world published 48,185 CVEs — about 132 a day, the latest in a run of record years unbroken since 2017. No human reads all of that.
The harder truth is that almost none of it matters to you. Across a quarter-million CVEs studied by Cyentia and FIRST, only about 6% have ever been seen exploited in the wild. So the job was never "read faster." It's "find the handful that matter — every day, honestly, without burning people out." It’s a signal-filtering problem: cutting through 94% noise to isolate the tiny fraction of exploitability before a threat actor does.
Comparing notes, I realized we were each holding half of the answer. You.com does real-time search: it can surface a fresh advisory the moment it lands. TypeSafe's Jev makes fast, calibrated decisions: it can judge that advisory against what you actually run. They complement perfectly to build an agent that solves the signal-filtering problem mentioned earlier.
I built a small, open-source agent to prove it. What follows is the what, the why, and the how.
One agent, three moves.
It runs on a schedule and does exactly three things: pull the day’s advisories for your stack, decide what each one means for you, and route only what’s worth a human’s attention — recording every step.
The judgment is four atomic questions — TypeSafe’s Choice, Score, and Noul primitives — asked in parallel against a description of your stack:
| Question | Type | What it decides |
|---|---|---|
affects_us | Noul | Is an affected component actually in your stack? |
severity | Score | informational → critical, for your stack |
exploited | Noul | Actively exploited, or a public exploit exists? |
action | Choice | patch_now · schedule · monitor · ignore |
Three problems, stacked.
Any tool that ignores one of these quietly fails. Together they explain why triage needs both fresh evidence and real judgment — two different capabilities.
The volume is unwinnable by hand
48,185 CVEs in 2025, up ~21% on 2024’s own record of 40,009. January 2025 alone set a single-month record. This is not a backlog you staff your way out of.
Severity is the wrong sort key
The industry default — CVSS — measures how bad a bug would be if exploited, not whether it will be. FIRST, which maintains CVSS, says so in its own docs: “CVSS Base scores are not risk.”
The data is stark. Sort by severity and most of your “critical” queue describes attacks that never come.
When something does matter, it’s already fast
Google’s Mandiant tracks time-to-exploit, and it has collapsed. In 2024–25 the average went negative — the typical exploited vulnerability is now weaponized around or before the day its patch is public.
Meanwhile the feed everyone leans on — NIST’s NVD — hit a documented enrichment backlog in 2024 and now fully enriches only a risk-based subset of new CVEs. The reference data is slower and less complete than it used to be.
So anything that decides from a model’s frozen training memory, or last week’s export, is deciding about yesterday’s world.
Two engines, one clean seam.
Here’s the heart of it: a real, important feature — one a security team would actually depend on — was built from retrieval by You.com and decisioning by Jev, and it works precisely because neither tries to be the other.
| Aspect | You.com · RetrievalFresh & grounded | Jev · DecisioningTyped & calibrated |
|---|---|---|
| What it does | Real-time open web — today’s advisory, not last year’s training data | Four decisions in one parallel call, in milliseconds |
| What it guarantees | Every signal carries a source and a date | Honest confidence on every answer |
| What it won’t do | Supplies the state; never guesses it | Can’t leave the schema — no hallucinated actions |
The instinct is to throw one big frontier LLM at this. It demos well and fails in production: slow and costly at ten-thousand-a-week, knowledge frozen at training time, and you still have to parse a decision out of prose.
Jev is built for the opposite. You hand it state and typed questions; it returns typed decisions with calibrated probabilities, in one parallel pass, and generates no text at all. The choice we most admire as engineers is what it deliberately gives up: no free-form output, no tools. Because it can’t emit a value outside your schema, a whole class of failure — malformed output, hallucinated fields, rogue tool calls — is simply off the table.
That constraint is the feature. It’s what makes Jev fast (70–500 ms), cheap ($0.042 / 1M input tokens, output free), and — the part that matters most here — honest.
And that same choice is why it needs a partner. A decision model with no senses is only ever as current as the state you hand it. TypeSafe is direct about it: if a fact isn’t in the state, Jev can’t use it. That’s the seam we fill. You.com is the senses.
The decision call is small enough to read at a glance — four questions, one round trip:
# You.com built `advisory`; STACK describes what you run answers = jev.system_one( state={"advisory": advisory, "our_stack": STACK}, questions={ "affects_us": noul("An affected component is in `our_stack`"), "severity": score("Severity for a team running `our_stack`", ["informational","low","medium","high","critical"]), "exploited": noul("Actively exploited, or a public exploit exists"), "action": choice("Recommended action", ACTIONS), }, ) # answers["severity"] → {score, confidence, probabilities} — nothing to parse
Confidence is the whole game
The judgments come from Jev; the policy stays in your code, where “what counts as urgent for us” belongs. I only auto-page when the action is patch_now, severity is high or above, and confidence clears 0.90.
That last clause earns its keep. On our first live run, You.com surfaced a real critical Redis advisory. Jev said it affects the stack (0.96) and recommended patching now — but its confidence in that action sat just under 0.90, so it landed in the review queue instead of paging someone at 2 a.m.
A tool that flattens every decision to a boolean can’t make that distinction. A calibrated one treats a 0.7 and a 0.98 differently — which is the entire point.
Read everything, show almost nothing
The machine reads the whole firehose; a person sees only the short list that survives. Full coverage, almost none of the noise.
- Every advisory for your stackretrieved live
- affects_us ≥ 0.5most dropped automatically
- Review queuea human glances
- Patch nowpaged
And every advisory, every source, and every decision — probabilities and confidence included — is written to an append-only audit log. When someone asks in three months why did we patch that and not this, the answer is a line in a file, not a shrug.
Under the hood
The whole system is small — Python standard library, no framework — with clean seams between retrieval, decisioning, and the plumbing around them. The data flows straight down:
affects_us · severity · exploited · action — answered in one parallel call, each with a calibrated confidence. Nothing to parse.alert · review · monitor, write a dated report and an append-only audit log, and notify Slack or email — on a schedule.Every boundary where outside data enters is parsed defensively and escaped — a malformed reply degrades gracefully, it never takes down a 3 a.m. run.
Retrieval isn't perfect: verify a CVE id against NVD before you act. And Jev's calibration is a property of groups of predictions, not a promise about any single one — keep a person on the patch_now calls.
The agent's job isn't to be an oracle. It's to make sure that person is looking at the right six things instead of all six hundred.
A building block, not an integration.
The use case is CVE triage, but the shape is everywhere: any decision made at volume, against fresh reality, over a knowable set of choices. Is this lead worth a call? Does this transaction look like fraud? Did the service we depend on just change?
Each is the same pipeline with different questions — retrieve with You.com, judge with Jev — and the two halves stay cleanly separated, so each can improve without breaking the other.
That's why it felt worth writing up: not an integration, but a pattern — one we expect builders to point at problems we haven't thought of yet. This first agent is open source and MIT-licensed precisely so you can take it apart and do exactly that.
→ github.com/youdotcom-oss/cve-triage-agent
Grab a You.com key and a TypeSafe API key, point it at your stack, and tell us what you build.
Sources: Jerry Gamblin CVE reviews | Cyentia/FIRST | NIST NVD | FIRST CVSS & EPSS | Mandiant M-Trends.
LI Test
LI Test
Share Article:
Related resources.

Pydantic AI and You.com—Delivering Powerful Web Intelligence and Agent Orchestration
August 27, 2026
Blog

The Model Is the Cheapest Part of a Grounded Answer
August 13, 2026
Blog
.png)
Real-Time Web Intelligence for Autonomous Agents: You.com × Sapiom
August 7, 2026
Blog

Your Trading Agent Should Read Before It Buys: Accessing You.com Over x402 on Base
August 4, 2026
Blog

Every Model, Every Agent, the Live Web: The You.com MCP Server Comes to Warp
July 27, 2026
Blog
