one flow.
one request.
one verdict.
An intrusion detection system that asks TypeSafe's Jev two typed questions about one network flow and gets a verdict back, with no text to parse.
Put to the test on NSL-KDD, the classic intrusion detection benchmark, against an LLM (Gemini 3.6 Flash), a Random Forest and an Isolation Forest: 2,000 flows, three seeds, k from 0 to 8.
-
7.7×
faster per verdict
than Gemini 3.6 Flash -
22×
cheaper per verdict
than Gemini 3.6 Flash -
0.95
precision, the same
as Gemini 3.6 Flash -
18×
fewer false alarms
than a Random Forest
The LLM is the better detector on F1, 0.880 against 0.856 at k = 1 and ahead at every k. Jev is second, at a fraction of the time and cost, and catches more zero-day attacks than the LLM from k = 1 to k = 8.
input
flow one connection record from NSL-KDD, plus k labeled examples
request: two typed questions to jev
is_attack noul Is the connection an intrusion attempt?
category choice Which category does it belong to?
output
is_attack → 0.82
category → dos, confidence 0.84
verdict → attack p_attack ≥ 0.5, the same cut for every detector
the request
Jev IDS runs on a new kind of AI, the System One Model (SOM). Instead of generating text, a SOM takes a state and standardized questions about it and answers each one with a probability. Jev IDS hands it one monitored network flow and gets back one verdict, attack or normal, with the probability behind it.
-
input
one flow
The monitoring record of a single connection: duration, protocol, service, bytes in and out.
-
jev
one request
State: the instructions, a few labeled examples and the flow. Questions:
is_attackandcategory. -
output
one verdict
attack · p_attack 0.82
dos · confidence 0.84
Everything Jev is told lives in one versioned file,
prompts/nsl-kdd/jev.json. Python adds only the flow and the examples, and the file's sha256
rides with every prediction, so any verdict traces back to the exact
prompt. Examples are labeled by category only; attack names never
reach a model. Both questions are answered in parallel over the same
state, in one round trip.
Three baselines follow the same protocol: an LLM agent with a JSON output schema, a scikit-learn Random Forest trained on the same examples, and an Isolation Forest fitted on benign traffic alone. Same frozen split, same examples, same seeds.
evidence
Paper split of NSL-KDD: 2,000 flows, 1,126 of them attacks and 300 of those zero-day, of a kind absent from the example pool. Three seeds of examples, so 6,000 predictions per detector and k. Means over the three seeds, from the runs of 2026-09-22; the full tables and figures are in docs/results.md.
| detector | k | f1 | precision | recall | zero-day recall | false alarms / 874 | latency | cost / 1M flows |
|---|---|---|---|---|---|---|---|---|
| jev (jev-1.13.0) | 1 | 0.856 | 0.953 | 0.778 | 0.747 | 43 | 315 ms | $74 |
| gemini 3.6 flash (gemini-3.6-flash) | 1 | 0.880 | 0.942 | 0.826 | 0.713 | 57 | 2,421 ms | $1,651 |
| random forest (100 trees) | 1 | 0.748 | 0.598 | 1.000 | 1.000 | 764 | 3 ms | local |
| jev | 8 | 0.854 | 0.942 | 0.783 | 0.721 | 56 | 335 ms | $295 |
| gemini 3.6 flash | 8 | 0.881 | 0.949 | 0.823 | 0.702 | 50 | 1,991 ms | $3,035 |
| random forest | 8 | 0.865 | 0.794 | 0.950 | 0.912 | 278 | 3 ms | local |
| random forest (whole pool) | all | 0.765 | 0.971 | 0.631 | 0.263 | 21 | 3 ms | local |
| isolation forest (whole pool) | all | 0.761 | 0.974 | 0.624 | 0.573 | 19 | 3 ms | local |
Jev against Gemini over all 6,000 paired verdicts at k = 1: 508 differ, Jev is right in 195 and Gemini in 313 (McNemar p < 0.001); Gemini leads at every k. On the 900 zero-day attacks at k = 1: 118 differ, Jev is right in 74 and Gemini in 44 (p = 0.007); Jev also leads at k = 2, and the two are indistinguishable at k = 4 and 8. Jev against the Random Forest over all flows at k = 1: 2,909 differ, Jev is right in 2,161 (p < 0.001). A forest trained on five rows calls almost everything an attack, which is why its recall is perfect and its precision is not: on the 874 benign flows it raised 764 false alarms, Jev 43.
limits worth knowing
- Cost is tokens times list price, not what was billed: $0.042 per million input tokens for Jev, output free; $0.75 input and $3.75 output for Gemini through 2026-12-31, doubling on 2027-01-01, with thinking tokens billed as output.
- Gemini's latency was measured with five processes in parallel on a congested day (transient 429, 500 and 504 answers, all repaired), so it is a real-day figure, not a floor. Jev's is the wall clock of one direct call to TypeSafe's API.
- Gemini 3.x cannot switch thinking off; it ran at the lowest level, about 200 reasoning tokens per flow.
- Every number is NSL-KDD's. NF-UQ-NIDS-v2 has a card and a preparation script and no run yet.
-
Jev's
noulanswer carries no confidence; only thechoiceanswer does.
try it
git clone https://github.com/jev-sec/jev-ids.git
cd jev-ids
uv sync
.env: TYPESAFE_API_KEY for Jev. Gemini: gcloud auth application-default login,
GOOGLE_GENAI_USE_VERTEXAI=true, GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATION=global.
Download
NSL-KDD
into data/raw/nsl-kdd/ and prepare it once.
uv run python -m scripts.prepare_nsl_kdd
uv run jev-ids run --dataset data/nsl-kdd/dataset.json --detector jev --split smoke --k 0,1
The smoke split is five flows. The run writes
results/<timestamp>-nsl-kdd-jev-smoke/ with
config.json, one JSON row per flow in
predictions.jsonl and the raw answers in
responses.jsonl.
bring your own flows
Jev IDS is an independent research prototype, not a product. Its numbers come from one benchmark, NSL-KDD. It can still sit inside a commercial solution, and this is how.
where jev fits
A commercial IDS already has sensors, a flow exporter and a signature engine feeding a SIEM. Jev IDS replaces none of them. It sits beside the pipeline, off the packet path, and judges one flow record at a time.
-
Second opinion on alerts. Send Jev the flow behind
each alert the signature engine raised.
p_attackranks the queue, and the SOC reads the top first. On the paper split at k = 1, Jev raised 43 false alarms on 874 benign flows where a Random Forest raised 764. - A net behind the signatures. Signatures miss what they have never seen. Sample the flows the engine passed as clean, or every flow to a critical asset, and let Jev judge them: it caught 75% of the zero-day attacks at k = 1, where the LLM caught 71%.
-
A category for the playbook. The
choiceanswer names the category with a confidence, so the SIEM routes dos, probe, r2l and u2r, or your own taxonomy, to different runbooks with no parser in between. - Coverage from day one. A new site or tenant has no training set. Jev needs one labeled flow per category, so it covers the segment while a classical model is still collecting data.
Half a second per verdict and a rate-limited gateway make this an
asynchronous side channel, fed from the exporter (NetFlow, IPFIX, Zeek
conn.log) through a queue, never an inline filter.
six steps
-
Get a key. Jev is served by TypeSafe; set
TYPESAFE_API_KEYin.env. Pricing and terms are TypeSafe's. -
Describe your flows. Write a dataset card like
data/nsl-kdd/dataset.json: the columns your flow exporter emits, in order, which of them are symbolic, your categories and which one is benign. -
Write the request. Copy
prompts/nsl-kdd/jev.json, replacecolumnsand the category descriptions with yours, and keep the two questions. - Pick examples. One labeled flow per category from your own network is enough to start; k = 1 is where the benchmark's F1 reaches its plateau.
-
Call the detector.
JevDetector(load_prompt(path)).predict(flow, examples)returnsp_attack,category_predandconfidencefor one flow in about half a second. Routep_attackto your alerting with the cut your alarm budget allows; 0.5 was the benchmark's choice, not a rule. -
Measure before you trust. Run
jev-ids runandjev-ids metricson a labeled split of your own flows. The numbers above are NSL-KDD's, not yours.
Only flow features leave your network, never payloads, but they do leave it: every request goes to the gateway.