Eleven days ago TypeSafe AI came out of stealth with a model called Jev and a new category name to go with it. Founder Diogo Almeida helped build the methods at OpenAI that became the research behind ChatGPT, and his team built a new model architecture, a parallel sampler, and a training method they call Reinforcement Learning for Calibrated Decisions.
The idea is elegant. You send Jev program state plus typed questions, and you get all the answers back in one parallel pass in 70 to 500 milliseconds, at $0.042 per million input tokens with output free. TypeSafe calls this a System One model, after Kahneman. I think they have the taxonomy right, and Jev can play important roles as a part of machine speed intelligence for cyber.
That said, while Jev may well help your detection engineering and a long list of security workflows, Jev cannot do your detecting. Detection needs its own System One, one that arrives already fluent in telemetry and that runs at full volume inside your environment.
That is LogLM.
Precedent
Two examples of machine System One AI in production are probably pretty familiar to everyone.
Waymo splits its stack along Kahneman's line. A pretrained sensor fusion encoder fuses camera, lidar, and radar over time and produces embeddings for fast reactions, while a driving VLM handles complex semantic reasoning. Waymo cared enough about the fast half to build silicon for it; its Hot Chips 2026 talk described a processor designed to run the sensor fusion encoder, with the VLM reserved for tasks that tolerate more latency.
Stripe did the same for payments. Its engineers built a self-supervised foundation model, trained on tens of billions of transactions, that distills each charge into a single embedding. A classifier ingests sequences of those embeddings and decides whether a slice of traffic is under attack. Card-testing detection on large users jumped from 59% to 97%, on the critical path, in under 100 milliseconds.
Same pattern twice: an encoder over native signal, small classifiers on the embeddings, and a reasoning model downstream for the hard, slow, rare cases. That is also our Intelligent Defense Platform: LogLM is a self-supervised foundation model trained on billions of logs, and downstream work is facilitated, controlled, and improved over time with the help of Vigil plus whatever AI you prefer.
As will be discussed, Jev and Vigil are like peanut butter and chocolate. (I’m hungry as I write this)
Train me on Pretraining
An encoder foundation model like LogLM earns its keep before anyone asks it a question. The recipe starts with an enormous corpus of unlabeled native data. Hide part of each sequence and train the model to reconstruct what is missing from the context on both sides. No labels, no analysts, no rules. To succeed, the model must internalize the grammar of the domain: which tokens travel together, which orderings recur, which rhythms are ordinary. The result is an embedding space in which typical behavior occupies dense, well-mapped regions and a novel sequence lands somewhere sparse and conspicuous. Stripe's Emily Sands describes a card tester's two hundred nearly identical requests lighting up as an island in that space. The island exists before anyone trains a classifier to find it.
We pretrain LogLM this way on cybersecurity telemetry drawn from diverse environments and telemetry types with the help of design partners including BNY. LogLM therefore arrives already knowing what ordinary telemetry looks like across many environments, and what a concerning sequence looks like against that backdrop. That is why LogLM detects at 99% zero-shot with 1% or fewer false positives, with no baselining period, and why purpose-built classifiers on its embeddings yield compound detections. Instead of flagging anomalies as incidents, a LogLM also determines whether those unusual behaviors are also similar to known attacks.
Incidentally - this is why the ramp up of agentic traffic in many environments is not causing a flood of false positives to LogLM users. In our research we do see new areas of embedding space being populated however they are not, in the vast majority of cases, malicious. These patterns are agentic Evan doing the job differently than meatbag Evan, and with no more ill intent than old grumpy.
Jev arrives with something different. TypeSafe trained it to make calibrated decisions over whatever state you supply, across many domains. We are already seeing some cool uses of Jev by security researchers. For example, see the jev-ids project, available here and authored by Paulo Severo, Silvio Quincozes and Amanda Dias. The workflow is simple: jev-ids prompts Jev with a handful of labeled flows with each request, and the quality of those labels largely determines whether Jev returns a useful and accurate classification. Jev's own model card lists hex and numeric values, dates, and counting as weak spots. Welp - that’s pretty much what telemetry is, for example IPs, ports, hashes, timestamps, and counts. The key point is that the pretraining and the architecture are what determine the efficacy, and Jev is not purpose built to identify hard to see attacks, unlike our LogLM.
Quick Descriptions

Encoders. Bidirectional attention over native tokens, self-supervised pretraining, and an output that is a vector rather than a sentence. Purpose-built classifiers sit on top. LogLM runs on premises, air-gapped, or in your own cloud, in batch or as a continuous stream.
Reasoning models. Autoregressive decoders pretrained on text, then post-trained for human preference and verifiable rewards. Chain of thought, tool use, long-running agents. They investigate, hunt, write and test detection logic, and explain themselves. They are also slow per decision; TypeSafe cites 3 to 329 seconds end to end for frontier models, and agentic investigations run far longer. Vigil users bring their own.
Jev. A single endpoint takes state and a map of questions, and a choice question supports up to 255 options. RLCD trains for calibrated probabilities rather than pleasing prose. The input is serialized program state, and someone decides what goes into that state and which question to ask.
Records and sequences
That last sentence is worth another read. Detection is the act of deciding what warrants investigation. Attacks live in sequences: a beacon cadence, a credential used from a host that has never touched that service, a slow exfiltration spread across a week of DNS. A per-record question carries that context only if someone packages it first, and packaging it is the detection.
Another quick point. Tokenomics are improved by Jev but not fixed. The open source jev-ids project packs roughly 1,800 tokens into each flow request, about $74 per million verdicts. A large enterprise can easily emit a billion flow records in a day. That comes to $74,000 a day, about $27 million a year, before identity, endpoint, DNS, and cloud logs, at a sustained 11,600 calls per second against a hosted API. Compare this to the LogLM which can handle that load on a single modern GPU.
The jev-ids authors see this division of tasks clearly; they see Jev as an asynchronous side channel fed from a queue, judging alerts or sampled flows, never as an inline filter. Right instinct. It also concedes the point: something upstream has to choose what to send.
Evidence
Security practitioners moved fast, which I love. What they have found so far:
- jev-ids on NSL-KDD. Mentioned above. At one example per category, Jev posted an F1 of 0.859 against 0.776 for GPT-5.6, and raised 24 false alarms on 420 benign flows where a Random Forest raised 362. The authors flag the limits themselves: pilot numbers, a rate-limited gateway where 1,350 of 1,800 Jev rows needed a retry, and flow features that leave your network on each request. The numbers have also moved between revisions; one version reports the LLM ahead on F1. https://github.com/jev-ids/jev-ids
- Cribl. Its team reports that Jev held above 92% agreement with a committee of LLM judges at roughly 1% of the cost on tasks such as log type identification. https://cribl.io/blog/what-typesafes-jev-means-for-telemetry/
- Ken Huang. He wired Jev into an agentic SOC pipeline in which typed answers gate containment before any language model writes the analyst brief. Details are only behind a paywall. https://kenhuangus.substack.com/p/jev-returns-typed-probabilities-at
All of the use cases for cyber mentioned, including triage, routing, parser selection, guardrails, and response gating, are downstream of detections and are in support of the analyst. That is where Jev belongs, not upstream in the detection stream itself.
Harness
Jev returns a probability. Someone still has to decide what 0.72 means at three in the morning. Thresholds, approvals, escalation, audit, and rollback belong to a harness, with humans in the loop for consequential actions and for overall system management and on the loop for the rest. Jev does not replace that harness. Vigil is the leading open source one, skills-native, with declared intent and policy stored as Markdown under change control and a console for the analysts who supervise it. Inside Vigil, a typed verdict becomes one more governed signal rather than an ungoverned action.
Vigil
We expect one or more contributors to bring Jev into Vigil, and we would welcome it. If you want to be first, please reach out to one of the project leads and/or attend the weekly office hours or just YOLO a PR; you can even just drop an issue or two into the Vigil Factory and see if it’ll do the coding for you. One possible shape: a typed-decision skill behind a stable contract with a pluggable backend, so a team can point it at Jev or at a local model, much as triagedy does. Good first jobs include ATT&CK tagging of detection rules, alert scoring, case routing, parser selection, and guarding agent actions. Vigil already supports OpenRouter, and OpenRouter now serves Jev, so the first contributor is one API key and a small client away.
We will get there ourselves. Right now, though, 0.6 is out and 1.0 is around the corner, and our users want open source running on premises, air-gapped, or in their own cloud. Jev today is a hosted API with an undisclosed architecture and no weights, and the users we are closest with have no plans to ship telemetry to it.
Contribute
Now is the time to write your name on the side of the Vigil rocket as it clears the pad heading towards safe, scalable, intelligent autonomic cybersecurity. Build a Vigil integration for Jev, publish an eval, and so on.
TL;DR:
Jev might help your detection engineering. It won't do your detecting. For that you need a System One that already knows your telemetry. And for downstream tasks you need something like VigilSoc.org as your cyber harness more than ever.
