Detect what rules miss, from the first sequence
LogLM is a patent-pending encoder-only foundation model pretrained on security logs and telemetry. Security teams run it inside their own environment to find concerning sequences zero-shot: no baselining period, no labeling project, no telemetry leaving the building.
Definitions and test conditions: Results
The right model for the volume
Large language models generate text one token at a time, priced per token. LogLM does something narrower and faster: it reads sequences of telemetry and produces a dense representation of behavior over time.
That design decides the economics. An encoder scores the full stream on modest hardware; a reasoning model is worth its cost on the handful of cases that are rare and semantically strange. Teams run LogLM on the volume and reserve reasoning models in Vigil for the escalations.
From raw telemetry to a finding
Normalize
Flow, DNS, proxy, identity, cloud, and endpoint telemetry is normalized automatically through a scale-out pipeline.
Sequence
Events become behavioral sequences between entities over time: what followed what, at what cadence, in whose company.
Embed
The encoder maps each sequence to a high-dimensional behavior embedding.
Classify
Purpose-built classifiers read the embeddings and emit compound detections: malicious, not merely unusual.
Zero-shot on arrival
LogLM was pretrained across many environments and telemetry types, so it generalizes on arrival. There is no per-customer baseline to build and no learning period to wait out.
Where an environment warrants it, your team adapts the classifiers inside your own boundary, in minutes, without retraining the base model.
Findings start with the first sequences scored.
Pretraining is self-supervised; your analysts do not build a training set.
Runs on infrastructure you already operate. Many deployments need no GPU.
Score historical telemetry in batches or the live stream continuously.
Findings your SOC can act on
Each detection arrives as structured JSON: the entities involved, the time window, the sequence that triggered it, a confidence score, and the MITRE ATT&CK tactics and techniques it maps to.
Your SIEM, your SOAR, or Vigil consumes it directly. No parser project, no proprietary console between your analysts and the evidence.
{
"finding_id": "f-20261003-7c2e",
"verdict": "malicious",
"confidence": 0.97,
"entities": {
"src": "10.4.12.37",
"dst": "185.220.101.14"
},
"window": "14:02:11Z/14:19:40Z",
"mitre": [
{ "tactic": "TA0011",
"technique": "T1071.001" }
],
"sequence_ref": "seq-88f1"
}Behavior you can search
Detect now
Classifiers score sequences as they arrive and emit findings in near real time. Measured mean time to detect is under one minute at Deutsche Telekom and a top-four global bank.
Search later
Embeddings persist beside your telemetry. When a new report lands, hunters search by behavior rather than by indicator: find sequences that resemble this one across months of history. Indicators rotate; technique persists.
Runs where your telemetry already lives
On premises or fully air-gapped, inside your data lake, or in your own cloud or Kubernetes. There is no hosted tier. Telemetry, model weights, and findings stay inside your boundary.
In your data lake
Score sequences where the data rests, in Snowflake, Databricks, or object storage. A purpose-built Spark pipeline scales ingestion horizontally.
In your Cribl pipeline
Detect in flight, then route findings to the SIEM and bulk telemetry to low-cost storage.
Upstream of the SIEM
Splunk, Sentinel, QRadar, and Elastic keep running and receive findings instead of raw volume. That shift is the source of SIEM savings of up to 45%.
See LogLM on your telemetry
Send us a sample to assess, or run the assessment on premises so nothing leaves your environment. Either way you receive a report of what your current stack missed, with the evidence behind each finding.
