LogLM

Detect what rules miss, from the first sequence

LogLM is a patent-pending encoder-only foundation model pretrained on security logs and telemetry. Security teams run it inside their own environment to find concerning sequences zero-shot: no baselining period, no labeling project, no telemetry leaving the building.

99.3%
Recall on malicious activity, Deutsche Telekom Phase II evaluation
0.08%
False positive rate, zero-shot, in production evaluation
<1 min
Mean time to detect at Deutsche Telekom and a top-four global bank

Definitions and test conditions: Results

Encoder, not chatbot

The right model for the volume

Large language models generate text one token at a time, priced per token. LogLM does something narrower and faster: it reads sequences of telemetry and produces a dense representation of behavior over time.

That design decides the economics. An encoder scores the full stream on modest hardware; a reasoning model is worth its cost on the handful of cases that are rare and semantically strange. Teams run LogLM on the volume and reserve reasoning models in Vigil for the escalations.

 
LogLM (encoder)
General LLM (decoder)
Job
Score every sequence
Reason over escalated cases
Cost profile
Fixed compute, no per-token fees
Per token, rising with volume
Output
Embedding plus MITRE-mapped finding
Prose
Zero-shot detection accuracy
99% in production evaluations
Not built for raw telemetry
Where it runs
Inside your data path
Often an external API
How it works

From raw telemetry to a finding

1

Normalize

Flow, DNS, proxy, identity, cloud, and endpoint telemetry is normalized automatically through a scale-out pipeline.

2

Sequence

Events become behavioral sequences between entities over time: what followed what, at what cadence, in whose company.

3

Embed

The encoder maps each sequence to a high-dimensional behavior embedding.

4

Classify

Purpose-built classifiers read the embeddings and emit compound detections: malicious, not merely unusual.

No adaptation period

Zero-shot on arrival

LogLM was pretrained across many environments and telemetry types, so it generalizes on arrival. There is no per-customer baseline to build and no learning period to wait out.

Where an environment warrants it, your team adapts the classifiers inside your own boundary, in minutes, without retraining the base model.

No baselining period

Findings start with the first sequences scored.

No labeled data

Pretraining is self-supervised; your analysts do not build a training set.

No special hardware

Runs on infrastructure you already operate. Many deployments need no GPU.

Batch or stream

Score historical telemetry in batches or the live stream continuously.

Output

Findings your SOC can act on

Each detection arrives as structured JSON: the entities involved, the time window, the sequence that triggered it, a confidence score, and the MITRE ATT&CK tactics and techniques it maps to.

Your SIEM, your SOAR, or Vigil consumes it directly. No parser project, no proprietary console between your analysts and the evidence.

Example findingJSON
{
  "finding_id": "f-20261003-7c2e",
  "verdict": "malicious",
  "confidence": 0.97,
  "entities": {
    "src": "10.4.12.37",
    "dst": "185.220.101.14"
  },
  "window": "14:02:11Z/14:19:40Z",
  "mitre": [
    { "tactic": "TA0011",
      "technique": "T1071.001" }
  ],
  "sequence_ref": "seq-88f1"
}
Detect now, search later

Behavior you can search

Detect now

Classifiers score sequences as they arrive and emit findings in near real time. Measured mean time to detect is under one minute at Deutsche Telekom and a top-four global bank.

Search later

Embeddings persist beside your telemetry. When a new report lands, hunters search by behavior rather than by indicator: find sequences that resemble this one across months of history. Indicators rotate; technique persists.

Deployment

Runs where your telemetry already lives

On premises or fully air-gapped, inside your data lake, or in your own cloud or Kubernetes. There is no hosted tier. Telemetry, model weights, and findings stay inside your boundary.

In your data lake

Score sequences where the data rests, in Snowflake, Databricks, or object storage. A purpose-built Spark pipeline scales ingestion horizontally.

In your Cribl pipeline

Detect in flight, then route findings to the SIEM and bulk telemetry to low-cost storage.

Upstream of the SIEM

Splunk, Sentinel, QRadar, and Elastic keep running and receive findings instead of raw volume. That shift is the source of SIEM savings of up to 45%.

See LogLM on your telemetry

Send us a sample to assess, or run the assessment on premises so nothing leaves your environment. Either way you receive a report of what your current stack missed, with the evidence behind each finding.