Results

Measured, defined, and linked

Each figure on this site traces to an evaluation, with the metric defined and the test conditions stated. Where a customer asked us not to name them, we describe the environment instead. Where a benchmark is public, we link the data so you can reproduce the result.

99.3%

Recall on malicious activity

Deutsche Telekom, Phase II

99 F1

At a 0.08% false positive rate

Zero-shot, a very large internet service provider

<1 min

Mean time to detect

Deutsche Telekom and a top-four global bank

Up to 45%

Lower SIEM cost

LogLM upstream of the SIEM

Evaluations

Customer and partner evaluations

Evaluation

Environment

Metric

Result

Conditions

Deutsche Telekom, Phase II

Telecom network

Recall on malicious activity; mean time to detect

99.3%; under 1 minute

Production telemetry inside the customer environment; classifiers adapted in minutes

A top-four global bank

Financial services

False negative rate; mean time to detect

0.7%; under 1 minute

Zero-shot; compared against the bank's own bespoke models for each attack family

BNY

Financial services

False negative rate; mean time to detect

0.9%; under 3 minutes

Customer telemetry; BNY is DeepTempo's first design partner and co-developer of the encoder work

A mega CDN vendor

Internet infrastructure

False positive rate; mean time to detect

0.08%; 1 minute

Zero-shot, with no customer-specific training

Technology Advancement Center

OT water-plant range, Modbus TCP

Executed attacks detected

100%, zero-shot

Passive monitoring on premises; the TAC is the NSA-established OT cyber center

Public benchmarks

Results anyone can reproduce

CIC dataset, zero-shot

LogLM scored NF-CSE-CIC-IDS2018-v2, the NetFlow edition of the Canadian Institute for Cybersecurity benchmark, with no exposure to the dataset before inference and no tuning: 8.4 million flows across seven attack scenarios.

F1 (attack)

96.9%

Precision (attack)

99.4%

Recall (attack)

94.5%

False positive rate

0.085%

SOCBench

SOCBench is the open benchmark for AI in security operations. It scores any detection stack, ours included, on precision, recall, false positives, MITRE ATT&CK coverage, cost, and drift. Methods and data are published so practitioners can rerun the comparison themselves.

Definitions and test conditions

What the numbers mean

Recall

The share of malicious activity the system detected: true positives divided by true positives plus false negatives.

False negative rate

The share of malicious activity the system missed. Recall and false negative rate sum to 100%.

False positive rate

The share of benign activity flagged as malicious: false positives divided by false positives plus true negatives.

F1

The harmonic mean of precision and recall, reported on the attack class.

Mean time to detect

Elapsed time from malicious activity appearing in telemetry to a finding emitted by LogLM.

Zero-shot

Scored with no customer-specific training, labels, or baselining before inference.

SIEM cost

The reduction in SIEM ingest and licensing when LogLM runs upstream and bulk telemetry routes to lower-cost storage while findings flow to the SIEM. Savings depend on log mix and SIEM pricing; up to 45% is the ceiling observed, not a guarantee.

Conditions

Customer evaluations ran on the customer's own telemetry, in the customer's environment or on a sample provided under agreement, and were measured by the customer or jointly. Each result describes one evaluation and is not a guarantee for another environment.

Your telemetry, your numbers

The only evaluation that settles the question is one on your own data. Send us a sample to assess, or run the assessment on premises so nothing leaves your environment.