Buyer's guide

How to evaluate an AI SOC: ten questions to ask before you buy

An AI SOC is a bet on how your security operations will run for years. Most evaluations reward the best demo. This guide rewards the architecture: what grounds the detections, whether your team can read the code, how autonomy is earned, and what evidence the vendor can show. Use it with any vendor, including us.

Ask about the detections

A chat layer over your existing alerts inherits every gap in those alerts. Ask what detects, not only what summarizes.

Ask to read the code

If analysts cannot read an agent or a playbook, they cannot trust it, extend it, or audit what it did.

Ask for evidence

Precision, recall and false positive rates on your telemetry, measured with a method you can repeat.

The trap

Most AI SOC evaluations measure the demo, not the system.

Three patterns produce confident purchases and disappointing years. Each is easy to avoid once named.

The scripted demo

A vendor replays a known incident through a polished console. That shows the interface, not the detection. Ask to run the system on a sample of your own telemetry, including the incidents you did not know about.

Triage without detection

Many products reason over alerts your SIEM and EDR already raise. They can speed up triage, but they cannot find what rules missed. Separate the two capabilities and score each on its own.

Vendor-reported metrics

Percentages without definitions are marketing. Ask what was measured, on which data, over what period, and whether the method is published so you can repeat it.

Questions 1 to 5

Detection and code

The first five questions decide whether you are buying intelligence or a user interface.

1. What grounds the detections?

Ask whether detections come from a model purpose-built for security telemetry, from rules, or from a general-purpose LLM prompted over alerts. A strong answer names the model, what it was pretrained on, and how it is evaluated.

2. Does it work before tuning?

Ask for results on day one, with no labels and no baselining period. A strong answer gives zero-shot figures and shows them on your data, not a reference dataset.

3. Can our analysts read the code?

Ask whether agents, playbooks and integrations are open source under a license you can audit, or closed. A free download is not open source. A strong answer is a repository and a license.

4. What happens to our existing detections?

You have years of rules in Splunk, Elastic or Sentinel. Ask whether the product runs them, manages them and measures their coverage, or merely consumes their output as alerts.

5. Which language model, and whose?

Ask whether you can bring your own LLM, local or hosted, under your keys and policies. A strong answer lets you choose the least powerful model that does each job.

Questions 6 to 10

Autonomy, evidence and control

The second five decide whether you keep control as the system takes on more of the work.

6. How is autonomy earned?

Ask what the system may do on its own, how that permission expands, and which gates remain. A strong answer describes declared permissions, confidence thresholds, cost ceilings and human approval points per workflow.

7. Where does it run, and what leaves?

Ask whether telemetry, model weights and verdicts stay inside your environment, and whether on-premises or air-gapped deployment is supported. Private cloud often means the vendor's cloud.

8. Who owns the learning loop?

When the system improves from your analysts' decisions, ask where that improvement lives. A strong answer keeps it in your environment, under your control and exportable.

9. How do you prove it works?

Ask for precision, recall, false positives, MITRE ATT&CK coverage and cost, measured with a method you can repeat. An open benchmark is stronger than a vendor report.

10. What does it cost at scale?

Token spend, per-alert pricing and ingestion fees compound. Ask for a cost model at your volume, including a quiet week and a loud one.

Scorecard

Score each answer from 0 to 2. The pattern matters more than any single row.

Dimension
Weak answer
Strong answer
Detection grounding
Weak:General LLM prompted over existing alerts
Strong:Purpose-built model pretrained on security telemetry, evaluated and published
Day-one performance
Weak:Weeks of baselining or tuning before results
Strong:Zero-shot results shown on your own telemetry
Code
Weak:Closed, or a free download without source
Strong:Open source under an auditable license
Existing detections
Weak:Consumed as alerts
Strong:Run, managed and measured for coverage
Language model
Weak:Vendor-selected and vendor-hosted
Strong:Your choice, local or hosted, under your keys
Autonomy
Weak:A toggle per workflow
Strong:Earned per skill with permissions, thresholds and human gates
Deployment
Weak:Vendor SaaS only
Strong:On premises, air-gapped, your data lake or your own cloud
Learning loop
Weak:Improves the vendor's model
Strong:Stays in your environment and belongs to you
Evidence
Weak:Vendor-reported percentages
Strong:Open benchmark with a repeatable method
Cost model
Weak:Per alert or per token, unbounded
Strong:Predictable at your production volume

A product scoring under 12 of 20 is a triage assistant. That can still be worth buying, but price it and position it as one.

Evaluation plan

Run a 30-day evaluation on your own telemetry.

Thirty days is enough to answer the ten questions with evidence rather than slides. Insist on your data, not a reference set.

Week 1: Sample and baseline

Choose a representative telemetry sample, including a period with a known incident if you have one. Record what your current stack alerted on during that period.

Week 2: Detect

Run the candidate on the sample with no tuning. Count what it found that your stack missed, and what it raised that you would dismiss. Those are recall and false positives on your data.

Week 3: Investigate

Hand the findings to analysts inside the product. Measure time to a decision, the quality of the evidence presented, and whether analysts can read and change the workflow that produced it.

Week 4: Decide

Score the ten questions. Compare the cost model at production volume. Ask for the evaluation method in writing so you can repeat it next year, with this vendor or another.

Hold us to the same ten questions.

Vigil is the leading open source AI SOC, Apache 2.0, and runs the detections you already trust. LogLM is a foundation model pretrained on security telemetry: 99% zero-shot, 1% or fewer false positives. SOCBench is the open benchmark we publish. Start with Vigil, or run a free detection assessment on your own telemetry.