How to evaluate an AI SOC: ten questions to ask before you buy
An AI SOC is a bet on how your security operations will run for years. Most evaluations reward the best demo. This guide rewards the architecture: what grounds the detections, whether your team can read the code, how autonomy is earned, and what evidence the vendor can show. Use it with any vendor, including us.
A chat layer over your existing alerts inherits every gap in those alerts. Ask what detects, not only what summarizes.
If analysts cannot read an agent or a playbook, they cannot trust it, extend it, or audit what it did.
Precision, recall and false positive rates on your telemetry, measured with a method you can repeat.
Most AI SOC evaluations measure the demo, not the system.
Three patterns produce confident purchases and disappointing years. Each is easy to avoid once named.
A vendor replays a known incident through a polished console. That shows the interface, not the detection. Ask to run the system on a sample of your own telemetry, including the incidents you did not know about.
Many products reason over alerts your SIEM and EDR already raise. They can speed up triage, but they cannot find what rules missed. Separate the two capabilities and score each on its own.
Percentages without definitions are marketing. Ask what was measured, on which data, over what period, and whether the method is published so you can repeat it.
Detection and code
The first five questions decide whether you are buying intelligence or a user interface.
Ask whether detections come from a model purpose-built for security telemetry, from rules, or from a general-purpose LLM prompted over alerts. A strong answer names the model, what it was pretrained on, and how it is evaluated.
Ask for results on day one, with no labels and no baselining period. A strong answer gives zero-shot figures and shows them on your data, not a reference dataset.
Ask whether agents, playbooks and integrations are open source under a license you can audit, or closed. A free download is not open source. A strong answer is a repository and a license.
You have years of rules in Splunk, Elastic or Sentinel. Ask whether the product runs them, manages them and measures their coverage, or merely consumes their output as alerts.
Ask whether you can bring your own LLM, local or hosted, under your keys and policies. A strong answer lets you choose the least powerful model that does each job.
Autonomy, evidence and control
The second five decide whether you keep control as the system takes on more of the work.
Ask what the system may do on its own, how that permission expands, and which gates remain. A strong answer describes declared permissions, confidence thresholds, cost ceilings and human approval points per workflow.
Ask whether telemetry, model weights and verdicts stay inside your environment, and whether on-premises or air-gapped deployment is supported. Private cloud often means the vendor's cloud.
When the system improves from your analysts' decisions, ask where that improvement lives. A strong answer keeps it in your environment, under your control and exportable.
Ask for precision, recall, false positives, MITRE ATT&CK coverage and cost, measured with a method you can repeat. An open benchmark is stronger than a vendor report.
Token spend, per-alert pricing and ingestion fees compound. Ask for a cost model at your volume, including a quiet week and a loud one.
Score each answer from 0 to 2. The pattern matters more than any single row.
A product scoring under 12 of 20 is a triage assistant. That can still be worth buying, but price it and position it as one.
Run a 30-day evaluation on your own telemetry.
Thirty days is enough to answer the ten questions with evidence rather than slides. Insist on your data, not a reference set.
Choose a representative telemetry sample, including a period with a known incident if you have one. Record what your current stack alerted on during that period.
Run the candidate on the sample with no tuning. Count what it found that your stack missed, and what it raised that you would dismiss. Those are recall and false positives on your data.
Hand the findings to analysts inside the product. Measure time to a decision, the quality of the evidence presented, and whether analysts can read and change the workflow that produced it.
Score the ten questions. Compare the cost model at production volume. Ask for the evaluation method in writing so you can repeat it next year, with this vendor or another.
Hold us to the same ten questions.
Vigil is the leading open source AI SOC, Apache 2.0, and runs the detections you already trust. LogLM is a foundation model pretrained on security telemetry: 99% zero-shot, 1% or fewer false positives. SOCBench is the open benchmark we publish. Start with Vigil, or run a free detection assessment on your own telemetry.
