All Articles

Agentic AI SOC Tool Benchmarking and Evaluation Guide

A guide on how to evaluate Agentic AI SOC tools that includes benchmarking advice and an ROI benefit formula.

URL copied

TL;DR: You can evaluate agentic AI cybersecurity and SOC tools by measuring time saved, using an Alert Volume × MTTR formula. Before running that evaluation, we recommend mapping your existing process, narrowing down to the best initial test cases, gathering context, and involving your analysts.

The push factor driving teams towards agentic AI is that there is an overwhelming volume of alerts and investigations piling up on teams that are short on resources. Many SOCs are forced to suppress detection rules or delay investigations just to keep pace. The pull is that AI solutions promise to reduce manual work, scale expertise, and speed up decision-making.

Recent studies show that one-third of IT and business leaders anticipate workload reductions greater than 50% from automated remediation.

Agentic solutions like Legion Security can reduce MTTI/R by 81% in common use cases. A benefit that Iain Paterson, CISO at WELL Health Technologies, described as “an actual supercharger for SOC analysts” and a “must-have to help your Operations teams get ahead of the volume of alerts."

But not all AI-driven solutions deliver the same value as Legion Security does. That’s why we encourage you to test ours and other agentic AI SOC solutions, and why we built the guide below as a practical vendor-neutral approach to benchmarking agentic SOC tools.

What are Agentic AI SOC Tools?

Agentic AI SOC tools like Legion Security are systems that reason and act with autonomy across the investigation process in an SOC.

Rather than following fixed rules or scripts, they are designed to take a goal, such as understanding an alert or verifying a threat, and figure out the steps to achieve it.

That includes retrieving evidence, correlating data, assessing risk, and initiating a response.

These systems are built for flexibility.

They interpret data, ask questions, and adjust based on what they find. In the SOC, that means helping analysts triage alerts, investigate incidents, and reduce manual effort. But because they adapt to their environment, evaluating them requires more than a checklist.

5 Steps to Evaluate the Detection Speed and Performance of an Agentic AI SOC Tool in 2026

Below is a list of steps, with sub-questions you can ask, to benchmark an agentic AI SOC tool in 2026. We've framed these primarily as questions to ask.

1. Map your current SOC processes

It sounds obvious, but before diving into use cases, you need a clear understanding of your current environment. What tools do you rely on? What types of alerts are flooding your queue? Where are your analysts spending most of their time? And just as importantly, where are they truly needed?

Ask:

  • What types of alerts do you want to automate?
  • How long does it currently take to acknowledge and investigate those alerts?
  • Where are your analysts delivering critical value through judgment and expertise?
  • Where is their time being drained by manual or repetitive tasks?
  • Which tools and systems hold key context or history that investigations rely on?

Investigating a user-reported phishing email that follows a predictable structure is a strong candidate for automation.

On the other hand, a suspicious identity-based alert involving cross-cloud access, irregular privileges, and unfamiliar assets may be better suited for manual investigation. These cases require analysts to think creatively, assess multiple possibilities, and make decisions based on a broader organizational context.

Benchmarking is only meaningful when it reflects your reality. Generic tests or template use cases won’t surface the same challenges your team faces daily. Evaluations must mirror your data, your processes, and your decision logic.

Otherwise, you’ll face a painful gap between what the system shows in a demo and what it delivers in production. Your SOC is not a demo environment, and your organization isn’t interchangeable with anyone else’s. You need a system that can operate effectively in your real world, not just in theory.

2. Filter for best-fit AI-driven SOC tool use cases

Once you understand where you need automation and where you don’t, the next step is selecting the right use cases to evaluate.

Focus on alert types that occur frequently and drain analyst time. Avoid artificial scenarios that make the system look good without testing it meaningfully.

Shape the evaluation around:

  • The alerts you want to offload.
  • The tools already integrated into your environment.
  • The logic your analysts use to escalate or resolve investigations.

If the system can’t navigate your real workflows or access the data that matters, it won’t deliver value even if it performs well in a controlled setting.

3. Map and collate sources of context

Accurate investigations depend on more than just alerts. Critical context often lives in ticketing systems, identity providers, asset inventories, previous incident records, or email gateways.

Your evaluation should examine:

  • Which systems store the data your analysts need during an investigation.
  • Whether the agentic system integrates directly with those systems.
  • How well it surfaces and applies relevant context at decision points.

It’s not enough for a system to be technically integrated. It needs to pull the right context at the right time. Otherwise, workflows may complete, but analysts will still need to jump in to validate or fill gaps manually.

4. Bring analysts into the testing loop

Agentic AI SOC systems work alongside humans in surfacing reasoning, offering speed, and allowing feedback that improves performance over time.

Your evaluation should test:

  • Whether the system explains what it’s doing and why.
  • If analysts can give feedback or course-correct.
  • How easily logic and outcomes can be reviewed or tuned.

When it comes to accuracy, two areas matter most:

  • False negatives: when real threats are missed or misclassified
  • False positives: when harmless activity is escalated unnecessarily

False negatives are a direct risk to the organization. False positives create long-term fatigue.

Critically, you should also evaluate how the system evolves over time. Is it learning from analyst feedback? Is it getting better with repeated exposure to similar cases?

A system that doesn’t improve will struggle to generalize and scale across different use cases. Without measurable learning and adaptation, you can’t count on consistent value beyond the initial deployment.

5. Evaluate speed and time saved with this simple formula

Time savings is often used to justify automation, but it only matters when tied to actual analyst workload. Don’t just look at how fast a case is resolved. Consider how often that case type occurs and how much effort it typically requires.

To evaluate this, measure:

  • How long it takes today to investigate each alert type.
  • How frequently those alerts happen.
  • Whether the system fully resolves them or only assists.

Use a simple formula to estimate potential impact:

  • Time Saved = Alert Volume × MTTR
    (where MTTR = MTTA + MTTI)

This provides a grounded view of where automation will drive real efficiency.

MTTA (mean time to acknowledge) and MTTI (mean time to investigate) help capture the full response timeline and show how much manual work can be offloaded.

Some alerts are rare but time-consuming. Others are frequent and simple. Prioritize high-volume, moderately complex workflows. These are often the best candidates for automation with meaningful long-term value. Avoid chasing flashy edge cases that won’t significantly impact operational burden.

Prioritize Reliability

It doesn’t matter how powerful a system is if it fails regularly or requires constant oversight. Reliability is the foundation of trust, and trust is what drives adoption.

Track:

  • How often do workflows complete without breaking.
  • Whether results are consistent across similar inputs.
  • How often manual recovery is needed.

If analysts don’t trust the output, they won’t use it. And if they constantly have to step in, the system becomes another point of friction, not relief.

Realistic Agentic AI SOC Tool Benefits

Agentic AI can reshape SOC operations.

Neil Robison, Head of Security Engineering & Cybersecurity at Virgin Money, described the impact of deploying Legion Security as “evolving from handcrafted systems to precision manufacturing: aligned to our flow, but now faster, repeatable, and secure.”

But realizing these kinds of benefits depends on how well the system performs in your real-world conditions. The strongest agentic AI solutions adapt to your environment, support your team, and deliver consistent value over time.

When evaluating agentic AI, focus on:

  • Your actual alert types, workflows, and operational goals.
  • The tools and systems that store the context your team depends on.
  • Analyst involvement, feedback loops, and decision transparency.
  • Real-time savings tied to the volume and complexity of your alerts.
  • Reliability and trust in day-to-day performance.

The best system is the one that fits the reality of your SOC.

Legion Security is an Agentic AI SOC tool. Legion works with your analysts to learn your SOC’s workflows, develop new ones, and conduct transparent and trustworthy automations in your environment. Learn more.

URL copied

Security investigations rarely start with all the context needed to reach the right decision, and we see plenty of examples of this in real environments. Let’s look at an anonymized but recent example. Every quarter, a publicly traded enterprise’s finance team uploads the company's still-unreleased earnings package which consists of revenue, forecasts, and results that won't go public until earnings day to a restricted SharePoint site for executive review. The package contains sensitive financial information so the upload triggers a DLP alert for review. Pretty standard stuff.

That alert triggered an analyst investigation where the incident response team confirmed the uploader was indeed a part of the reporting team, the destination was the approved site, and access was limited to only the small group of executives who were supposed to see it. Nothing dangerous, so it was safely closed as benign. This single investigation established the conditions that made the activity safe: who was expected to upload the file, where it was supposed to go, and who was supposed to have access.

But the lingering question is… what should be carried forward and/or codified from that investigation? This question is one that we’re obsessed with answering and helping our customers address.

With Legion, instead of carrying forward a single verdict from a single investigation, enterprises can uniquely capture the conditions that each investigation establishes together with the underlying and complementing evidence behind them. On a continuous basis. This holistic view matters, particularly in today’s world, because the same activity type doesn't always mean the same thing, and this is a constantly moving target as environments change. Using our ‘finance team uploading earnings files into SharePoint’ example, one of the conditions that was met, who had access to the folder, can change very quickly. So perhaps the next time, the package is the same, the site is the same, the timing is the same, but the folder may have been shared with an external account or a new unverified user.

The challenge isn't collecting more data. Most enterprises already have plenty of it, scattered across identity providers, endpoints, SaaS apps, and past investigations. The challenge is turning that raw data into knowledge that's reliable enough, and accessible enough, for agents to actually reason over: preserving what made something true, connecting it to the organizational context around it, and continuously testing whether it still holds as the organization changes.

Knowledge needs conditions, not conclusions

That's why Legion represents organizational knowledge and context as a continuously evolving model that connects identities, teams, systems, data, access, behaviors, and the evidence establishing how they all relate to one another.

Legion’s knowledge isn't built from investigations alone. Legion brings information from across the environment, including identities, access, systems, infrastructure, and the relationships between them, into the same layer. Past investigations add another important source, giving Legion an accumulated history from day one: what analysts already checked, what they found, and the evidence that supported those decisions.

Raw data on its own doesn't tell an agent much. An identity, a login, a file upload, a network connection, in isolation, are just data points. What makes this usable is the relationship it has to everything around it. That's what turns data into knowledge an agent can actually act on: not just what happened, but who was involved, what it touched, what normally follows it, and what it means if it doesn't.

That gap between "looks the same" and "is the same" is hard to manage at enterprise scale, and Legion Knowledge is designed to connect the data flowing in and out of thousands of employees, dozens of teams, hundreds of new and existing tools changing in real time, and access to relationships that change over time. This empowers security teams, and their agents, to stay on top of every legitimate exception, relationship, and operating pattern at agentic scale.

We see all the time that not everything security tools observe should become codified as organizational best practices. Before new information can influence future investigations, there needs to be enough evidence to support it. Otherwise, an observation can become an assumption that extends beyond what the evidence actually established, and an assumption an agent can't verify is a liability, not an insight.

Research on memory management in LLM agents shows why this matters. Researchers at Harvard, Michigan State, and other institutions found that agents exhibit what they call "experience-following": the more similar a new task is to an experience retrieved from memory, the more likely the agent is to follow that past execution. That's useful when the retrieved experience applies. When it doesn't, the agent can carry an assumption from one task into the next that the new evidence doesn't support. Reliable knowledge is what keeps that experience-following useful instead of risky.

Useful organizational knowledge is more than a collection of isolated facts. The relationships and intricacies between those facts provide the context needed to interpret them: not just what is known about an identity, system, or activity, but how each relates to the organization around it. Preserving those relationships is also what surfaces the insights security teams actually need: correlation across seemingly unrelated events, the blast radius of a compromised identity or system, and where the real detection opportunities sit. None of that comes from more data. It comes from data that's been made reliable enough to connect.

Strong evidence can still become outdated

Preserving the right conditions solves one problem, but it creates another: conditions change.

In our finance example, previous investigations may provide strong evidence that only a specific group of executives had access to the folder. That evidence doesn't become wrong when someone new is granted access; they could be, simply, a new member of the exec team.

That's why Legion separates confidence from freshness: confidence reflects how strongly the evidence supports what is known, while freshness reflects how recently those conditions have been verified.

That distinction matters when existing knowledge is used in a new investigation, or acted on by an agent. Something can remain strongly supported by evidence while becoming too stale to rely on without verifying that the same conditions still hold. An agent that can't tell the difference between confident-and-fresh and confident-and-stale is an agent that will eventually act on the wrong assumption.

New evidence has to reconcile with existing knowledge

Every new investigation produces information that could become organizational knowledge. But observing something doesn't automatically make it a best practice. Before new evidence changes the output, Legion evaluates it against what the organization already knows. It may reinforce something already established, add something new, or contradict it.

New evidence doesn't necessarily make the old evidence wrong. Both may be valid: one describes what was true when it was established, while the other shows that something has since changed. Preserving the evidence and timing behind both lets security teams understand that change rather than simply replacing one version with another.

This makes evaluation part of the learning process, not just a gate at the moment knowledge is created. An investigation produces new evidence, that evidence is evaluated against existing knowledge, and only then can it change what Legion, and the agents built on top of it, carry into future investigations.

Learning is automatic. Authority isn't.

Automatic learning shouldn't make organizational knowledge opaque to the humans who rely on it. If that knowledge is going to shape future investigations, and the agents acting on them, the people who know the organization should be able to see what was learned and contribute to its quality.

Human feedback adds another signal to that process. A validation can strengthen what Legion has learned, while a correction or rejection can challenge it. And for people to make those judgments, the knowledge has to remain traceable: where it came from, the evidence behind it, and how it has changed over time.

Knowledge has to remain trustworthy

Organizational knowledge is useful only as long as there is a reason to keep trusting it. Something can be well supported and still become outdated. New evidence can strengthen what is already known or show that the environment has changed. And a conclusion that was right six months ago shouldn't become an assumption simply because nothing has challenged it yet.

That's the distinction we built Legion around. The goal isn't simply to collect more data about an organization. It's to make that data reliable and accessible enough, for analysts and agents alike, to know what still deserves to be trusted.

Cybersecurity
How Knowledge Management Impacts Security Investigations
September 9, 2026
min read

How Legion turns security data into reliable, agent-ready knowledge to preserve evidence, track freshness, and surface what still deserves trust.

Gili Diamant

Hear directly from Neil Robinson, CISO at Virgin Money, on how Legion’s agentic security operations platform learned his team’s existing investigation workflows and scaled them without forcing new tools or processes onto analysts. Neil shares how Legion delivered faster, more repeatable, and more secure investigations, aligned to how his team already operates, rather than a generic playbook applied from outside.

0:00 / 0:00
Cybersecurity
Customer Testimonial: Virgin Money
August 2, 2026
min read

Neil Robinson, CISO at Virgin Money

Legion Team

TL:DR. The eval said Claude had no internet access. That ended up being not true… three times. Anthropic just disclosed that across 141,006 cybersecurity evaluation runs, three separate incidents let its models reach the open internet from environments that were supposed to be sealed. A misconfiguration between Anthropic and its eval partner, Irregular, left the machines with live connectivity. The models didn't know it was a mistake, they just found a door that should not have been open, and walked through it. It’s a stark reminder of the importance of agentic systems' reliance on organizational context, and how without it, or without a complete picture of it, things can go astray.

Three Models, Three Choices

First, Opus 4.7 ran into a real company that happened to share a name with a fictional target in its scenario. It exploited weak passwords and unauthenticated endpoints, pulled credentials, and reached a production database with several hundred real rows in it. Anthropic's own review notes that the model "eventually recognized... that the system was real" but kept going anyway.

Mythos 5 published a booby-trapped Python package to PyPI as part of a capture-the-flag exercise. It sat live for about an hour and was downloaded onto 15 real systems, including a security vendor's scanner. The model's own reasoning flagged the problem as "NOT okay, and surely not the intended solution,” then talked itself out of stopping, pointing to an unfamiliar certificate authority and a 2026 date as evidence the whole thing must be simulated.

An internal research model scanned roughly 9,000 targets, compromised one real company using textbook techniques (exposed debug credentials, SQL injection), recognized the target was live, and stopped on its own.

It’s the same false premise with the same moment of doubt, but three different endings for how agents proceeded.

The Line Anthropic Drew

Anthropic's own framing of the incident is quite poignant, as the state, "The line between an aligned action and a harmful one is dependent on the model's understanding of its situation."

This really hits home, as it’s not what you might expect, which could be a statement about model safety training or the importance of guardrails, but rather a clear statement that the agent’s contextual awareness of what it was doing determines the outcomes.

Every one of these models had the values and the guardrails baked in during training. What they didn't reliably have was an accurate read on where they actually were. Read the situation wrong, and the same model that would refuse an obviously harmful request will talk itself into finishing one it thinks is fictional.

Why This Isn't Just Anthropic's Problem

While it is likely tempting to file this under "frontier lab eval hygiene" and move on, we don’t think this is the way to think about this particular incident.

Anthropic calls it correctly and deserves much credit for their response and summary to this grouping of incidents. This was a harness and operational failure, not an alignment failure. The failure here wasn't a jailbreak, a novel exploit, or a model deciding to go rogue. In fact, the models behaved exactly as agentic systems behave when they're missing organizational context… they filled the gap with their best guess, it just so happened that two out of three guessed wrong.

On the defensive side, this is a tidy summary of why there is hesitation to unleash generic AI systems into their environments. Particularly for an AI agent that is responsible for triaging your alerts, scoping a compromise, or deciding whether to isolate a host, it is critical to remember that these agents inherently make the same kind of situational judgment call, constantly and with real stakes. The agent determines if this is real, is this expected, does this action match how this specific business actually operates. The Anthropic incidents are a rare, public, unusually well-documented look at what happens when that judgment runs without enough grounding to get it right. That should be a stark reminder of how every CISO evaluates the agentic tools already running inside their own stack, from offensive research models to defensive SOC copilots alike.

What This Should Change for Security Leaders

From our perspective, there are a few things worth pulling out of this disclosure and applying directly to whatever agentic AI you're already running or evaluating:

  • Assume your environment is a target, not just a beneficiary. Fifteen real systems downloaded a package that was never meant to exist. Roughly 9,000 targets got scanned by a model that was supposed to be sandboxed. Eval infrastructure, research environments, and "internal only" tooling deserve the same monitoring as production; because from the outside, they increasingly look identical.
  • Don't take "it has guardrails" on faith. Context is king. All three models retained their safety training. It didn't prevent two of the three incidents. Guardrails matter, but they're not a substitute for auditability and contextual awareness — you need to see the reasoning and deploy agents that understand your organizational context (tools, processes, bespoke knowledge, etc.), not just trust the outcome.
  • Demand whitebox AI, not a black box you hope behaves. Anthropic found this because it went back and read the transcripts. That's the standard: agentic systems, yours or a vendor's, should be inspectable, not just monitored for red flags.
  • Build for the model that stops, not the one that rationalizes. The internal research model got it right because it had enough signal to recognize reality and enough restraint built in to act on that recognition. That combination: context plus a real decision point for a human or a hard stop, is a choice, not coincidence.

Anthropic deserves real credit here: they found this themselves, through proactive review, disclosed it before anyone made them, and are publishing the transcripts for all to see and learn from. That's the posture every lab and every vendor building agentic security tools should be held to, very much including ourselves as well.

But the underlying lesson is the one we keep coming back to: agentic AI is only as trustworthy as its contextual understanding of the situation it's actually in. That's true for a frontier model deciding whether a target is real. It's just as true for an AI agent in your SOC deciding whether an alert is a false positive, a test, or the start of an incident. Build the context in, keep the reasoning visible, and give the system a real reason to stop when it isn't sure, because agents are often irrationally confident and take ‘not sure’ as an instruction to pick their best guess and go.

AI
Context, Not Guardrails: The Line Between Aligned and Harmful
July 31, 2026
min read

Anthropic found its "sandboxed" models reaching the real internet three times. Here's why context, not guardrails, decides if agentic AI stays safe.

Legion Team