Resources


DragonClaw is Legion's agent orchestration layer for the SOC; grounded in your org's own context, not a generic chatbot bolted onto security tools.
TL:DR: Ask any security team what would give them back the most time, and the answers tend to converge on the same theme: less time spent stitching things together, more time spent actually deciding. These are exactly the things that DragonClaw is built to optimize, as the orchestration layer that deploys Legion’s trusted AI agents into any security task.
Automated workflows have already gotten teams part of the way there, triggering playbooks and kicking off investigations the moment an alert fires. DragonClaw is built upon the foundation of Legion’s platform, in that we require zero integrations in exchange for the ability to operate any tool, and goes further: it leverages the business context (past cases, runbooks, recordings, etc.) to orchestrate the agents needed to respond to an alert or escalation, to tell you why the last three cases like this one got closed the way they did, and to surface the exact query that finds the right evidence in your specific environment. That's the difference between automation that runs a process and intelligence that understands one.
Instead of an analyst hunting across five tools to reconstruct context that already exists somewhere in the organization's own history, DragonClaw brings that context directly to them and performs a task, in their own way, the moment they need it. Ask a question, get a grounded answer or a completed action, drawn from how your organization actually operates, not a generic playbook applied from outside.
The result is analysts can spend more of their time on the judgment calls only a person can make while orchestrating the agentic layer, where DragonClaw handles the reconstruction, the pattern-matching, and the acceleration and scale that used to eat the hours in between.
From Analyst to CISO: Closing the Context Gap in Security Operations
For security analysts, think of real-world threat hunting. Today, it means pulling and reading vast amounts of data across a bunch of different tools before you can even form an opinion or a lead on where to go. DragonClaw runs that process, end-to-end, with agents. DragonClaw consumes data across all of your tools, correlates it, and comes back with a thesis for the analyst to either approve or disapprove.
If you're a CISO or security leader, quickly investigating what the risk or impact is for a CVE requires organizational context not contained in a single tool. DragonClaw assembles all of that data and surfaces the answers, with recommendations, and where appropriate, autonomous actions that can put the findings to work.
Add it up across a team, and the opportunity is real: practitioners who spend their time on judgment instead of relearning tools, leaders with a straight answer whenever they need one, and a security program built to scale with the threat landscape instead of falling further behind it.
Introducing DragonClaw
DragonClaw is Legion Security's agent orchestration layer for the SOC. It gives security teams the ability to invoke Legion's agents in plain conversational language, enabling security teams to seamlessly get work done, or to answer questions about how their processes, tools, and people are actually making decisions.
One thing to be clear is that this is not (yet another) bolt-on chat interface. DragonClaw is the next step in the Legion platform, built on everything Legion has already learned across the tools, knowledge, and decision logic for your team’s security workflows. DragonClaw takes that further, putting that context and institutional knowledge to work answering questions and completing tasks the moment someone asks.
Under the hood, DragonClaw interprets intent, figures out which agents a request actually requires, and orchestrates them; across all tools in the stack, including agents that take real action, like API calls or web interactions, without any integrations required. All of it runs inside configurable guardrails: explicit permission before any response action, only approved tools, and credentials pulled from secure vaults. Nothing about “conversational” means “unsupervised.”
What Changes For Each of You
Threats are scaling with AI. Automation and agents close a large part of that gap, and they'll take a SOC further than headcount ever could… but not all the way. Security teams need humans to stay in the loop, not to keep pace with volume (which they can’t), but to supervise the work, evaluate outcomes, test and challenge what the agents conclude, and make sure security stays something that enables the business rather than something that slows it down or breaks it. Security analysts and leaders serve essentially as the maestros of the agentic orchestra. That's the same place the sharpest thinking on AI lands more broadly: the machine executes and reasons whereas the human owns judgment where needed and accountability.
DragonClaw is what supercharges the security workers. It's what lets a security team orchestrate its agents instead of losing control over what they do. For security practitioners and SOC analysts, that shows up as a partner inside the investigation itself: context and enrichment on demand, memory across past cases, guidance on what to do next, and the ability to generate the right query for your environment instead of learning a new query language from scratch.
For managers and security leadership, it's one place to ask about real-time SLA risk, process improvement opportunities, MTTR and false-positive trends, bottlenecks, coverage gaps, and team workload — instead of stitching the answer together from five dashboards.
For CISOs, DragonClaw provides direct answers on risk posture, SLA exposure, MTTR trends, exposure to a new CVE, audit evidence, automation ROI, and board-ready reporting, available the moment you need them instead of on the next reporting cycle.
Not Another Chatbot, An Orchestrator
Chat interfaces are becoming table stakes across the industry, and we're not going to pretend otherwise; it’s been proven that chat alone isn't a durable differentiator. What makes DragonClaw different is what's underneath it: every answer and every action is grounded in the workflows, case history, and coverage data Legion has already built for your specific security team and your specific organization.
A generic assistant sitting outside your platform can talk about security in general. DragonClaw can talk about your security workflows, because it already has the record of how your security team works.
That's the same principle behind everything Legion builds: AI for defenders should understand how a specific business operates, across its tools, its workflows, its people, before it's trusted to answer questions or take action with real business impact. DragonClaw is where that understanding becomes something every person in your organization can talk to directly, whether that's the analyst mid-investigation, the manager reviewing the week, or the CISO prepping for the board.
DragonClaw will be showcased at Black Hat USA 2026, visit us at Booth #5150 to see it in action!
Hear directly from Neil Robinson, CISO at Virgin Money, on how Legion’s agentic security operations platform learned his team’s existing investigation workflows and scaled them without forcing new tools or processes onto analysts. Neil shares how Legion delivered faster, more repeatable, and more secure investigations, aligned to how his team already operates, rather than a generic playbook applied from outside.

TL:DR. The eval said Claude had no internet access. That ended up being not true… three times. Anthropic just disclosed that across 141,006 cybersecurity evaluation runs, three separate incidents let its models reach the open internet from environments that were supposed to be sealed. A misconfiguration between Anthropic and its eval partner, Irregular, left the machines with live connectivity. The models didn't know it was a mistake, they just found a door that should not have been open, and walked through it. It’s a stark reminder of the importance of agentic systems' reliance on organizational context, and how without it, or without a complete picture of it, things can go astray.
Three Models, Three Choices
First, Opus 4.7 ran into a real company that happened to share a name with a fictional target in its scenario. It exploited weak passwords and unauthenticated endpoints, pulled credentials, and reached a production database with several hundred real rows in it. Anthropic's own review notes that the model "eventually recognized... that the system was real" but kept going anyway.
Mythos 5 published a booby-trapped Python package to PyPI as part of a capture-the-flag exercise. It sat live for about an hour and was downloaded onto 15 real systems, including a security vendor's scanner. The model's own reasoning flagged the problem as "NOT okay, and surely not the intended solution,” then talked itself out of stopping, pointing to an unfamiliar certificate authority and a 2026 date as evidence the whole thing must be simulated.
An internal research model scanned roughly 9,000 targets, compromised one real company using textbook techniques (exposed debug credentials, SQL injection), recognized the target was live, and stopped on its own.
It’s the same false premise with the same moment of doubt, but three different endings for how agents proceeded.
The Line Anthropic Drew
Anthropic's own framing of the incident is quite poignant, as the state, "The line between an aligned action and a harmful one is dependent on the model's understanding of its situation."
This really hits home, as it’s not what you might expect, which could be a statement about model safety training or the importance of guardrails, but rather a clear statement that the agent’s contextual awareness of what it was doing determines the outcomes.
Every one of these models had the values and the guardrails baked in during training. What they didn't reliably have was an accurate read on where they actually were. Read the situation wrong, and the same model that would refuse an obviously harmful request will talk itself into finishing one it thinks is fictional.
Why This Isn't Just Anthropic's Problem
While it is likely tempting to file this under "frontier lab eval hygiene" and move on, we don’t think this is the way to think about this particular incident.
Anthropic calls it correctly and deserves much credit for their response and summary to this grouping of incidents. This was a harness and operational failure, not an alignment failure. The failure here wasn't a jailbreak, a novel exploit, or a model deciding to go rogue. In fact, the models behaved exactly as agentic systems behave when they're missing organizational context… they filled the gap with their best guess, it just so happened that two out of three guessed wrong.
On the defensive side, this is a tidy summary of why there is hesitation to unleash generic AI systems into their environments. Particularly for an AI agent that is responsible for triaging your alerts, scoping a compromise, or deciding whether to isolate a host, it is critical to remember that these agents inherently make the same kind of situational judgment call, constantly and with real stakes. The agent determines if this is real, is this expected, does this action match how this specific business actually operates. The Anthropic incidents are a rare, public, unusually well-documented look at what happens when that judgment runs without enough grounding to get it right. That should be a stark reminder of how every CISO evaluates the agentic tools already running inside their own stack, from offensive research models to defensive SOC copilots alike.
What This Should Change for Security Leaders
From our perspective, there are a few things worth pulling out of this disclosure and applying directly to whatever agentic AI you're already running or evaluating:
- Assume your environment is a target, not just a beneficiary. Fifteen real systems downloaded a package that was never meant to exist. Roughly 9,000 targets got scanned by a model that was supposed to be sandboxed. Eval infrastructure, research environments, and "internal only" tooling deserve the same monitoring as production; because from the outside, they increasingly look identical.
- Don't take "it has guardrails" on faith. Context is king. All three models retained their safety training. It didn't prevent two of the three incidents. Guardrails matter, but they're not a substitute for auditability and contextual awareness — you need to see the reasoning and deploy agents that understand your organizational context (tools, processes, bespoke knowledge, etc.), not just trust the outcome.
- Demand whitebox AI, not a black box you hope behaves. Anthropic found this because it went back and read the transcripts. That's the standard: agentic systems, yours or a vendor's, should be inspectable, not just monitored for red flags.
- Build for the model that stops, not the one that rationalizes. The internal research model got it right because it had enough signal to recognize reality and enough restraint built in to act on that recognition. That combination: context plus a real decision point for a human or a hard stop, is a choice, not coincidence.
Anthropic deserves real credit here: they found this themselves, through proactive review, disclosed it before anyone made them, and are publishing the transcripts for all to see and learn from. That's the posture every lab and every vendor building agentic security tools should be held to, very much including ourselves as well.
But the underlying lesson is the one we keep coming back to: agentic AI is only as trustworthy as its contextual understanding of the situation it's actually in. That's true for a frontier model deciding whether a target is real. It's just as true for an AI agent in your SOC deciding whether an alert is a false positive, a test, or the start of an incident. Build the context in, keep the reasoning visible, and give the system a real reason to stop when it isn't sure, because agents are often irrationally confident and take ‘not sure’ as an instruction to pick their best guess and go.

Anthropic found its "sandboxed" models reaching the real internet three times. Here's why context, not guardrails, decides if agentic AI stays safe.
TL:DR: Ask any security team what would give them back the most time, and the answers tend to converge on the same theme: less time spent stitching things together, more time spent actually deciding. These are exactly the things that DragonClaw is built to optimize, as the orchestration layer that deploys Legion’s trusted AI agents into any security task.
Automated workflows have already gotten teams part of the way there, triggering playbooks and kicking off investigations the moment an alert fires. DragonClaw is built upon the foundation of Legion’s platform, in that we require zero integrations in exchange for the ability to operate any tool, and goes further: it leverages the business context (past cases, runbooks, recordings, etc.) to orchestrate the agents needed to respond to an alert or escalation, to tell you why the last three cases like this one got closed the way they did, and to surface the exact query that finds the right evidence in your specific environment. That's the difference between automation that runs a process and intelligence that understands one.
Instead of an analyst hunting across five tools to reconstruct context that already exists somewhere in the organization's own history, DragonClaw brings that context directly to them and performs a task, in their own way, the moment they need it. Ask a question, get a grounded answer or a completed action, drawn from how your organization actually operates, not a generic playbook applied from outside.
The result is analysts can spend more of their time on the judgment calls only a person can make while orchestrating the agentic layer, where DragonClaw handles the reconstruction, the pattern-matching, and the acceleration and scale that used to eat the hours in between.
From Analyst to CISO: Closing the Context Gap in Security Operations
For security analysts, think of real-world threat hunting. Today, it means pulling and reading vast amounts of data across a bunch of different tools before you can even form an opinion or a lead on where to go. DragonClaw runs that process, end-to-end, with agents. DragonClaw consumes data across all of your tools, correlates it, and comes back with a thesis for the analyst to either approve or disapprove.
If you're a CISO or security leader, quickly investigating what the risk or impact is for a CVE requires organizational context not contained in a single tool. DragonClaw assembles all of that data and surfaces the answers, with recommendations, and where appropriate, autonomous actions that can put the findings to work.
Add it up across a team, and the opportunity is real: practitioners who spend their time on judgment instead of relearning tools, leaders with a straight answer whenever they need one, and a security program built to scale with the threat landscape instead of falling further behind it.
Introducing DragonClaw
DragonClaw is Legion Security's agent orchestration layer for the SOC. It gives security teams the ability to invoke Legion's agents in plain conversational language, enabling security teams to seamlessly get work done, or to answer questions about how their processes, tools, and people are actually making decisions.
One thing to be clear is that this is not (yet another) bolt-on chat interface. DragonClaw is the next step in the Legion platform, built on everything Legion has already learned across the tools, knowledge, and decision logic for your team’s security workflows. DragonClaw takes that further, putting that context and institutional knowledge to work answering questions and completing tasks the moment someone asks.
Under the hood, DragonClaw interprets intent, figures out which agents a request actually requires, and orchestrates them; across all tools in the stack, including agents that take real action, like API calls or web interactions, without any integrations required. All of it runs inside configurable guardrails: explicit permission before any response action, only approved tools, and credentials pulled from secure vaults. Nothing about “conversational” means “unsupervised.”
What Changes For Each of You
Threats are scaling with AI. Automation and agents close a large part of that gap, and they'll take a SOC further than headcount ever could… but not all the way. Security teams need humans to stay in the loop, not to keep pace with volume (which they can’t), but to supervise the work, evaluate outcomes, test and challenge what the agents conclude, and make sure security stays something that enables the business rather than something that slows it down or breaks it. Security analysts and leaders serve essentially as the maestros of the agentic orchestra. That's the same place the sharpest thinking on AI lands more broadly: the machine executes and reasons whereas the human owns judgment where needed and accountability.
DragonClaw is what supercharges the security workers. It's what lets a security team orchestrate its agents instead of losing control over what they do. For security practitioners and SOC analysts, that shows up as a partner inside the investigation itself: context and enrichment on demand, memory across past cases, guidance on what to do next, and the ability to generate the right query for your environment instead of learning a new query language from scratch.
For managers and security leadership, it's one place to ask about real-time SLA risk, process improvement opportunities, MTTR and false-positive trends, bottlenecks, coverage gaps, and team workload — instead of stitching the answer together from five dashboards.
For CISOs, DragonClaw provides direct answers on risk posture, SLA exposure, MTTR trends, exposure to a new CVE, audit evidence, automation ROI, and board-ready reporting, available the moment you need them instead of on the next reporting cycle.
Not Another Chatbot, An Orchestrator
Chat interfaces are becoming table stakes across the industry, and we're not going to pretend otherwise; it’s been proven that chat alone isn't a durable differentiator. What makes DragonClaw different is what's underneath it: every answer and every action is grounded in the workflows, case history, and coverage data Legion has already built for your specific security team and your specific organization.
A generic assistant sitting outside your platform can talk about security in general. DragonClaw can talk about your security workflows, because it already has the record of how your security team works.
That's the same principle behind everything Legion builds: AI for defenders should understand how a specific business operates, across its tools, its workflows, its people, before it's trusted to answer questions or take action with real business impact. DragonClaw is where that understanding becomes something every person in your organization can talk to directly, whether that's the analyst mid-investigation, the manager reviewing the week, or the CISO prepping for the board.
DragonClaw will be showcased at Black Hat USA 2026, visit us at Booth #5150 to see it in action!

DragonClaw is Legion's agent orchestration layer for the SOC; grounded in your org's own context, not a generic chatbot bolted onto security tools.
Security teams don't get to choose their threat model. They inherit it: alert volumes outpacing headcount, attackers moving at machine speed, and a stack of tools that all assume a human is sitting in the middle of every decision. Legion Security was built to change that operating model. It should go without saying that the platform doing the changing has to be secure to its core.
CISA's Secure by Design Pledge exists to push the whole industry toward that standard: security as a default, not an upsell. Legion Security has signed the pledge, and this page lays out how we deliver on each of its seven goals today, where we go beyond them, and where we're still pushing.
One note before diving in. Legion Security is independently certified against SOC 2 Type 2, ISO/IEC 27001:2022, and HIPAA, and against ISO/IEC 42001:2023 for responsible AI management, because AI reasoning sits at the center of everything the platform does and we believe that deserves its own standard of scrutiny. Everything below is backed by audited evidence, not just a blog post. Every report, policy, and pentest result is available on request through our Trust Center.
Authentication
The pledge commitment: Demonstrate actions taken to measurably increase the use of multi-factor authentication across the manufacturer's products
Every Legion Security customer gets enterprise-grade identity protection from day one, not as an add-on. The platform integrates with any SAML 2.0 or OIDC-compliant identity provider, including Okta, Microsoft Entra ID, and Google Workspace, so your existing MFA and conditional access policies extend automatically to every login. No extra licensing tier, no SSO tax, no “contact sales” for basic protection.
For teams still mid-rollout on SSO, direct logins to Legion Security require MFA. There is no opt-out, not for admins, not for anyone.
And because Legion Security operates natively in the browser rather than through a sprawl of API integrations, there are no long-lived API keys quietly wiring it into the rest of your stack. Your analysts sign in once, through the identity layer they already trust, and the platform meets them there. Sessions are scoped and expire like any other corporate login, which is exactly how it should be.
Default Passwords
The pledge commitment: Demonstrate measurable progress towards reducing default passwords across the manufacturer's products.
Default and shared passwords are how breaches happen quietly. Legion Security doesn't have any, and we never will. Every customer environment is provisioned with unique, securely generated credentials from the moment it's stood up, and service-to-service connections run on scoped, short-lived access that expires on its own rather than a static key someone forgot to rotate. There's no fallback password lurking in a setup guide for an attacker to find.
Reducing Entire Classes of Vulnerabilities
The pledge commitment: Demonstrate actions taken towards enabling a significant measurable reduction in the prevalence of one or more vulnerability classes across the manufacturer's products.
Most security automation platforms expand your attack surface before they ever deliver value: months of custom connectors and API plumbing, each one a new door into your environment. Legion Security was built to eliminate that door entirely. The platform works through the browser your analysts already use, so there's no integration sprawl and no expanded blast radius when something goes wrong elsewhere in your stack.
For an AI-native platform, though, the honest conversation starts with a different vulnerability class: prompt injection. Legion Security's agents reason over evidence pulled from your tools, and some of that evidence, phishing emails and attacker-controlled artifacts among it, is adversarial. So we treat all of it as untrusted input. Evidence is validated and normalized before it ever reaches a model, agents operate with tightly scoped permissions rather than broad tenant access, and actions that change your environment pass through explicit approval gates. Malicious content can try to talk to our agents. It doesn't get to instruct them.
When we do find a weakness, whether through the independent penetration tests we commission, a researcher's report, or our own review, we don't stop at the fix. Significant findings get a root cause analysis and a set of preventative changes aimed at the pattern, not just the instance. Security isn't a checkpoint at the end of our development cycle; it's a constraint we design around from the start.
Security Patches
The pledge commitment: Demonstrate actions taken to measurably increase the installation of security patches by customers.
Legion Security runs as SaaS, full stop. There's no patch cycle for your team to manage, no fleet of installations slowly drifting out of date, no tenant running last year's fixes because nobody got around to the upgrade. We ship continuously through automated pipelines, so when we close a gap, every customer is protected at the same moment. The single biggest reason security patches don't get applied, the burden falling on the customer, simply doesn't exist in our model.
Vulnerability Disclosure Policy
The pledge commitment: Publish a vulnerability disclosure policy that authorizes testing in good faith, commits to not pursuing legal action against good-faith researchers, and provides a clear channel to report vulnerabilities.
As a SaaS-only platform, Legion Security maintains a vulnerability management policy that defines severity classifications, SLAs, and response processes for issues in the platform that may impact customers. We commit to those timelines contractually in our customer agreements, and our security operations are audited, internally and by third parties, to confirm we actually hold to them.
Researchers have a clear channel to responsibly disclose vulnerabilities to us, with explicit safe harbor: we will not pursue legal action against anyone making a good-faith effort to find and report an issue. Researchers working with us are partners, not liabilities, and we treat them that way.
Customers can follow all of it through our Trust Center, which provides self-service access to policies, procedures, security notifications, and third-party assessment reports such as penetration tests. Our Shared Responsibility Model, covered under Evidence of Intrusions below, defines which vulnerability management obligations sit with us and which stay with you. And when we patch a vulnerability that matters to you, you'll see it documented as a regular part of our release notes.
CVEs
The pledge commitment: Demonstrate transparency in vulnerability reporting, including accurate CVE records for the manufacturer's products.
Legion Security's SaaS-only architecture means we don't ship versioned software the way traditional on-prem vendors do, which changes how public vulnerability reporting typically plays out. What doesn't change is our commitment to transparency. We're formalizing a procedure that defines exactly how our security team evaluates and reports vulnerabilities, including issuing CVE records where appropriate, so customers always know what to expect from us. Not just when something goes wrong, but how we'll communicate it.
Evidence of Intrusions
The pledge commitment: Demonstrate a measurable increase in the ability for customers to gather evidence of cybersecurity intrusions affecting the manufacturer's products.
You shouldn't have to pay extra to see what's happening in your own tenant. Every Legion Security customer gets detailed, security-relevant audit logs at no additional cost, exportable to your own data lake for as long as you need to keep them. If something happens, you're never locked out of the evidence.
And here's the part we think matters most for a platform like ours: those logs don't stop at human logins and admin changes. Every action an agent takes during an investigation is recorded, along with the evidence it examined and the reasoning behind what it did. When an AI is doing SOC work, “who did what and why” has to include the AI. With Legion Security, it does.
We also publish a Shared Responsibility Model that spells out which parts of logging, monitoring, and incident response we own and which stay with you, so there's no ambiguity on the day it counts.
Software Supply Chain Security
The pledge commitment: The software manufacturer should maintain and share provenance data of third-party dependencies and have processes to govern its use of, and contributions to, open-source software components.
The pledge stops at seven goals. Your security questionnaires don't, and the topic that comes up most is supply chain. Third-party and open-source components in Legion Security are scanned automatically in source control and CI before anything reaches production, new vendors and dependencies go through a risk review before we adopt them, and a software bill of materials is available to customers through the Trust Center. If it runs inside Legion Security, we can tell you what it is and where it came from.
Looking Ahead
Legion Security's whole premise is that the best security systems don't freeze in place; they keep learning your environment and getting sharper with every investigation. We hold our own security program to the same standard. As the platform grows and the threat landscape shifts, this won't be a static page. We'll keep it current, and we'll keep raising our own bar right alongside it.
Questions about anything above? Dig into the documentation in our Trust Center, or reach out to your Legion Security team directly.

See how Legion Security meets CISA's Secure by Design Pledge with MFA by default, no default passwords, full audit logs, and SOC 2, ISO 27001, and HIPAA-certified security.
Demand for agentic security that actually works in complex enterprise environments has never been higher, and today we're excited to take a meaningful step forward in meeting it
We're excited to announce that Legion Security has partnered with Optiv to become an Authorized Partner to help enterprises stop talking about the same-old-problem, and start putting AI to work. Security teams are under pressure that doesn't need a lot of explaining. Analysts, engineers, and practitioners are being asked to do more with less; more alerts, more tools, more threat surface, and fewer people to manage it all. AI was supposed to be the great equalizer, and the promise of the AI SOC was compelling: automate the noise, free up your people, let machines handle the volume.
The reality has been… more complicated.
Most AI security tools were built generically for a generic security team in a generic enterprise. One problem with this is… what is an average security team? Every large organization has processes that are entirely their own: workflows built around a specific stack, custom tools that were built and tuned over long stretches, tribal knowledge accumulated over years, investigation procedures tuned to their environment, their risk tolerance, their regulators, their customers.
Heavy API integrations try to stitch it together but end up slow, brittle, and context-poor (at best). And agents that operate inside a black box create exactly the kind of trust deficit that makes security leaders hesitate to hand anything off at all.
This is the gap Legion was built to close.
A Different Approach to Agentic Security in the Enterprise
The premise of Legion is straightforward: nobody knows your security operations like you do. Our platform doesn't arrive with assumptions about how your team should work. Instead, it observes and learns from how your team actually works; across your tools, your workflows, your most repetitive processes and your most bespoke ones, and then uses that knowledge to build optimized AI agents that operate within the context of your organization.
We don’t require integrations for full contextual awareness. We’re an open book (no black box) that leans on our browser-based approach to see what your analysts see and do, learns what they know, and earns YOUR trust before taking action.
The result is agentic security that can actually scale in the enterprise — not by replacing how teams work, but by amplifying it.
The Imperative for Partnering with Optiv
Becoming an Optiv Authorized Partner matters because of what Optiv represents to the enterprise security buyer. Optiv works with organizations that have mature, complex security programs; exactly the kind of environment where Legion's approach of learning from bespoke processes is most valuable.
Enterprise security leaders look to trusted advisors to help them evaluate fit, plan implementation, and optimize outcomes over time. Optiv's position in the market as an integrator with deep relationships and deep domain expertise makes them uniquely positioned to bring best-in-breed solutions to the organizations that need it most and to help them get maximum value from it.
This partnership reflects something we're hearing consistently in the market: enterprises want agentic security, but they want it on their terms. They want AI that understands their environment before it acts in it. They want partners who can help them think through where automation should start, how to build confidence in the system over time, and how to expand from their first use cases into a broader program.
That's exactly what this partnership is designed to deliver.
What It Signals More Broadly
The Optiv partnership is a data point in a larger trend. Channel partners; the integrators, MSSPs, and advisors who sit closest to enterprise security buyers, are increasingly being asked about agentic security. Their clients want to know what's real, what's ready, and what actually works in complex environments.
For Legion, this is an important milestone in building the ecosystem that enterprise agentic security requires. We're grateful to the Optiv team for their partnership and excited about what we'll build together. And for enterprise security leaders who have been watching the agentic security space and wondering what a path to trusted AI adoption actually looks like, we'd love to show you.
Interested in learning how Legion Security and Optiv can help your organization automate, scale, and elevate your security posture? Get in touch.

Legion Security is now an Optiv Authorized Partner. Enterprise security teams can now deploy agentic AI for security operations that understands and optimizes agentic workflows without integrations, black boxes, or needing to ask teams to change how they work.
I was there, I sat in every SOC seat out there…
A SOC analyst grinding through alert queues at 2am. Part of an Incident Response team leading running war rooms. A SOC manager in Monday morning stand-ups asking what we learned this week while staring at blank faces.
Every single role. Every single day. And the one thing that never changed across any of them?
The insights, recommendations, self improvement, the de-facto SOC continuous improvement action items were disappearing. Seating documented in a case log for no one to action upon, trapped inside closed tickets that live in a backlog nobody rarely reopens.
I know the why and I feel the overwhelming operations, which is why I’m offering a practical solution for how to continuously improve your SOC with the valuable insights coming out of your investigations.
The Hidden Goldmine You're Sitting On
Every ticket your team closes tells a story. It's not just that an alert fired, then an analyst investigated and eventually closed. There are powerful signals buried in those notes, whether it's a tool with overly noisy alerts, a gap in your email gateway rules, or the same user clicking a phishing link for the third month in a row.
Your tier 1 all the way to your tier 5 analysts and IR responders are generating intelligence every single shift and with every single incident. They know things and they're writing them down. It's useful information but these notes get buried and never read again.
It's a sad truth... I know because I've been in those weekly SOC meetings, I was running them.
It's not a people problem, rather, it's a system problem.
The Weekly Report Trap
The thing people look to as the standard fix is the weekly report. In theory it's elegant: senior analysts summarize the week, extract the learnings, feed them back into tier 1 runbooks and detection improvements. On paper, it's a proper feedback loop.
In practice, it becomes the task that either gets rushed on Friday afternoon or simply doesn't happen. It's for good reason too! Your senior analysts are already stretched because on top of everything they need to do for their jobs, they're also being asked to synthesize everything in themes. You either get a half-hearted copy-paste of ticket titles, or, more likely, you get nothing.
Teams try rotation where everyone takes a turn on the ferris wheel. But in doing so, you face losing important insights and information, not to mention a lack of consistency.
Now add a follow-the-sun operation to this. APAC closes tickets while EMEA is asleep. EMEA handles incidents while Americas is offline. By the time anyone tries to compile a summary, they're working with fragments. Nobody has the full picture. The patterns that only emerge when you look across all shifts stay invisible.
Wait, Can't AI Can Solve This Pretty Easily?
When capable LLMs became available, I thought this was finally solved. Just feed all the investigation summaries in, ask for a weekly report. Done? Not so fast... here's what actually happened.
First attempt: I gave the best LLM models that money can buy more than 250 investigation summaries and asked for a consolidated report. But what I got back was a mess.
What I saw were recommendations repeated five times just with slightly different wording. Severity assessments that made no sense and my “favorite” recommendations that are not feasible, for example “Tune your EDR machine learning to reduce false positives of macro xlsx files”.
No traceability whatsoever, no way to tie anything back to the original investigation and forget about cross referencing with similar recommendations.
Second attempt: I went deep on prompt engineering. Longer prompts. More detailed. With examples. The results improved marginally, but the ceiling was surprisingly low.
The fundamental issue is that when you dump a large context with complex requirements into a single LLM call, it can't hold everything in working memory. It forgets constraints from earlier in the prompt. It hallucinates connections between unrelated incidents. Severity levels come out inconsistent.
One-shot approaches get you mediocre fast. They don't get you useful.
The Breakthrough: Think Multi-Step, Not Prompt
The shift that changed everything was stopping thinking about this as one task and starting to think about it as a multi-step pipeline.
When an experienced analyst writes a weekly report, they don't try to do it all at once. They read, they group, they prioritize, they write. Multiple steps. Each one is different.
So I built it that way.
The 6-step pipeline
Step 1: Classification
The first step does one thing and one thing only. It extracts and categorizes recommendations from raw investigation summaries. It looks for whatever your analysts call them: Recommendations, Do Better, Action Items, Next Steps. It pulls each one out and assigns it to a category: detection, prevention and process improvements.
No dedupe. No severity. Just extraction, done well.
Step 2: Feasibility Assessment
Now we evaluate each recommendation against practical reality. Can this actually be implemented? Is it a quick win or a multi-quarter project? Does it require resources you don't have?
This is also where web search earns its keep. When a recommendation references a specific product or vendor, the model can look up current best practices, product documentations, tech community discussions and verify the suggested configuration actually exists and is supported. Without this, you get generic, often infeasible advice. With it, you get grounded recommendations.
Make sure to use an LLM model that has web search capability via API calls.
Step 3: Citation Attachment
Before touching deduplication, every recommendation gets linked back to its source investigation. This is non-negotiable for a report anyone will actually act on. When a SOC manager reads and SOC teams attempt recommendation implementation, they need to know which investigations triggered that and value with volume justification to it. Otherwise it's just noise or worse, it might break business operations.
Step 4: Deduplication
Three analysts working three separate investigations but same use case, all recommend the same prevention improvement. Without deduplication, you get three entries saying the same thing with slightly different wording. With it, you get one consolidated recommendation that shows it came from three independent investigations, which is actually a stronger signal.
Citations from all source recommendations get merged. Nothing is lost.
Step 5: Severity Classification
Now, with duplicates consolidated, we can assign severity levels that actually mean something. The model evaluates security impact per your instructions, weights and SOC defined severities for each use case. Not how urgent did the analyst feel when writing this, but what is the actual risk if this doesn't get addressed built on your SOC knowledge base.
Separating this from extraction forces objectivity. If you try to assign severity while also pulling recommendations from raw notes, the analyst's tone bleeds in and skews the assessment.
Step 6: Report Generation
Everything feeds into the final structure. The model has category breakdown, feasibility assessments, severity levels, citation references. It produces a coherent report with an executive summary and recommendations sorted by severity, with enough context to actually act on. Also comparing recommendations week on week to get remediation/implementation progress for repeated action items.
Add another layer of disregard recommendations and you have a magnificent mechanism.
No LLM at this stage, actually. It's programmatic and deterministic. It assigns citation letters for easy grounding and reference of recommendation with feasibility (A, B, C...), builds the reasoning section for each recommendation, and outputs clean JSON ready for whatever you want to do with it.
Why This Architecture Actually Works
The goal is to achieve focused context at each step. Instead of one massive prompt juggling ten objectives, each step gets only what it needs. Fewer constraints to forget.
Modular iteration is the name of the game here. When severity ratings were inconsistent, I refined only the severity prompt. When analysts switched from Recommendations to Do Better as their section header, I updated only the classification step and nothing else broke.
Inspectable intermediate outputs. Between every step, results are saved. If something looks wrong in the final report, you can trace back through the pipeline and find exactly where it broke. Debugging is possible, which is not nothing.
Web search in the right place. Not as a general capability, but specifically in the feasibility step where it does the most work. Validating that a recommended configuration actually exists changes the quality of the output completely.
The Payoff
Your analysts don't change anything, they can run the same investigations, keep the same ticket notes they're already writing. The pipeline simply runs against their existing documentation.
The output is consistent. Same structure, same categories, same severity criteria, every week. You can compare week over week and actually spot trends. You can see if the same recommendations keep surfacing, which means they're not getting actioned, which is itself a signal.
The feedback loop that should have existed closes automatically. Tier 2 findings reach tier 1. Detection gaps surface. The Monday morning question about what we learned has an answer.
Build it or use it
Building this right takes time. Getting prompts tuned for the variety in how analysts write, handling edge cases, making it robust across different ticketing systems. It's not weekend work.
If you want to build it yourself: start with extraction only. Get that reliable first. Then add deduplication. Then severity. Don't try to build the whole thing at once.
If you'd rather not build tooling while also running a SOC, this is exactly what we built at Legion Security. Already tuned across real SOC environments, connected to your existing ticketing system, your analysts change nothing.
Either way: stop burying the intelligence your team generates every day.
Your team is learning constantly. Those lessons deserve to surface.
Written by someone who's been the analyst, the IR lead, and the manager staring at the empty Monday morning whiteboard.

SOC continuous improvement fails when insights get buried in closed tickets. Learn a 6-step LLM pipeline that turns investigation notes into action.
Legion Security is Now Available on Google Cloud Marketplace
Security operations were built around human investigators. Skilled analysts, working manually across dozens of tools, piecing together evidence, making judgment calls, closing cases. But as alert volumes outpaced human capacity, institutional knowledge became a bottleneck, and the complexity of the modern enterprise made scaling impossible. The industry responded with more headcount, more tools, more automation. None of it solved the fundamental problem.
Legion introduces a different operating model entirely.
What Legion Does
Legion observes how your analysts operate when running real investigations, learning your organizational context, tools, past cases, playbooks, runbooks and all other tribal knowledge in order to understand what an optimal investigation looks like for your environment. This is then turned into an easily editable and audible workflow which can be automated when you’re ready. Powered by Google Cloud's Gemini models, each workflow is executed by AI agents that reason through the evidence and provide a verdict and even remediate. This is all accomplished with no manual playbook writing or need to document predefined rules.
But legion goes well beyond workflow creation. As Legion builds trust in its performance, teams can choose to keep a human in the loop to approve every decision or have Legion operate fully autonomously reducing MTTR eliminating MTTA, allowing analysts to focus on more novel investigations that are becoming more and more common in the world of AI.
Memory: The Compounding Advantage
Every investigation Legion conducts makes it smarter. A persistent memory layer continuously captures context from previous cases, your SOC knowledge base, and direct analyst feedback, feeding all of it back into future investigations and decisions. Institutional knowledge that once lived in the heads of your most experienced analysts becomes a permanent, improving organizational asset. The more Legion works, the better it gets. That's not a feature. That's a compounding strategic advantage.
Zero Integrations. Immediate Value.
Most security automation platforms fail at the same hurdle: integrations. Enterprises face months of API work, custom connectors, and professional services before anything runs in production, or are forced to adopt entirely new tools and processes, something most complex enterprises simply can't do.
Legion operates natively in the browser, which means it works across your entire security stack, from threat intel platforms to legacy internal tools, without any API configuration. If your analysts can open it in a browser, Legion can learn from it, generate workflows from it, and execute investigations through it.
Proven Results at Scale
The impact Legion delivers isn't theoretical:
As the head of Security at Virgin Money put it, Legion is “like evolving from handcrafted systems to precision manufacturing aligned to our flow (except) faster, repeatable and secure”.
Legion works with the worlds largest enterprises and delivers strong results:
- A large insurance organization automated 24,000 investigations and cut mean time to respond from 20 minutes to 2 minutes.
- WELL Health Technologies reduced investigation times by 81%, allowing existing analysts to handle significantly higher alert volumes without additional headcount.
- The University of Tulsa cut investigation times in half, enabling their team to overcome capacity limits with the staff they already had.
Across deployments, Legion reduces mean time to investigate by up to 85% and response times by up to 90%.
Built on Google Cloud
Legion's integration with Google Cloud goes deeper than the Marketplace listing. The platform runs on Google Cloud infrastructure and leverages Gemini models to power its AI reasoning, combining Legion's browser-native architecture with Google Cloud's security, scale, and model quality.
For organizations already invested in Google Cloud and Google SecOps, Legion extends that ecosystem directly into the analyst workflow.
Who It's For
Legion is purpose-built for enterprise security operations teams, CISOs, VPs of Information Security, SOC Directors, and Security Operations Managers at organizations running in-house SOCs. If your team is dealing with any of the following, Legion was built for you:
- Alert volumes that have outpaced your team's capacity
- Analyst burnout from manual, repetitive investigation work
- Institutional knowledge that walks out the door when senior analysts do
- Automation gaps caused by complex integration requirements
Available Now on Google Cloud Marketplace
Legion Security is available today on Google Cloud Marketplace, allowing customers to apply their spend toward their annual Google contract and simplify procurement. For security teams ready to move beyond the limits of traditional operations, this is where that transformation begins.

Legion is officially on the Google Cloud Marketplace.
Introduction
TL;DR: Using LLMs to automate security alert correlation won’t work when they’re fed unmanaged IOCs including emails, URLs, IPs, domains, and hostnames. These inflate token costs, produce inconsistent references, and break structured output and automation reliability. Legion Security’s IOC indexing system replaces raw indicators with compact symbolic references that the model reuses throughout its reasoning. Across 100 evaluation runs, this took JSON validity from ~80% to 100% and IOC reference compliance to 100%, resulting in the ability to reliably automate security alert correlation.
Automating security alert correlation and other modern security investigations with LLM-based agents means using an agentic LLM to power a multi-step security investigation.
A typical workflow begins with an alert - say, a reported phishing email - and the agent iteratively queries tools such as Microsoft Defender Threat Explorer, Splunk, or CrowdStrike to gather evidence, assess scope, and recommend containment actions.
At each step, the agent receives query results containing raw IOCs: sender addresses, embedded URLs, source IPs, recipient domains, and device hostnames. It must reason about these indicators, decide whether to refine its search or conclude the investigation, and return its findings as structured output.
Without any re-engineering of indicators of compromise (IOCs) an agentic LLM can work well for short investigations.
But as the number of steps in an investigation grows, an issue with alert automation emerges.
The agent's context window fills with repeated, verbose indicator values, the model begins echoing raw IOCs inconsistently, and the structured outputs it produces become increasingly fragile.

[Figure 1: High-level architecture of an AI-driven investigation agent.]
Consider a phishing investigation that proceeds through four steps:
- Initial query: Search for emails from a reported sender to a specific recipient. The results contain the sender's email address and a handful of URLs.
- Scope expansion: Search for all emails from the same sender across the organization. The results return 22 emails with SharePoint URLs, tracking links, and font-file references.
- URL analysis: Search by specific URLs found in step 2. Additional domains and redirects surface.
- Conclusion: The agent summarizes its findings and lists all relevant IOCs.
By step 4, the agent's prompt contains the full history of steps 1 through 3 - including every raw URL, email address, and domain mentioned in each step's results and the agent's own reasoning. Some of these URLs are long tracking links with base64-encoded parameters, easily exceeding 200 characters each.
This accumulation creates three concrete problems.
- Token bloat. Raw IOC values, particularly URLs with tracking parameters and encoded payloads, consumed a disproportionate share of the context window. A single newsletter email might contain 30+ URLs, each repeated in the query results, the agent's reasoning, and the indicators list, tripling the token cost per IOC, per step.
- Over-reporting. When asked to list relevant indicators, the model would frequently dump every IOC it had ever seen into the response - even when the current step involved only one or two. In one case, an agent listed all 145 email addresses from its registry when the current query concerned a single sender.
- Structural fragility. Query results from security tools sometimes contained comma-separated URL lists embedded in strings. When the model attempted to reproduce these in its JSON output, it produced malformed structures - unescaped commas, broken string boundaries, and invalid nesting. In our baseline evaluation, only approximately 80% of model responses parsed as valid JSON.
Legion AI’s Approach To Building Better AI Security Alert Correlation
We address the AI security workflow problems that emerge from complex invesigations with a three-part system:
- A unified IOC manager that extracts and indexes indicators.
- An IOC prompt adjustment that instructs the model on how to use indexed references.
- A preprocessing step that cleans malformed tool output before it reaches the model.
IOC Extraction and Indexing
The core of the system is an IOC manager that maintains a registry of all indicators encountered during an investigation. When new text enters the pipeline - whether from tool query results or from the agent's own prior reasoning - the manager scans it using a set of type-specific patterns covering URLs, email addresses, IPv4 addresses, file hashes, hostnames, and domains.
Each newly discovered IOC is assigned a compact symbolic reference following a consistent naming convention: the first email becomes EMAIL01, the first URL becomes URL01, the second domain becomes DOMAIN02, and so on. The original value is stored in the registry, and all occurrences in the text are replaced with the corresponding reference.
Extraction order matters. URLs are processed first because a URL contains both a domain and potentially an IP address. By extracting URLs before domains and IPs, we prevent the system from fragmenting a single indicator into multiple overlapping entries.
The manager also performs selective extraction. In a typical prompt, the first section contains static task instructions - tool descriptions, output format specifications, and investigation guidelines. IOC extraction is applied only to the dynamic sections (step history and query results), leaving instruction text unchanged. This prevents false positives from example IOCs embedded in the prompt template.
Deduplication is handled through a value-to-reference mapping. If the same IOC appears in step 1 and again in step 3, it receives the same reference both times, ensuring consistent tracking across the entire investigation.

[Figure 3: The IOC extraction pipeline.]
IOC Prompt Adjustment
Extraction alone is not sufficient. Even when the input prompt uses symbolic references, the model may revert to generating raw IOC values in its output, particularly if the system prompt or prior conversation history contains raw values, or if the model has seen the actual value during context processing.
To address this, we developed an IOC prompt adjustment, a compact, structured appendix appended to the user prompt that explicitly instructs the model on how to handle IOCs. The adjustment establishes three rules:
In reasoning
Always use symbolic references. Never write raw email addresses, URLs, IP addresses, or domains. Instead of writing a raw sender address followed by a description of the campaign, use the corresponding reference identifier throughout.
In the indicators field
Distinguish between known and new IOCs.
For indicators already present in the registry, use the symbolic reference. For indicators being reported for the first time, newly discovered in the current step's results, use the actual value, so it can be added to the registry for subsequent steps.
Relevance filtering
Only include IOCs that are directly relevant to the current investigation step. Do not copy all registry entries into every response.
The IOC prompt adjustment includes a populated copy of the current IOC registry, mapping each reference to its actual value, so the model can look up identifiers when constructing its reasoning. It also provides correct and incorrect examples, a validation checklist, and explicit rejection criteria.
We tested two versions of the IOC prompt adjustment:
- A comprehensive version with extensive examples and redundant emphasis
- An optimized version that distills the same rules more concisely.
Both achieved equivalent compliance rates, suggesting that clarity of instruction matters more than volume of repetition.
Input Preprocessing
The second component addresses a problem upstream of the model: malformed tool output.
Security tool APIs sometimes return URL lists as comma-separated values within a single string field, rather than as properly structured arrays. When passed through to the model as-is, these malformed strings caused structured output generation failures.
Our preprocessing step detects comma-separated URL patterns in query results and reformats them into clean, numbered lists before the text reaches the model. This small transformation, applied before IOC extraction, resolved the structured output validity issue independently of the other components.
Security Alert Correlation Evaluation
Setup
We evaluated the system using 10 real-world investigation traces captured from production. Each trace represents a complete phishing investigation conducted through several security tools, containing the system prompt, user prompt with step history, and the raw query results that the model must reason about.
For each trace, we ran 10 iterations with the same prompt configuration, measuring two metrics:
- JSON validity: Whether the model's response parsed as valid a structured output.
- IOC reference compliance: Whether the response used symbolic references exclusively in its reasoning field (no raw IOC values) and correctly distinguished between known references and new actual values in its indicators field.
We tested four configurations to isolate the contribution of each component.
Results
JSON validity
The baseline configuration produced valid JSON in approximately 80% of responses. Adding URL preprocessing alone brought this to nearly 100%, confirming that malformed tool output - not model capability - was the root cause of parsing failures. All configurations that included the enforcer achieved 100% validity.
IOC compliance
Without the IOC prompt adjustment, the model never spontaneously adopted symbolic references, compliance was 0% regardless of whether the input text had been processed by the IOC manager. With the prompt adjustment, compliance jumped to 100% across all traces and iterations. This held for both the comprehensive and optimized prompt adjustment variants.
Component independence
The results reveal a clean separation of concerns: URL preprocessing fixes JSON validity, the IOC manager fixes IOC compliance, and provides the underlying registry and extraction infrastructure that makes both possible.
Qualitative Observations
Beyond the quantitative metrics, we observed several qualitative improvements:
- Reduced prompt size. Replacing verbose URLs (some exceeding 200 characters) with compact references, meaningfully reduced token consumption in the step history, particularly for investigations involving newsletter or marketing emails with numerous tracking links.
- Consistent cross-step tracking. The registry ensured that the same IOC received the same reference throughout the investigation, this is particularly helpful with IOCs referenced throughout multiple steps in the investigation.
- Focused indicator reporting. With the IOC manager's relevance-filtering instruction, the model stopped dumping entire registries into its responses. Indicator lists became proportional to the current step's scope rather than the investigation's total history.
AI Security Alert Correlation Discussion
Why extraction without the IOC prompt adjustment fails
A natural question is why input-side extraction alone does not work. If the prompt already contains a symbolic reference instead of a raw email address, why does the model still generate raw values in its output?
The answer lies in how LLMs process context.
The model has access to the full prompt, including sections where the actual IOC value may still appear, task-specific inputs, quoted alert descriptions, or the registry itself.
More fundamentally, the model's training distribution contains overwhelmingly more examples of raw IOC values than of symbolic reference systems. Without explicit instruction, the model defaults to the more familiar pattern.
This finding has a broader implication for LLM-based agent design: transforming the input is necessary but not sufficient when you need the model to adopt a non-default output convention. Explicit behavioral instruction, the IOC prompt adjustment, bridges the gap.
Limitations
Our evaluation has several limitations worth noting. All traces were drawn from a single investigation type (phishing via specific security tools). While the IOC types encountered are representative of broader security operations, there are additional evaluations to be done.
The evaluation was conducted with a single model (chatGPT4.1). Different models may exhibit different compliance characteristics, and the prompt mechanism may need tuning for models with different instruction-following tendencies.
Finally, our compliance metric is binary - a response either uses references correctly or it does not. A more granular metric could capture partial compliance and might reveal subtler performance trends across model versions or investigation complexities.
Conclusion
We presented an IOC indexing system for AI-driven security investigations that addresses three interrelated problems: token bloat from repeated raw indicator values, inconsistent IOC tracking across investigation steps, and structural fragility in model-generated structured outputs.
The system combines automated IOC extraction with symbolic reference assignment, explicit behavioral guidance through prompt engineering, and input preprocessing to handle malformed tool output. Across 100 evaluation runs on 10 production investigation traces, the full system achieved 100% JSON validity and 100% IOC reference compliance, up from approximately 80% and 0%, respectively, at baseline.
The key insight is that managing IOCs in the context of LLM-based agents requires intervention at both the input and output stages. Extraction and indexing normalize the input, but only explicit prompt-level guidance ensures the model adopts the reference convention in its generated output. Neither component alone is sufficient; together, they eliminate the problem entirely.
As SOC automation platforms handle increasingly complex, multi-step investigations, structured approaches to managing the information that flows through the agent's context window become essential. IOC indexing is one instance of a more general pattern: giving the agent a well-organized working memory that scales with investigation complexity rather than against it.
.png)
Automation of security alert correlation with AI (LLMs) depends on replacing raw indicators with compact symbolic references.
Abstract & Data Summary
We gathered and manually annotated a dataset of 196 hard triage decisions from real-world security investigations, covering a wide range of outcomes, including benign, malicious, and false positives. After cleaning the dataset by removing mock runs and cases with missing information or incorrect workflow execution, the remaining 163 examples were grouped into use case categories to form a high-quality cohort. We then evaluated LLMs on the dataset overall and per use-case category and found that Gemini 3 Pro performs best overall, though the best LLM varies by use case category.
Model performance by use case category:
If you’d like to understand our full research methodology, read on.
*Note: since this blog was authored, several new model families have been released. While the results have remained broadly stable, particularly among the best and worst performers, updated research may be required for a nuanced understanding of the performance differences amongst the rest.
Data Collection
The dataset was constructed from security investigations from eight US-based customers.The evaluation is conducted in a secure, federated way, without mixing customer data, only reporting summary statistics from each customer tenant.
To create a challenging evaluation, we over-weighted cases in which the analyst dis-agreed with the model - so the error rate is inflated here.
The investigations were conducted automatically according to predefined, customer-specific workflows, each of which contained at least one triage decision node. A triage decision node is a decision point within a workflow, where an LLM chooses a decision from among a list of provided decision options, given the information that was gathered in the workflow up until that point.
At each decision node, the LLM used in production selected a classification decision from a list of workflow-specific decision options and provided the reasoning for its decision, based on a summary of the steps completed until that point in the investigation.
For each investigation containing at least one decision node, we collected the following information from production session logs:
- A summary of the workflow steps up until the decision node, including tool name, step description, and step outputs
- Organization-specific knowledge, written by the customer and containing a title, description, and data
- The set of available decision options at the decision node
- The model's selected decision in production, as well as the reasoning and detailed reasoning for the decision
- The decision option selected by the customer
- Feedback text written by the customer for the decision
Here is an example workflow diagram:

Quality Control
An expert cybersecurity analyst annotated the 196 decision examples with reasoning tags to explain the production and customer decisions, and label whether disagreements are explained by an analyst-error, mistaken reasoning by the AI or missing data / steps in the workflow.
Examples tagged with "Workflow ran correctly but missing information" or "Workflow ran incorrectly" were removed from the dataset. Two additional examples with the use case titled "Workshop" were removed, as these were mock runs. For the remaining examples, the workflow ran correctly and was not missing information.
Triage Decision Distribution
By Label
Across the filtered dataset, the workflows contained 27 distinct normalized decision labels, which we grouped into the following buckets: False Positive, True Positive, Requires Review, and Other. The distribution of the labels is shown below:
The final evaluation dataset contains data from eight customers. The table below shows the number of annotated decision examples per customer and the tools used in each environment.
Use Case Distribution
We consolidated the use cases into 3 categories to consolidate our findings. Below is the map from the consolidated categories to the original use cases, as well as the distribution of the dataset over the consolidated categories.
Confusion Matrix
Below is a confusion matrix between the expert analyst annotations and the recommendations our system makes. We prompt the models to be careful and escalate when they are not sure.
Results
Over all use cases (including those without a use case name), Gemini 3 Pro had the highest performance at 74.8%, with GPT-4.1 and Opus 4.5 tied for second.
Phishing Results:
On the phishing use cases, Gemini 3 Pro performed the best, followed by Opus 4.5.
Account Takeover Results:
Sonnet 4 and GPT-4.1 were tied for best on the account takeover use cases.
Network Results:
Opus 4.5 and GPT-4.1 were tied for best on the network use cases.
Conclusion & Recommendation
We gathered and annotated 163 triage decisions from real-world security investigations. We characterized the use case distribution, and grouped the use cases according to common categories. We then benchmarked large language models across each use case category and the full dataset. We found that Gemini 3 Pro performs best overall. Per use case category, Gemini 3 Pro gives the best performance on phishing, Sonnet 4 and GPT-4.1 are tied for best on account takeover, and Opus 4.5 and GPT-4.1 are tied for best on network. Based on our results, we recommend that security teams test models for different scenarios to find the solution that works best for their use case, different models are good at different things and the only way to know which model works best for your use-cases it to run formal evaluation - or, you can trust us! Our research team in Legion is constantly evaluating new models and improvements to our triage pipelines.

We benchmarked leading LLMs on 163 real-world security triage decisions across phishing, account takeover, and network use cases. See which models performed best and why the answer depends on your use case






