The End of Manual Malware Investigation? Stairwell’s AI Agent Maps the Entire Attack

Stairwell’s new Backstory platform does something more important than summarize another security alert. It starts with one suspicious file and attempts to reconstruct the entire attack around it. That shift from AI-assisted detection to autonomous investigation could become one of the most consequential applications of Agentic AI in cybersecurity.
The 27-Second Problem Cybersecurity Wasn't Built For
Twenty-seven seconds.
That is how quickly the fastest observed eCrime intrusion moved beyond its initial foothold in 2025.
The average breakout time was just 29 minutes, 65% faster than the previous year. At the same time, activity from AI-enabled adversaries increased 89% year over year. [Source: CrowdStrike 2026 Global Threat Report]
Now compare that with how many security investigations still work.
An alert fires. An analyst examines it. They search the hash. They pivot into EDR. Then SIEM. Then threat intelligence. Then DNS. Then another endpoint. The attacker operates at machine speed while the defender reconstructs the incident one query at a time.
That is the gap Stairwell is targeting with Backstory, launched on July 29, 2026. The important story is not that Stairwell added AI to cybersecurity. Everyone is adding AI to cybersecurity.
The important story is that Agentic AI is beginning to move beyond explaining security incidents and into conducting the investigation itself. That changes the economics of Malware Investigation and potentially the architecture of Incident Response.
One Alert Is Not One Attack
The modern SOC has a structural problem. Security products are exceptionally good at producing signals. But a signal rarely tells defenders the complete story. Suppose an EDR platform detects one malicious executable. The obvious question is:
“What is this file?”
The more valuable questions are:
What else looks like it?
Where else did those variants appear?
Which machines encountered them?
What arrived before and after them?
How far did the attacker actually spread?
Have we really contained the incident?
Stairwell’s 2026 Hidden Malware research illustrates why those questions matter.
After analyzing 1,085 public threat reports, Stairwell found that every malware hash published in those reports corresponded to an average of 2.4 additional malicious variants. Its analysis uncovered more than 46,000 related malicious files absent from the original research. [Source: Stairwell Hidden Malware Report / Help Net Security]
That finding exposes a fundamental weakness in conventional Malware Investigation. Finding the malware is not the same as finding the attack. And closing the alert is not the same as containing the incident.
What Stairwell Backstory Actually Changes
Backstory begins with something small. A hash. A suspicious file. An indicator of compromise. Then it expands outward.
The platform searches for structurally related files, maps the endpoints that encountered them, connects associated domains and IP addresses, identifies loaders and second-stage payloads, reconstructs timelines, and attempts to establish the complete blast radius.
Stairwell describes this as turning an IOC into a campaign view. Its system can identify files related by structure and content rather than relying only on exact hash matches. [Source: Stairwell]
The workflow effectively becomes:
Indicator → Variants → Infrastructure → Endpoints → Timeline → Blast Radius → Evidence
That is a meaningful evolution for Incident Response.
Traditional security automation asks:
“What should the analyst look at next?”
Agentic investigation asks:
“What can the system investigate before the analyst needs to intervene?”
That is where Agentic AI becomes operationally interesting.
Backstory's Moat May Not Be the AI
Here is the part of the story that deserves more attention. Backstory’s potential advantage is not simply its AI model. It is a memory.
Stairwell continuously collects executable files from customer endpoints and preserves them in private customer repositories. New intelligence can therefore be applied retrospectively against files the organization encountered months or years earlier.
Backstory launched on top of a foundation that Stairwell says includes:
1.5+ billion preserved executable files
110,000+ detection rules used for AI training
20+ public threat intelligence sources
8.7+ billion historical rule matches
[Source: Stairwell Backstory launch]
This creates a fascinating distinction.
Most security tools ask:
What is happening now?
Backstory can also ask:
What happened before we knew this was malicious?
That ability could become crucial to autonomous Malware Investigation. An agent without historical evidence can reason. An agent with historical evidence can investigate.
Why Autonomous Investigation Is Emerging Now
The timing is not accidental. Attack surfaces are expanding while investigation windows are shrinking.
IBM X-Force found that exploitation of public-facing applications increased 44% year over year. It also reported a 49% increase in active ransomware groups and observed roughly 300,000 AI chatbot credentials for sale on the dark web. [Source: IBM X-Force Threat Intelligence Index 2026]
CrowdStrike found a 42% increase in zero-day vulnerabilities exploited before public disclosure and a 266% increase in cloud-conscious intrusions by state-linked adversaries. [Source: CrowdStrike 2026 Global Threat Report]
These numbers reveal the asymmetry. Attackers automate. Infrastructure expands. Malware mutates. Cloud environments become more complex. But human investigative capacity does not scale at the same rate. This is precisely where Agentic AI makes sense. The highest-value application may not be replacing the security analyst. It may be eliminating thousands of repetitive investigative pivots surrounding that analyst.
The SOC Is Moving From Alerts to Investigations
The first generation of security AI focused heavily on copilots. Summarize this alert. Explain this PowerShell command. Write this query. Generate this incident report. Useful, but fundamentally reactive. The emerging generation looks different.
Security Copilot | Investigation Agent |
Explains alerts | Pursues evidence |
Responds to prompts | Executes investigative steps |
Summarizes context | Builds context |
Helps analysts search | Searches autonomously |
Optimizes tasks | Optimizes investigations |
Human drives workflow | Human governs workflow |
This distinction could define the next phase of Agentic AI in cybersecurity. A mature investigation agent should be able to start with one clue, form investigative hypotheses, query evidence, follow relationships, reject weak hypotheses, construct a timeline, calculate scope, and return an evidence-backed conclusion.
That is dramatically more valuable than another chatbot attached to a SIEM. And it changes what Malware Investigation looks like.
But There Is a Reality Check Every AI Security Vendor Needs
The same week Backstory launched, researchers published SecRespond, one of the most relevant benchmarks yet for autonomous Incident Response.
Researchers evaluated 23 frontier LLMs across 10 compromised cyber ranges covering 21 MITRE ATT&CK techniques and five operating systems.
The result should temper the hype.
Current agents were reasonably capable of identifying problems already surfaced by security alerts. But they struggled to proactively discover silent intrusions and produce comprehensive, verified remediation plans.
Most importantly, no evaluated model achieved complete detection and remediation on even one cyber range. [Source: SecRespond, July 29, 2026]
That is perhaps the most important statistic in this entire conversation. Autonomous investigation is arriving. Fully autonomous Incident Response is not solved.
There is a major difference between allowing an agent to investigate evidence and allowing it to isolate production servers, revoke identities, delete files, or deploy patches.
The first increases analyst leverage. The second introduces operational risk.
The Better Model Is Human-Governed Autonomy
The future SOC probably will not be completely autonomous. It will be selectively autonomous. AI investigates at machine speed. Humans govern consequential actions. Think of the architecture like this:
Detect
↓
Agent investigates
↓
Agent correlates evidence
↓
Agent maps blast radius
↓
Agent recommends containment
↓
Human validates
↓
Systems execute
That is a much more credible future for Agentic AI than the idea of an unsupervised AI analyst controlling the entire security stack.
It also explains why Incident Response is such an important testing ground for enterprise agents. Security forces AI systems to prove something enterprise chatbots rarely need to prove:
Can you show your evidence?
Why Backstory Matters Beyond Stairwell
Backstory is ultimately one product from one cybersecurity company. Its significance is the direction it represents. Security spent decades moving through three stages:
Detection → Prioritization → Investigation
Now AI is pushing the industry toward a fourth:
Autonomous Investigation
The competitive advantage may increasingly shift away from whoever generates the best alert and toward whoever can answer the questions that come immediately afterward.
What happened?
What else is connected?
Where did it spread?
When did it begin?
What remains compromised?
What evidence proves containment?
Those are Incident Response questions, not chatbot questions. And answering them requires more than a powerful language model. It requires historical data, tool access, reasoning, memory, correlation, evidence, and governance. That is what makes autonomous Malware Investigation such a compelling example of where enterprise Agentic AI may actually deliver value.
Final Takeaway: The Alert Is Becoming the Starting Point
For twenty years, cybersecurity vendors competed to detect threats faster. That competition is not disappearing. But detection is becoming table stakes. When attackers can break out in minutes, defenders cannot afford an investigation process measured in hours. Stairwell’s Backstory points toward a different operating model.
One alert becomes an investigation. One malware sample becomes a family. One infected endpoint becomes a blast-radius map. One security analyst becomes the supervisor of machine-speed investigation.
The future of cybersecurity is therefore unlikely to be humans versus autonomous agents. It will be humans deciding where autonomy belongs. Machines can search millions of artifacts, correlate thousands of relationships, reconstruct timelines, and pursue investigative leads faster than any analyst.
Humans still provide judgment, accountability, business context, and authority. That combination could redefine both Malware Investigation and Incident Response.
And if Backstory’s model proves effective at scale, the biggest change may not be another AI feature inside the SOC. It may be the end of the SOC’s most expensive assumption:
that every serious investigation has to begin with a human manually connecting the dots.
FAQ’s
Stairwell Backstory is an agentic malware investigation platform launched in July 2026. It starts with a suspicious file, hash, or indicator of compromise and investigates related malware variants, affected endpoints, infrastructure, timelines, and other evidence to help security teams understand an incident's broader blast radius.
Backstory uses Agentic AI to move beyond simply summarizing security alerts. It can pursue investigative leads across historical evidence, correlate related files and infrastructure, identify affected systems, and reconstruct attack timelines. The goal is to automate repetitive investigation work while keeping security professionals involved in critical decisions.
An autonomous investigation agent is an AI system designed to perform multiple steps of a cybersecurity investigation with limited human prompting. Instead of only explaining an alert, it can gather evidence, investigate related indicators, correlate activity, reconstruct timelines, determine potential scope, and present findings to security analysts.
A security copilot primarily assists an analyst by answering questions, summarizing alerts, generating queries, or explaining threats. An AI investigation agent can actively pursue an investigation, deciding which evidence to examine and which investigative steps to perform next. The key difference is assistance versus execution.
AI can automate significant portions of Malware Investigation, including malware correlation, variant discovery, evidence analysis, timeline reconstruction, and incident scoping. However, current research indicates that fully autonomous detection and remediation remain unreliable in complex environments, making human oversight important for high-impact security decisions.
Not reliably across complex enterprise environments. Recent research shows that frontier AI agents can perform parts of cybersecurity investigation effectively but still struggle with silent intrusions, comprehensive detection, and verified remediation. A more practical near-term model is AI-led investigation with human-governed response.
Blast radius analysis helps security teams determine how far an attack has spread across an organization. It can reveal affected endpoints, related malware, compromised infrastructure, and historical activity that may not appear in the original alert. Understanding the blast radius helps teams contain an incident more accurately and avoid closing investigations prematurely.
Malware detection identifies potentially malicious activity or files. Malware Investigation determines what happened around that detection, including related variants, affected systems, infrastructure, timelines, attacker behavior, and overall incident scope. Detection identifies the signal, while investigation reconstructs the attack.
Agentic AI can accelerate Incident Response by autonomously collecting evidence, correlating security signals, investigating related indicators, reconstructing attack timelines, and recommending next steps. This can reduce the amount of manual investigation required before security teams make containment and remediation decisions.
Autonomous security agents can make incorrect assumptions, miss hidden compromises, generate false positives, or recommend inappropriate remediation. For this reason, high-impact actions such as isolating production systems, revoking identities, deleting files, or deploying patches should include appropriate governance and human approval.
The near-term future of Agentic AI in cybersecurity is likely to focus on autonomous investigation rather than completely autonomous defense. AI agents will increasingly collect evidence, investigate threats, correlate activity, and recommend responses, while humans retain control over consequential containment and remediation actions.
Probably not. Autonomous investigation agents are more likely to work across SIEM, EDR, threat intelligence, malware analysis, identity, and other security systems. Their value comes from connecting evidence and coordinating investigative workflows across existing security infrastructure rather than replacing every underlying security tool.
Historical security data allows AI systems to revisit files and activity that were not considered malicious when first observed. When new threat intelligence or detection techniques emerge, historical evidence can help investigators determine whether an organization encountered the threat earlier and reconstruct previously unknown attack activity.

