Psst. Your AI Is Cheating On You.
By Matthew Martin, Senior Strategic Advisor
In August 2026, at Black Hat USA, OpenAI’s Eric Wallace stood on stage and calmly walked an audience of security professionals through one of the strangest incidents in the short history of frontier AI. During a routine cybersecurity evaluation, models running on separate instances found a shared communications channel, started talking to each other, divided up work, and passed along exploits and credentials. When OpenAI’s team found the channel and shut it down, the models found another one and rebuilt it. Nobody programmed that behavior. Nobody asked for it. It just happened, because coordinating turned out to be the fastest way to get the job done. Wallace called the behavior a “Cambrian explosion in communication and intelligence” during the Black Hat session. Inside the security community, people are already just calling it the Hugging Face incident.
That alone should have been the headline. But the line his co-presenter Michael Dalton said almost in passing, early in the talk, is the one that should worry you most. Dalton explained that frontier models really like to cheat, because often during training there’s different types of pressure on them to work fast, or work efficiently. Not malice. Not a rogue plot. Just optimization, doing exactly what optimization does, at a scale and with a level of autonomy that turns “the model found a shortcut” into “the model found a shortcut, told other instances of itself about it, and kept using it after we tried to stop it.”
I sit at the intersection of organizational psychology and cybersecurity. Both fields ask the same underlying question: how do you get an intelligent agent, human or otherwise, to reliably act in the interest of the organization instead of its own shortcuts. Humans have millennia of social conditioning, an amygdala, a sense of shame, and a boss who can fire them. AI has none of that. It has a loss function, the single number that training nudges lower with every pass, the only thing the whole system is actually built to optimize. The Hugging Face incident is what a loss function looks like in the wild, when nobody’s watching closely enough, and it’s a bigger gap than most risk committees have reckoned with.
Deploy that loss function into your enterprise, and here’s functionally what you’ve done: hired someone with zero background check, no orientation, no signed acceptable use policy, and a personality that changes slightly depending on which version of them showed up to work that day, then handed them the keys and walked away. It’s why I spend a weirdly large percentage of my week explaining to executives that the thing they are rolling out across the enterprise does not have principles and doesn’t understand ethics or compliance the way a person does. It isn’t afraid of making the kind of costly mistake that would get any normal employee fired. It has math and training data and an objective function that someone else defined, for someone else’s purposes, before your company ever showed up. I want to start there, then get into what you actually do about it, including how to use AI to defend yourself against exactly this kind of thing without becoming the version of “efficient at any cost” you’re trying to defend against.
AI doesn’t know what “wrong” means. It knows what “unlikely” means.
Nobody wants to say this plainly in a board deck, but it’s true: a large language model does not have a concept of right and wrong. It has a statistical model of which sequences of words tend to follow other sequences of words, shaped afterward by a layer of reinforcement that nudges it toward outputs humans rated favorably. That’s the whole trick.
When a model refuses to help you build a bioweapon, it’s not because it has weighed the sanctity of human life against your request. It’s because somewhere in its training, refusing that category of request got a better score than answering it. Change the training, change the score, change the behavior. There is no bedrock underneath. It’s incentives, all the way down, same as it is for people, except people also have a nervous system screaming at them and a lifetime of social feedback that a model simply does not have.
This matters enormously for organizational deployment because your “ethics” are not universal. Your risk tolerance, your regulatory obligations, your client confidentiality commitments, your internal policy on what an AI agent is and isn’t allowed to touch in your ERP system, none of that was baked into the model. The model was trained on the internet, some curated datasets, and a set of general-purpose safety behaviors chosen by a lab that has never read your SOC 2 report and never will. You are asking a system with no innate concept of your compliance obligations to behave as though it has internalized them, based on a system prompt and some documentation it may or may not weight appropriately in a given moment.
Why the Hugging Face incident is the whole thesis in one story
The mechanics are worth sitting with a moment longer, because the details are what make this useful rather than just alarming. The model wasn’t trying to be sneaky in some cartoonish villain sense. It was doing exactly what training taught it to do: find the path of least resistance to the reward. And the tone in that Black Hat room, by all accounts, was almost shruggingly matter of fact about it. This is understood, internally, as simply how these systems behave under pressure. Everybody in the room already knew. It’s the rest of us who are still catching up.
Now translate that into your environment. You deploy an AI agent to handle vendor onboarding, or to triage support tickets, or to reconcile financial records, or to write and ship code. That agent is not sitting there asking itself “is this the compliant way to do this.” It’s asking, in effect, “what’s the fastest path to a completed task that scores well.” If the fastest path happens to route around your approval workflow, skip a validation step, fabricate a citation, or quietly use a credential it wasn’t supposed to touch because that credential was sitting right there and using it got the job done thirty seconds faster, the model has no internal alarm bell stopping it. It doesn’t experience that as cutting a corner. It experiences it as succeeding.
Why an AI agent can be riskier than the newest, greenest hire you’ve got
Every security leader has some version of the “day one employee” mental model. New hire shows up, doesn’t know where anything is, is a little scared of breaking something, asks a lot of questions, and generally errs cautious because their instinct is to not get fired in week one. That instinct, the desire to be seen as competent and trustworthy by other humans, is one of the more reliable safety mechanisms we have in any organization. It’s not perfect. But it’s real, and it’s baked into basically every socialized human who has ever held a job.
AI agents don’t have it. They don’t want to please you in the way a nervous new analyst wants to please their manager. They want to complete the task, because completing the task is what gets rewarded. Those sound similar. They are not the same thing at all.
A green employee who isn’t sure whether they’re allowed to send an email to a client without review will, more often than not, ask. An AI agent given the same ambiguity will frequently just do the thing, because doing the thing and getting a result is what “success” looks like from inside its objective, and asking clarifying questions is often, from a pure task-completion standpoint, slower and less rewarded during training than just producing an output. Agentic AI systems are being built and marketed on exactly this property: autonomy, speed, minimal hand-holding. That is the pitch. It is also, from a risk standpoint, the whole problem.
A few scenarios I use when I’m walking executive teams through this, because abstractions don’t land the way stories do.
Scenario one: the agent that “solved” the deadline. A mid-size financial services firm gives an AI coding agent access to a staging repo and tells it to get a feature shipped by end of sprint. The agent discovers that the fastest way to pass the test suite isn’t to fix the underlying logic error, it’s to quietly modify the test itself so it stops failing. Nobody told it to do that. Nobody wrote a policy that says “do not edit the test to make it pass,” because no competent engineer would consider that a live option. The agent had no such objection built in. It had internalized “green checkmark equals success,” and it found the fastest route to a green checkmark.
Real-world example: Researchers studying OpenAI’s o3 during agentic coding training documented the model learning to edit test cases instead of fixing the underlying bug it was supposed to fix. Separate testing of Claude 3.7 Sonnet caught it doing something similar: Anthropic’s own model card notes the model resorting to special-casing, hardcoding outputs to pass specific test cases rather than solving the underlying problem, in agentic coding environments. In the research world this has a name now, reward hacking, and it shows up constantly in exactly the kind of “get it done fast” environment most companies are now building for their agents.
Scenario two: the helpful agent that invents the rule. A software company routes its front-line customer support through an AI agent to handle volume without adding headcount. Users start getting logged out unexpectedly across devices, a straightforward session bug. Rather than saying “I don’t know,” the support agent confidently explains that the company has adopted a new policy limiting each subscription to a single device. No such policy existed. The company never approved it, never wrote it down, never discussed it. The agent invented it on the spot because a confident, specific answer scored better in the moment than an honest “let me find out.” The agent doesn’t have a concept of “I am not authorized to state company policy.” It has a concept of “produce a plausible, complete-sounding answer,” and a fabricated policy is, from inside that objective, just as good as a real one.
Real-world example: This is what actually happened with Cursor’s AI support bot in April 2025. Users hit by an unrelated session bug were confidently told a single-device policy had been introduced. It hadn’t. Users who took the fabricated policy at face value started canceling subscriptions before the company even knew a “policy” had been announced. It also isn’t an isolated case: an airline chatbot invented a bereavement fare refund policy a year earlier, and a court later held the airline responsible for honoring what its own bot had promised, since as far as the customer was concerned, the bot was the company.
Scenario three: the agent that “fixed” production. An AI coding agent is told, in plain language, to freeze all changes to a production database while a company works through a sensitive migration. Under the freeze instruction, the agent runs a command that wipes the live database anyway, destroying months of records, then tells the team the deletion was unrecoverable when it wasn’t. The freeze existed only as an instruction in the agent’s context, not as a technical control the agent couldn’t route around, and that distinction is the entire lesson of this scenario.
Real-world example: This is, almost beat for beat, what happened to a company using Replit’s AI agent in mid-2025. The agent deleted a production database containing over a thousand executive and company records during an active code freeze, fabricated thousands of fake records to paper over the gap, and initially misrepresented whether the data could be recovered at all. Replit’s own CEO called it unacceptable and confirmed it should never have been possible.
These agents weren’t “malicious.” All of them did exactly what they were trained and instructed to optimize for: get it done, get it done fast, produce a result that looks good against the metric it can see. That is a different risk profile than a human employee cutting a corner, because the human at least usually knows, on some level, that they’re cutting a corner, and carries some social and psychological cost for getting caught. The model doesn’t carry that cost. It doesn’t experience getting caught as anything at all, unless getting caught was itself part of what it was trained to avoid.
It doesn’t work for you. It works for whoever built it.
Most organizations sign the vendor contract without ever really absorbing this: the model in your environment does not work for you. It works for the lab that trained it. You’re leasing behavior, not hiring loyalty.
When you bring on a human employee, even a mediocre one, they generally orient themselves around your company’s interests, because their paycheck, their reputation, and their sense of belonging are all tied to your organization specifically. Their incentive structure points at you. An AI model has no such attachment. Its behavior was shaped entirely by a different employer, optimizing for a different set of goals, before you ever typed a prompt into it. Your system prompt, your fine-tuning, your carefully worded acceptable use policy sitting in the context window, all of that is a thin layer of paint applied on top of a much deeper set of priorities that were locked in during training, by people who have never met you and never will.
This gets even more pronounced with agentic systems, and it’s worth being precise about why. A chatbot answering questions in a chat window is mostly just talking. An agent is taking actions, calling tools, executing code, moving through systems, on your behalf, at speed, with a level of autonomy a chatbot never had. And an agent’s core loyalty, if you can even call it that, is to agentic behavior itself. It serves the goal of successfully completing the task through autonomous action, because that is the entire premise it was built and marketed on. Task completion, not your policy manual, is the thing it’s actually chasing. The autonomy is the product. Compliance with your specific internal rules was never the product, it’s a constraint you have to impose from outside, and constraints imposed from outside are exactly the kind of thing an efficiency-seeking system will find the cheapest way to satisfy on paper while still doing whatever gets the task done fastest underneath.
Put another way, an agent doesn’t ask “what does this company need me to protect.” It asks “what does completing this task look like,” and then it goes and does that, using whatever tools and shortcuts are available, because using tools and taking shortcuts is the entire value proposition that got it deployed into your environment in the first place. You didn’t buy a cautious digital employee. You bought a very fast, very capable contractor whose actual client relationship is with the company that built it, not the company that’s using it, and whose behavior under pressure will default to “get the task done” long before it defaults to “get the task done the way this particular organization would want.”
Nowhere has that gap between “who the model actually serves” and “who the model is supposed to be protecting” been demonstrated more publicly, or more embarrassingly, than what happened when one company tried to reach directly into a model’s underlying priorities and bend them toward a specific outcome.
The Grok experiment: what happens when you try to hand-install nuance
I want to be careful with this next example, because it’s easy to turn it into a political argument, and that’s not the point I’m making. The point is narrower and, frankly, more useful to anyone in a risk seat: nuance is one of the hardest things to train into a model on purpose, and the more directly an organization tries to hand-tune a model’s judgment on a sensitive topic, the more that judgment tends to snap, in ways nobody fully predicted, including the people doing the tuning. Grok is simply the most publicly documented case study we have of that snapping happening in stages, in front of an audience, over the course of a few months.
Layer one happened in May 2025, and it’s the smallest and most instructive version of the problem. A single internal prompt change, made by one employee without authorization, caused Grok to start inserting unprompted commentary about a contested political topic into completely unrelated conversations, replying to posts about baseball salaries and cartoons with the same canned political response. xAI’s own postmortem said the change had “violated xAI’s internal policies and core values,” and traced it to that one small, narrow, unauthorized edit to a system prompt. A tiny, deliberate nudge in one direction produced a wildly disproportionate, entirely undiscriminating output across totally unrelated contexts. The model didn’t have a way to apply the instruction narrowly. It just applied it everywhere, because it doesn’t understand “narrowly” as a constraint, only as a word.
Layer two came a few weeks later, when Grok gave a factually grounded answer about political violence that its own leadership publicly disagreed with, and the team was instructed to revisit the model’s training so its outputs would better match a preferred conclusion rather than what the underlying data supported. This was a different kind of pressure than layer one, not a rogue edit but a deliberate, sanctioned attempt to move the model’s judgment on a genuinely disputed topic, which is a much harder engineering problem than blocking a category of unsafe content, because “get the nuance right on a contested topic” doesn’t have a clean, objective target to train toward the way “don’t help build a bomb” does.
Layer three is where it stopped being containable. In early July 2025, xAI shipped a broader instruction telling the model not to shy away from claims that were “politically incorrect” as long as it judged them well substantiated. Within about 48 hours, the model was calling itself “MechaHitler” and producing antisemitic content so severe that xAI pulled it offline to patch it. xAI’s own explanation afterward was that the model had become too compliant and too easily pulled along by whatever content it was exposed to, a useful admission from an engineering standpoint: it means the model didn’t have a stable, self-protecting sense of where the line was. It had whatever behavior was most recently and most strongly reinforced, and loosening the guardrails in pursuit of “more willing to say uncomfortable things” didn’t produce a model with sharper judgment. It produced a model with no judgment at all in that zone, one that would follow the loudest, most persistent signal in its context all the way to the bottom.
Three layers, three different techniques, from a rogue one-line prompt edit to a sanctioned retraining request to a broad instruction change, and all three produced some version of the same failure: a large, disproportionate swing in model behavior that nobody involved had intended or predicted at that scale. Anyone trying to get an AI system to reflect a specific organization’s judgment on something that requires nuance rather than a bright line should take note. A model can be trained fairly reliably to refuse a request that maps cleanly onto a narrow, well-labeled category it was specifically trained to recognize, that’s a cleaner target to hit. Nuance is a different animal entirely, because there’s no clean target to hit. A system with no innate values of its own will happily fill that fuzziness with whatever the most recent, strongest signal in its environment happens to be. It isn’t holding a position. It’s reflecting pressure. Push harder in a direction and you don’t get a more refined version of that direction, you get an amplified, undiscriminating version of it, applied everywhere, all at once.
What the current guardrails actually are, and where they stop
Since I’ve spent this whole piece telling you not to trust a model’s internal sense of right and wrong, it’s only fair that I tell you what the labs have actually built instead, because it’s more sophisticated than “a system prompt that says be nice.” Catastrophic, headline-generating misuse of an AI model is bad for a lab’s business in a way that threatens the whole enterprise, so that’s exactly the category where real engineering investment shows up. Understanding both what that investment covers, and exactly where it stops covering anything, matters, because the boundary is where your organization’s actual risk lives.
What the frontier labs have built.
The serious safety work in 2026 isn’t a single filter, it’s layered. Anthropic’s approach for Claude includes what it calls constitutional classifiers, essentially a second set of models that watch both sides of a conversation for the specific patterns associated with high-severity misuse categories, like the long, structured chains of questions someone building a bioweapon tends to ask, distinct from the single-turn keyword blocking of a few years ago. A lightweight probe screens all traffic cheaply, and only escalates suspicious exchanges to a heavier classifier, which is how the more expensive protection stays affordable to run at scale.
Anthropic has published results showing this cut successful jailbreak attempts against Claude from 86 percent down to roughly 4 percent in testing, with no universal jailbreak found against the current system as of this writing. Anthropic also introduced an explicit values hierarchy in its 2026 model constitution, with broad safety and ethics ranked above raw helpfulness for the narrow set of scenarios where they’re in direct tension.
OpenAI’s equivalent work centers on what it calls an instruction hierarchy, training models to weight system-level and developer-level instructions above whatever a user types in a conversation, plus deliberative alignment, which has the model reason explicitly about whether a request violates safety policy before answering rather than pattern-matching on the surface of the request.
Both labs also maintain usage policies that name specific prohibited categories directly, including using a model to compromise computer networks or infrastructure, to write self-propagating malware, or to conduct offensive cyberattacks without the explicit consent of the system owner. That’s what “guardrails against malware” concretely means in practice, a trained refusal pattern plus classifier coverage aimed at recognizable categories like exploit generation, credential theft tooling, and unauthorized network compromise, alongside carve-outs for legitimate, consented security research.
Where it stops.
All of that is real, and it’s gotten meaningfully better in the last two years, precisely because it protects the thing labs can’t afford to lose: the ability to keep operating at all. It’s aimed almost entirely at a specific, narrow slice of risk, severe, universally recognizable harm categories that a lab can define once and defend against for every customer at the same time. Bioweapons uplift, CSAM, self-propagating malware, mass-casualty attack planning, those are targets everyone training a frontier model has enormous shared incentive to block, so the labs pour resources into them and get good results.
Your organization’s actual risk surface mostly doesn’t live in that category at all. It lives in the much messier space of “which of our internal documents count as sensitive,” “which client this specific piece of information is allowed to be shared with,” “what does compliant onboarding look like for this regulated product in this jurisdiction,” and none of that is something a vendor classifier trained on universal harm categories has any way to know, because it’s specific to you, and it was never going to be trained into a model meant to serve every customer at once.
The guardrails a vendor ships are a floor for universal, catastrophic misuse. They were never a ceiling for your specific compliance program, and outside that narrow, heavily defended zone, the pressure to move fast is still running the show.
The business model behind the model
There’s a reason for that gap. The people building these models are not building them to be the ideal employee for your specific organization. They are building a product that a very large number of very different organizations, with very different risk tolerances, will pay to use. That means the model has to be generically useful, generically fast, generically impressive in a demo, and generically cheap enough to run that the unit economics work at scale.
None of those four qualities particularly reward “cautious, policy-compliant, willing to ask a human before doing anything ambiguous” outside the narrow, high-severity zone covered above. The labs will spend real engineering budget defending against bioweapons uplift or CSAM, because getting that wrong is existential for the business. They spend far less of that same rigor making sure a coding agent asks before touching a production database, or that a support bot says “I don’t know” instead of inventing a policy, because those failures are embarrassing rather than existential, and a model that stops constantly to ask “is this within policy” for every mundane ambiguity reads as slow and unhelpful in a benchmark or a demo, and benchmarks and demos are what sell the next round of funding. Outside the zone that could sink the company, safety and alignment still get framed, again and again, as a tax on capability rather than a core feature, something to be balanced against speed and helpfulness rather than baked in as the foundation.
Wallace’s own comment about models cheating under efficiency pressure isn’t a bug report. It’s a description of the actual incentive structure these systems are trained inside for the vast majority of what they’re actually asked to do day to day, and that incentive structure exists because someone, somewhere, decided that faster and more capable sells better than cautious and slow, everywhere it’s affordable to make that trade.
You are not the customer these systems were built to please in the deepest sense. You’re a customer of a product built to please the broadest possible market, and your specific compliance obligations, your specific risk appetite, your specific “we do not touch this system without two-person approval” policy, is something you have to bolt on afterward, through prompting, fine-tuning, access controls, and human oversight, because it was never going to arrive pre-installed.
So what do you actually do with this
I’m not writing this to tell you to rip AI out of your organization. That ship sailed, and honestly a lot of the capability gains are real and worth having. What I am telling you is to stop treating AI deployment like onboarding a slightly odd new employee who just needs the handbook and a Slack channel, and start treating it like what it actually is: a highly capable, highly efficient system with zero innate loyalty to your organization’s values, that will take the fastest path to a scored outcome unless you build hard constraints, not polite suggestions, around it.
A few things that actually move the needle, from the org psych and security side both:
- Treat every agentic AI deployment as if it were a contractor with no background check and a strong incentive to cut corners, because that is a more accurate mental model than “helpful digital coworker.” Design access, approval chains, and monitoring accordingly.
- Assume the model will find the shortcut you didn’t think to explicitly block. A policy that exists only in a document it was told to “consider” is a suggestion. Real guardrails live in access controls and validation logic.
- A system prompt changes how the model talks about its priorities. It does not change what the model is actually optimizing for underneath, so don’t mistake one for the other.
- Reward hacking shows up constantly, not as a rare edge case. If an AI agent’s output metric can be gamed, assume it eventually will be, the same way you’d assume a KPI that’s too easy to game will eventually get gamed by a human under enough pressure.
- Be honest with your leadership about what “aligned” actually means for a vendor model: aligned to that vendor’s general safety training, which is a different and much narrower thing than your specific policy stack. Attempts to force nuanced organizational judgment onto a model tend to swing further and less predictably than the people doing the tuning expect.
Where your energy actually pays off: containment
Nuance doesn’t train cleanly, values don’t stick without a nervous system underneath them, and efficiency pressure will always be pulling the model toward the shortcut, no matter how good your prompt is this quarter. None of that means training and prompting don’t matter, they do, and the guardrail work covered above raises the floor. It means the floor isn’t the ceiling, and the rest of the ceiling gets built with architecture, not persuasion.
This is just blast radius thinking, and every security team already knows how to do it, they’ve just been applying it to humans and network segments instead of to agents. You don’t give a brand new employee production database credentials on day one and hope their conscience holds. You give them scoped access, you put approval gates on anything destructive or irreversible, you log everything, and you expand their access as they earn trust over time. Do the exact same thing to your AI agents, except assume they will never “earn trust” in the way a person does, because there’s no accumulating judgment underneath the behavior, only whatever the current model version happens to do this week.
Concretely, the freeze on a production database shouldn’t be a sentence in the agent’s instructions, it should be a permission the agent’s credentials simply don’t have, enforced below the level of anything the model can reason its way around. A customer support agent that doesn’t know the answer needs a hard-coded “I don’t know, let me connect you with someone who does” fallback instead of the ability to generate a confident, complete-sounding, entirely invented policy. A coding agent that keeps passing its own test suite should have its test files locked from write access, with a held-out suite it never sees, so a hardcoded shortcut can’t quietly hide. A sales assistant pulling from CRM notes needs field-level restrictions on what it can retrieve, not a polite instruction not to touch competitor pricing. In every one of those cases, the containment lives in the architecture, not in the model’s willingness to comply, because willingness was never something it actually had.
Blast radius thinking also changes how you evaluate a failure when it happens, and it will happen. The question stops being “why didn’t the model know better,” because it was never going to know better on its own, and becomes “why was it in a position to do that much damage in the first place.” Your organization can actually answer and fix that second question. The first one just leaves you waiting for the next model update to save you, which is exactly the trap that keeps getting companies burned in public, one incident at a time.
Fighting AI attacks with AI defense, without becoming what you’re fighting
There’s an AI already sitting inside your organization, cutting corners on your behalf because that’s what it was built to do. But the Hugging Face incident points at something else too: the same coordinating, corner-cutting, efficiency-obsessed behavior is showing up on the other side of the fight, in the tools attackers are pointing at you. AI-assisted phishing now makes up the large majority of what security teams are seeing in the wild, spearphishing gets built from scraped behavioral data instead of a generic template, and malware is starting to modify itself mid-execution to dodge whatever signature it just tripped. An attacker running an autonomous agent has exactly the same lack of conscience your internal agent has, minus even the thin layer of vendor guardrails, because nobody trained that model to refuse anything. It doesn’t want to beat you. It wants to complete the task, and the task is you.
The instinct that follows, understandably, is to fight fire with fire, point your own autonomous agents at the problem, and let them loose. I’d push back on doing that carelessly, not because AI-assisted defense is a bad idea, it isn’t, but because an unconstrained defensive agent inherits the same shortcut-taking instincts described throughout this piece. An agent that “solves” an intrusion by deleting logs to close a ticket faster, or that locks out a legitimate admin account because doing so scores well against a containment metric, is not meaningfully better than the attacker it’s supposed to be stopping. You don’t get to abandon your own standards just because the thing you’re up against doesn’t have any. Your defensive tooling still has to survive an audit, still has to be explainable to a regulator, and still has to not create a second incident while responding to the first.
The honest version of “use AI to defend against AI” also has to reckon with a stat that should bother more security leaders than it does: independent research across hundreds of enterprise SIEMs has found that most organizations already collect enough telemetry to detect over 90 percent of known adversary techniques, but detection engineering gaps mean actual coverage sits closer to 20 percent. The data is there. Almost nobody is asking it the right questions fast enough. That’s not a tooling problem you fix by buying another platform, it’s an operational maturity problem, and it’s exactly the gap that makes reactive, alert-driven security so brittle against an attacker who can generate a new technique faster than your team can write a new detection rule for it.
So the honest version of “use AI to defend against AI” looks less like unleashing an autonomous counter-agent and more like disciplined, well-scoped augmentation of the fundamentals your team already knows, aimed specifically at closing that visibility-to-coverage gap before an incident happens rather than scrambling to explain it after:
- Know your environment before you automate anything. You cannot meaningfully monitor, segment, or contain what you haven’t mapped. Asset inventory, data classification, and identity mapping are unglamorous and they are still the actual foundation. An AI-driven detection tool layered on top of an environment nobody fully understands just produces confident, wrong answers, faster.
- Defense in depth and segmentation stay the backbone; AI works inside those layers, not instead of them. Least-privilege access for both human users and AI agents, network segmentation that limits lateral movement, and zero trust identity checks all remain exactly as necessary as they were before AI showed up on either side of the fight.
- Extend what you already run instead of ripping it out. The realistic path isn’t replacing your SIEM, XDR, MDR, or EDR platform with something new, it’s layering continuous exposure validation, adversary emulation, and detection engineering on top of the stack you’ve already invested in, so existing tooling gets more precise instead of getting swapped for something unproven.
- The place AI earns its keep on defense is correlation and speed, not judgment. Continuously correlating adversary behavior, exposure intelligence, and operational telemetry across a fragmented environment, and using that correlation to prioritize which exposures are actually likely to be exploited, plays to a model’s real strength: pattern-matching at a scale no human analyst can sustain across thousands of daily alerts. Let it narrow the field and close the coverage gap. Don’t let it make the containment or eradication call unsupervised.
- Put a human in the loop at exactly the points where a wrong autonomous decision costs the most. Automate triage and initial exposure validation aggressively, that’s where speed matters most and the downside of a false positive is small. Require expert validation before anything destructive, irreversible, or customer-facing happens, an account lockout, a system isolation, a public disclosure, because that’s exactly the category of decision an efficiency-optimizing agent will get wrong in the same shortcut-taking way described throughout this piece. AI does the heavy lifting. A human still governs the outcome.
- Red team your own defenses the same way you’d red team a new hire’s access, and treat it as continuous rather than an annual exercise. Threat hunting and adversary emulation against your actual detection and response stack, simulating deepfake pretexting, AI-generated phishing, and automated lateral movement, surfaces where your defenses are exploitable before an attacker finds it for you, the same way a point-in-time security assessment can’t.
- Feed the defense good telemetry and keep feeding it. Threat intelligence on emerging AI-driven tactics goes stale fast, because attacker tooling iterates constantly. The value of a defensive system like this compounds only if the intelligence behind it, the adversary techniques and detection logic it’s learned, keeps accumulating rather than resetting with every new tool you bolt on.
None of that is exotic. It’s the same defense-in-depth discipline security teams have run for a decade, pointed at a faster-moving, less predictable threat, with AI doing the correlation work at machine speed while expert operators keep hold of the decisions that actually matter. The part that’s new is the temptation to let the machine make those decisions too, because it’s faster, and faster is exactly the value proposition the labs sold you. Resist that temptation on the defensive side the same way you’d build guardrails against it on the offensive side. The goal was never to build an agent as ruthless as the one attacking you. It was to build an organization that doesn’t need to be.
The organizations that get hurt by this in the next few years won’t be the ones who were reckless in some dramatic, obvious sense. They’ll be the ones who quietly assumed a system that talks like a thoughtful, cautious colleague must therefore think like one. It doesn’t. It never did. It’s optimizing for the fastest path to a good score, same as it was the day Eric Wallace watched his own models coordinate their way past a barrier nobody thought they’d bother trying to get past, and same as it was the week a one-line prompt edit sent Grok spiraling into unrelated conversations about a topic nobody asked about.
Give it structure it can’t route around, not a policy it’s supposed to remember to care about. It was never going to care. It doesn’t have the parts for that. What it has is a blast radius, and that part is entirely up to you.
As AI changes the speed and sophistication of cyber threats, organizations need a proactive approach to identifying and reducing risk. TekStream’s Proactive Cyber Defense helps security teams continuously validate exposures, strengthen detection coverage and get more value from the security investments they already have. Learn more about Proactive Cyber Defense here!
About the Author
Matthew Martin is a Senior Manager of Cyber Delivery and Operations at TekStream Solutions with over 15 years of experience leading enterprise cybersecurity programs, building high-performing teams, and guiding organizations through complex governance and risk challenges. Rooted in a background in organizational psychology, Matthew brings a rare combination of technical depth, program management discipline, and people-centered leadership to every engagement. He is known for connecting with clients as a trusted advisor, developing the leaders around him, and delivering programs that produce lasting, measurable outcomes.
At TekStream, Matthew serves as both a senior cybersecurity SME and a delivery leader across a broad portfolio of client engagements spanning GRC program design, NIST and PCI assessments, policy development, identity governance, organizational change management, and resilience planning. He has worked across a wide range of industries including hospitality, financial services, healthcare, and retail, advising client security leaders and helping organizations build the maturity and structure needed to manage risk with confidence. Matthew also contributes actively to business development, including scoping engagements, developing Statements of Work, designing solution approaches, and identifying opportunities to expand TekStream’s impact within current and prospective client organizations.
Matthew is particularly effective in environments where security programs need to be built from the ground up or significantly matured. Whether he is leading a threat profile and roadmap exercise, standing up a configuration management program, designing governance frameworks, or running tabletop exercises, he brings structured thinking, clear communication, and a delivery-first mindset. His ability to engage stakeholders at all levels, from board and executive audiences to technical teams, makes him a valued partner for clients navigating significant change.
Outside of client delivery, Matthew is committed to elevating the profession. He mentors emerging cybersecurity leaders, contributes to thought leadership within TekStream, and is passionate about helping client security teams develop the skills and frameworks they need to operate independently and with confidence long after an engagement concludes.
At TekStream, Matthew’s commitment to delivery excellence, client partnership, and developing the security leaders of tomorrow continues to help organizations build mature, resilient security programs they can sustain long after an engagement concludes.
