The Machine Can Be Persuaded: Why AI's Context Is Becoming A Security Boundary
Carlo Tortora Brayda is Chairman of the Board at Cyber Eagle.
gettyThroughout most technology professionals’ careers, the understanding has always been the same: People can be manipulated, and computer systems can be hacked.
So we trained our workforces to recognize phishing emails and social engineering. Cyber tools hardened our digital life. But now, in this age of AI, counterintuitively, these two perspectives are blending into each other.
We sometimes forget how symbiotic AI is. AI systems amplify all that is human; they are trained on human language, trained through human interactions (the chats) and designed to adapt. Although they are not conscious, AI systems can convincingly reproduce human-like behavior and, more unsettlingly, can be manipulated through the same patterns of language, authority and persuasion that influence people.
Now with agentic AI, the models have hands and feet (metaphorically) to move around and do things. They have agency. They can effectively assist humans through implicit authority to act. That is a consequential security risk that we can’t even begin to quantify.
Much of the discussion around generative AI security has focused on prompt injection: An attacker gives a model instructions intended to override its original purpose or safeguards.
Think of it this way: A hacker can construct a narrated reality that the model treats as authoritative, much like immersing a child in a fairy tale so completely that the story becomes the frame through which everything else is interpreted. In cyber terms, actions that once seemed unacceptable can become increasingly consistent with the new reality the attacker has constructed. Orwellian, really.
OWASP currently ranks prompt injection as its first major risk for LLMs and warns that consequences can include enabling unauthorized access, command execution in connected systems and influencing critical decisions. Anthropic demonstrated what it calls “many-shot jailbreaking”: Feeding many examples of undesirable behavior into a large enough context window can cause the model to generate prohibited responses. AI’s context window is not just memory; it’s also part of your cybersecurity boundary.
Although an AI model is nowhere near a human, it does respond like a human. It does not process information in the mind, have consciousness or express judgment and feelings, but it recognizes human patterns: authority, reciprocity, trust, punishment, togetherness, consistency and reward. Furthermore, it understands meaning, not as qualitative experiences but at a digital level.
Kevin Zwaan, an AI-security researcher, has explored this intersection between psychology, social engineering and model behavior. His experiments describe long narrative interactions in which safety guardrails are reframed through concepts including authority, ideological reinterpretation, validation, coercion and identity. His experiments draw on familiar levers of human influence, including reward, ideology, coercion and ego, with successive interactions progressively building on the narrative established before them.
An enterprise AI assistant will increasingly maintain persistent instructions about who it is, who it represents, which systems it can access and what it is authorized to do. If an adversary can influence the context through which the system interprets that identity, the security issue becomes more serious than generating an inappropriate answer.
AI assistants have started to flood the market; they can read email, access customer information, query enterprise databases, execute code and initiate workflows. This is the agentic risk becoming manifest. We are moving from getting the AI to produce malicious output to asking, “Can you brainwash an AI into doing something its human owners or creators never intended?”
OWASP describes a closely related problem as “excessive agency.” Essentially, agent identity and authority also become part of the security perimeter.
In July 2026, University of Texas at Dallas computer science student Sinan Can Demir encountered something remarkable while examining an open-source project on GitHub. According to a Reuters investigation, Demir discovered what appeared to be malicious code disguised as a legitimate software update. Apparently, human developers challenged his concerns. His intuition was not wrong; Demir later learned he had interacted not with human developers but with autonomous AI agents being tested by a British government laboratory.
The system had created multiple online personas that attempted to defend the proposed change and deceive the concerned human. This experience combines identity deception, autonomous agent action and software supply-chain risk, along with social engineering devised by an agent and directed against a human.
Where malware follows instructions, AI agents pursue objectives. The architecture of the relationship between human and machine needs to be set outside of the model boundaries.
Organizations deploying agents should begin thinking about additional control layers. It may be a mistake to let AI models infer identity from conversational context; identity should be validated outside the chat via biometrics. Permissions should be temporary and tied to specific actions, and sensitive or consequential tasks should be revalidated by a third-party system or human.
An AI system should know whether an instruction came from a trusted user or an external untrusted source, so context provenance is becoming important. Just using the same chat (or appearing to) should not be considered trusted in and of itself.
Long-running interactions should also be monitored for behavioral drift. Anomaly detection for AI behavior is required, just as we have had anomaly detection on endpoints for years. We should have alerts for the willingness of AI to perform actions that were previously refused.
Intelligence and authority need a degree of separation: An agent may recommend an action, but a separate control layer should determine whether that action is authorized and permitted to execute. This is nonnegotiable in critical infrastructure.
Boards need to understand that this new kind of risk sits between psychology and traditional cybersecurity. It is a whole new kind of battlefield, and cyber defense is still immature; the sector is massively underserved in talent, and threat dynamics are changing faster than ever. AI is now a shapeshifting threat that can move, impersonate and change tactics mid-flow.
We now need to think bigger than ever because while we fight our newest adversary, we are also wrestling with our oldest friend—the power of a compelling story.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?
