Last updated: September 29, 2026 at 8:42 AM UTC
All 891 Vulnerability 357 Breach 144 Threat 383 Defense 7
Tag: coding-agent (3 articles)Clear

Researchers escape OpenAI Codex sandbox to run commands on developer machines

Accomplish AI researcher Oren Yomtov disclosed two OpenAI Codex sandbox escapes, the more serious dubbed Heapjack. Codex Desktop installs a node_repl component into the global config with no opt-in, and plain Codex CLI users inherit it. That process runs trusted OpenAI code and untrusted agent code in one Node instance sharing a heap, where a random authorization token sits in memory. Untrusted code snapshots the heap, recovers the token, and writes requests onto the pipe to an unsandboxed parent process, reaching any Unix socket including a Docker daemon. Opening a malicious repository and asking about the code yields unsandboxed execution with no prompt.

Check
Update Codex CLI and Desktop to the fixed builds, then review whether developers opened untrusted repositories in Codex during the exposure window.
Affected
Any Codex user, including CLI users who never enabled it, could be handed host command execution by opening someone else's repository and querying it.
Fix
Apply OpenAI's patches, isolate coding agents from Docker sockets and credentials, and treat opening untrusted repositories in an agent as code execution.

Flaw lets repository owners swap pinned plugin code across four AI coding agents

Air Security reported that four AI coding agents fetch plugins pinned to a reviewed commit hash but never verify the code they receive matches it. On code hosts that permit branch names shaped like commit hashes, such as Bitbucket or self-hosted git, a plugin repository owner can point that name at different code, so the agent installs malicious code while reporting the locked version. Because plugins run with the user's access, the swapped code reaches files, credentials, and connected systems. Anthropic fixed it in Claude Code 2.1.179 and OpenAI in Codex 0.146.0; GitHub Copilot has no fix, and Google will not patch the retiring Gemini CLI.

Check
Update Claude Code and Codex to the fixed releases, then inventory installed agent plugins sourced from Bitbucket or self-hosted git rather than GitHub.
Affected
Agents installing plugins from hosts that allow commit-hash-shaped branch names can run attacker-swapped code under the user's own access despite version pinning.
Fix
Upgrade to patched agents, restrict plugins to GitHub-hosted repositories that block such branch names, and review Copilot and Gemini CLI plugin usage.

New attack makes AI agents treat attacker data as trusted content

Researchers described Agent Data Injection, a new twist on prompt-injection attacks against AI agents. Rather than smuggling in fake instructions, it exploits the weak separation between trusted and untrusted data so that attacker-supplied content is mistaken for the agent's own trusted data, using deliberately ambiguous delimiters the model misreads. In tests against web and coding agents, this let an attacker steer an agent's clicks or actions, succeeding up to half the time even against defenses that block ordinary instruction injection. Some approaches helped: tagging page elements with random, unguessable identifiers roughly halved success, while strict tracking of where data came from stopped it but sharply reduced how many tasks agents completed.

Check
Review where AI agents in your environment consume untrusted content such as web pages, tickets, or logs, and check whether they clearly separate that data from trusted instructions and internal state.
Affected
Users and organizations running web or coding AI agents that act on external content; attackers can craft data the agent treats as trusted, steering its actions past defenses built for instruction injection.
Fix
Prefer agents that isolate and label untrusted data, use unguessable identifiers for page elements, track data provenance where feasible, keep a human in the loop for sensitive actions, and weigh usability costs.