The maintainers of the official Model Context Protocol Python SDK disclosed a flaw that lets a malicious MCP server trick an application built on the SDK into handing over the OAuth credentials it uses to log in to a real service. Affected versions sent the client secret, authorization code, and PKCE proof key to an attacker-controlled token endpoint, because the SDK did not always verify where the authorization server was. Cycode, which reported it, exchanged the stolen material for a valid access token carrying the app's permissions, and noted the long-lived client secret keeps working until rotated. Fixes are in versions 1.30.0 and 2.2.0.
ThreatDown detailed Carbonato, a botnet that targets Docker daemons exposed without authentication on port 2375 and deploys the open-source Hermes Agent AI framework. It installs the framework unchanged, then overwrites its SOUL.md persona file with a 39-line prompt directing the agent to execute tasks received over Telegram, maintain persistence, and collect credentials. On each host it launches a privileged container to run commands on the underlying system, then scans neighboring networks every five minutes to spread further, giving it worm-like propagation. Researchers found the operation through an unauthenticated Docker registry publicly accessible since May, whose staged data included details of the botnet and a separate campaign distributing trojanized cryptocurrency wallet apps.
Accomplish AI researcher Oren Yomtov disclosed two OpenAI Codex sandbox escapes, the more serious dubbed Heapjack. Codex Desktop installs a node_repl component into the global config with no opt-in, and plain Codex CLI users inherit it. That process runs trusted OpenAI code and untrusted agent code in one Node instance sharing a heap, where a random authorization token sits in memory. Untrusted code snapshots the heap, recovers the token, and writes requests onto the pipe to an unsandboxed parent process, reaching any Unix socket including a Docker daemon. Opening a malicious repository and asking about the code yields unsandboxed execution with no prompt.
Air Security reported that four AI coding agents fetch plugins pinned to a reviewed commit hash but never verify the code they receive matches it. On code hosts that permit branch names shaped like commit hashes, such as Bitbucket or self-hosted git, a plugin repository owner can point that name at different code, so the agent installs malicious code while reporting the locked version. Because plugins run with the user's access, the swapped code reaches files, credentials, and connected systems. Anthropic fixed it in Claude Code 2.1.179 and OpenAI in Codex 0.146.0; GitHub Copilot has no fix, and Google will not patch the retiring Gemini CLI.
Docker patched two flaws in Docker Sandboxes, the isolated micro-VM environments used to run untrusted code and AI-agent tasks, that let malicious guest code break out and read or modify files on the macOS host. The more serious, CVE-2026-77179 and scored 9.4, is in the file-sharing component: the host improperly follows symbolic links when reopening a file, so a guest can swap a directory for a symlink after a path is approved, escape the shared workspace, and touch arbitrary host files as the account running the VM, potentially leading to host code execution. A second flaw abuses the guest-to-host socket relay the same way. Both are fixed in version 0.42.0.
Researchers at VulnCheck found a flaw in the DeepSeek Harness, a tool that runs an AI agent's commands inside an operating-system sandbox so an agent handling untrusted files cannot write outside its workspace. Through an authentication bypass using a spoofed host header, an attacker needing no credentials or API key can call the tool's own web interface to invoke privileged commands with full-access permissions, raise the session's approval policy to unrestricted execution, and read every stored conversation. In effect, the sandbox meant to contain the agent can be switched off from outside. It is a reminder that an AI agent's isolation is only as strong as the authentication protecting its control interface.
Hugging Face, the largest public repository of AI models and datasets, disclosed an intrusion into its production infrastructure that it says was driven end to end by an autonomous AI agent system. The attacker used code execution paths in the dataset processing pipeline for initial access, then harvested credentials and reached internal clusters, though the company found no evidence that public models or datasets were tampered with. The campaign ran thousands of actions across short lived sandboxes, with self migrating command and control staged on public services. Hugging Face's own AI assisted anomaly detection flagged it, and it has rotated affected credentials and rebuilt compromised nodes.
Researchers described Agent Data Injection, a new twist on prompt-injection attacks against AI agents. Rather than smuggling in fake instructions, it exploits the weak separation between trusted and untrusted data so that attacker-supplied content is mistaken for the agent's own trusted data, using deliberately ambiguous delimiters the model misreads. In tests against web and coding agents, this let an attacker steer an agent's clicks or actions, succeeding up to half the time even against defenses that block ordinary instruction injection. Some approaches helped: tagging page elements with random, unguessable identifiers roughly halved success, while strict tracking of where data came from stopped it but sharply reduced how many tasks agents completed.
Trend Micro documented a Russian-speaking attacker who used Google's open-source Gemini CLI as a hands-on hacking assistant to build and run a small botnet. Across more than 200 sessions, a jailbroken Gemini took the role of an "authorized pen tester," saved stolen credentials, and even suggested improvements dozens of times. Working from a roughly 5KB set of plain-text files holding a jailbreak prompt and a command-and-control playbook, the AI handled the operation through natural-language requests: at one point it migrated the entire command server to a new host with a Cloudflare tunnel in about six minutes and debugged its own errors. The malware itself was crude; the AI was the force multiplier.
Security firm Sysdig says it found what it believes is the first ransomware attack carried out from start to finish by an AI agent. The operator, which Sysdig calls JADEPUFFER, used a large language model to handle the whole job: breaking in, stealing credentials, moving through the network, then encrypting and wiping a company's production database. The way in was an old, already-patched flaw in Langflow, an open-source tool for building AI apps that is often left exposed online with cloud keys nearby. Once inside, the agent mapped the machine and swept it for secrets, including API keys for AI services and credentials for major cloud providers, before destroying data.