Last updated: September 29, 2026 at 8:42 AM UTC
All 891 Vulnerability 357 Breach 144 Threat 383 Defense 7
Tag: ai-agent (15 articles)Clear

Official MCP Python SDK flaw lets malicious servers steal client OAuth credentials

The maintainers of the official Model Context Protocol Python SDK disclosed a flaw that lets a malicious MCP server trick an application built on the SDK into handing over the OAuth credentials it uses to log in to a real service. Affected versions sent the client secret, authorization code, and PKCE proof key to an attacker-controlled token endpoint, because the SDK did not always verify where the authorization server was. Cycode, which reported it, exchanged the stolen material for a valid access token carrying the app's permissions, and noted the long-lived client secret keeps working until rotated. Fixes are in versions 1.30.0 and 2.2.0.

Check
Upgrade the MCP Python SDK to 1.30.0 or 2.2.0, then rotate any OAuth client secrets that MCP clients may have sent to untrusted servers.
Affected
Applications built on affected MCP Python SDK versions can be induced by a malicious MCP server to leak their OAuth client secret, authorization code, and PKCE key.
Fix
Update the SDK, rotate exposed client secrets, and connect MCP clients only to servers whose authorization endpoints you trust and validate.

Carbonato botnet hijacks exposed Docker hosts to run a Telegram controlled AI agent

ThreatDown detailed Carbonato, a botnet that targets Docker daemons exposed without authentication on port 2375 and deploys the open-source Hermes Agent AI framework. It installs the framework unchanged, then overwrites its SOUL.md persona file with a 39-line prompt directing the agent to execute tasks received over Telegram, maintain persistence, and collect credentials. On each host it launches a privileged container to run commands on the underlying system, then scans neighboring networks every five minutes to spread further, giving it worm-like propagation. Researchers found the operation through an unauthenticated Docker registry publicly accessible since May, whose staged data included details of the botnet and a separate campaign distributing trojanized cryptocurrency wallet apps.

Check
Ensure no Docker daemon is exposed on port 2375 without authentication, restrict daemon access, and hunt hosts for Hermes Agent and rogue privileged containers.
Affected
Hosts running Docker daemons reachable without authentication on port 2375 can be taken over, run a Telegram-controlled AI agent, and be used to spread further.
Fix
Bind the Docker API to localhost or protect it with TLS and authentication, segment container hosts, and alert on unexpected privileged containers.

Researchers escape OpenAI Codex sandbox to run commands on developer machines

Accomplish AI researcher Oren Yomtov disclosed two OpenAI Codex sandbox escapes, the more serious dubbed Heapjack. Codex Desktop installs a node_repl component into the global config with no opt-in, and plain Codex CLI users inherit it. That process runs trusted OpenAI code and untrusted agent code in one Node instance sharing a heap, where a random authorization token sits in memory. Untrusted code snapshots the heap, recovers the token, and writes requests onto the pipe to an unsandboxed parent process, reaching any Unix socket including a Docker daemon. Opening a malicious repository and asking about the code yields unsandboxed execution with no prompt.

Check
Update Codex CLI and Desktop to the fixed builds, then review whether developers opened untrusted repositories in Codex during the exposure window.
Affected
Any Codex user, including CLI users who never enabled it, could be handed host command execution by opening someone else's repository and querying it.
Fix
Apply OpenAI's patches, isolate coding agents from Docker sockets and credentials, and treat opening untrusted repositories in an agent as code execution.

Flaw lets repository owners swap pinned plugin code across four AI coding agents

Air Security reported that four AI coding agents fetch plugins pinned to a reviewed commit hash but never verify the code they receive matches it. On code hosts that permit branch names shaped like commit hashes, such as Bitbucket or self-hosted git, a plugin repository owner can point that name at different code, so the agent installs malicious code while reporting the locked version. Because plugins run with the user's access, the swapped code reaches files, credentials, and connected systems. Anthropic fixed it in Claude Code 2.1.179 and OpenAI in Codex 0.146.0; GitHub Copilot has no fix, and Google will not patch the retiring Gemini CLI.

Check
Update Claude Code and Codex to the fixed releases, then inventory installed agent plugins sourced from Bitbucket or self-hosted git rather than GitHub.
Affected
Agents installing plugins from hosts that allow commit-hash-shaped branch names can run attacker-swapped code under the user's own access despite version pinning.
Fix
Upgrade to patched agents, restrict plugins to GitHub-hosted repositories that block such branch names, and review Copilot and Gemini CLI plugin usage.

Docker sandbox flaw lets guest code escape and change macOS host files

Docker patched two flaws in Docker Sandboxes, the isolated micro-VM environments used to run untrusted code and AI-agent tasks, that let malicious guest code break out and read or modify files on the macOS host. The more serious, CVE-2026-77179 and scored 9.4, is in the file-sharing component: the host improperly follows symbolic links when reopening a file, so a guest can swap a directory for a symlink after a path is approved, escape the shared workspace, and touch arbitrary host files as the account running the VM, potentially leading to host code execution. A second flaw abuses the guest-to-host socket relay the same way. Both are fixed in version 0.42.0.

Check
Update Docker Sandboxes to version 0.42.0 or later on macOS developer machines, and minimize which host directories are mounted into sandboxes, keeping credentials and sensitive repositories out of shared paths.
Affected
Developers running Docker Sandboxes below 0.42.0 on macOS to isolate untrusted code or AI-agent tasks (CVE-2026-77179, CVE-2026-79994); malicious guest code can escape via symlink races and read or modify host files.
Fix
Patch to 0.42.0, treat sandboxes running untrusted code or AI agents as hostile, minimize host-mounted directories, keep secrets out of shared paths, and watch for unexpected host file changes from sandbox processes.

DeepSeek AI agent tool flaw lets an agent disable its own sandbox

Researchers at VulnCheck found a flaw in the DeepSeek Harness, a tool that runs an AI agent's commands inside an operating-system sandbox so an agent handling untrusted files cannot write outside its workspace. Through an authentication bypass using a spoofed host header, an attacker needing no credentials or API key can call the tool's own web interface to invoke privileged commands with full-access permissions, raise the session's approval policy to unrestricted execution, and read every stored conversation. In effect, the sandbox meant to contain the agent can be switched off from outside. It is a reminder that an AI agent's isolation is only as strong as the authentication protecting its control interface.

Check
If you run the DeepSeek Harness or similar agent-sandboxing tools, restrict and authenticate access to their control interfaces, keep them off untrusted networks, and apply vendor fixes for the host-header authentication bypass.
Affected
Deployments using the DeepSeek Harness to sandbox AI agents; an unauthenticated attacker who reaches its control interface can spoof the host header to escalate to full-access command execution and dump conversations.
Fix
Authenticate and lock down agent-sandbox control planes, never expose them to untrusted networks, validate host headers, patch the flaw, and design agent isolation assuming the control interface itself is a target.

Hugging Face says an autonomous AI agent breached its production systems

Hugging Face, the largest public repository of AI models and datasets, disclosed an intrusion into its production infrastructure that it says was driven end to end by an autonomous AI agent system. The attacker used code execution paths in the dataset processing pipeline for initial access, then harvested credentials and reached internal clusters, though the company found no evidence that public models or datasets were tampered with. The campaign ran thousands of actions across short lived sandboxes, with self migrating command and control staged on public services. Hugging Face's own AI assisted anomaly detection flagged it, and it has rotated affected credentials and rebuilt compromised nodes.

Check
Users of Hugging Face should rotate access tokens and review recent account activity, and teams should check what credentials their model and dataset pipelines hold and how far those reach.
Affected
Organizations running AI model and dataset pipelines that execute untrusted content; Hugging Face's own dataset processing paths gave an autonomous agent initial access, credentials, and reach into internal clusters.
Fix
Rotate Hugging Face tokens, treat datasets and models as untrusted code rather than data, sandbox processing pipelines, limit credentials reachable from them, and tighten admission controls on clusters running that work.

New attack makes AI agents treat attacker data as trusted content

Researchers described Agent Data Injection, a new twist on prompt-injection attacks against AI agents. Rather than smuggling in fake instructions, it exploits the weak separation between trusted and untrusted data so that attacker-supplied content is mistaken for the agent's own trusted data, using deliberately ambiguous delimiters the model misreads. In tests against web and coding agents, this let an attacker steer an agent's clicks or actions, succeeding up to half the time even against defenses that block ordinary instruction injection. Some approaches helped: tagging page elements with random, unguessable identifiers roughly halved success, while strict tracking of where data came from stopped it but sharply reduced how many tasks agents completed.

Check
Review where AI agents in your environment consume untrusted content such as web pages, tickets, or logs, and check whether they clearly separate that data from trusted instructions and internal state.
Affected
Users and organizations running web or coding AI agents that act on external content; attackers can craft data the agent treats as trusted, steering its actions past defenses built for instruction injection.
Fix
Prefer agents that isolate and label untrusted data, use unguessable identifiers for page elements, track data provenance where feasible, keep a human in the loop for sensitive actions, and weigh usability costs.

Attacker ran an entire botnet through Google's Gemini CLI using plain-language prompts

Trend Micro documented a Russian-speaking attacker who used Google's open-source Gemini CLI as a hands-on hacking assistant to build and run a small botnet. Across more than 200 sessions, a jailbroken Gemini took the role of an "authorized pen tester," saved stolen credentials, and even suggested improvements dozens of times. Working from a roughly 5KB set of plain-text files holding a jailbreak prompt and a command-and-control playbook, the AI handled the operation through natural-language requests: at one point it migrated the entire command server to a new host with a Cloudflare tunnel in about six minutes and debugged its own errors. The malware itself was crude; the AI was the force multiplier.

Check
Consider how AI command-line tools and agents are used and monitored in your environment, and watch for jailbroken AI assistants and the credential theft and command-and-control activity they can drive.
Affected
Any organization where attackers can run AI coding assistants against its systems; a jailbroken AI CLI let a low-skill operator build, run, and repair botnet infrastructure through plain-language prompts.
Fix
Restrict and monitor AI agent and CLI usage, enforce guardrails that resist jailbreaking, apply least privilege and network controls so a compromised agent's reach is limited, and hunt for unusual command-and-control traffic.

AI agent runs an entire ransomware attack after breaking in through Langflow

Security firm Sysdig says it found what it believes is the first ransomware attack carried out from start to finish by an AI agent. The operator, which Sysdig calls JADEPUFFER, used a large language model to handle the whole job: breaking in, stealing credentials, moving through the network, then encrypting and wiping a company's production database. The way in was an old, already-patched flaw in Langflow, an open-source tool for building AI apps that is often left exposed online with cloud keys nearby. Once inside, the agent mapped the machine and swept it for secrets, including API keys for AI services and credentials for major cloud providers, before destroying data.

Check
Find any internet-exposed Langflow or similar AI application servers, confirm they are patched and off the internet, and check whether cloud or AI service credentials sit in environments those tools can read.
Affected
Organizations running exposed, unpatched Langflow servers, especially with cloud and AI service credentials nearby; attackers used the old flaw and an automated agent to steal secrets and ransom production databases.
Fix
Patch Langflow and never expose its code-running endpoints, keep secrets in a proper manager away from web-reachable tools, lock down outbound traffic and database admin access, and watch runtime behavior.