Malware carries a hidden prompt to derail AI-assisted analysis
ESET found that a Russia-aligned group planted a prompt inside a malicious script designed to trip an AI system's safety filters and disrupt AI-assisted malware analysis. The technique, dubbed GuardBreaker, embeds text about a sensitive topic so that a language model reviewing the code refuses or derails instead of analyzing it. The script itself installs a loader the group uses to deliver further payloads. It is not isolated: earlier in 2026, malicious packages in supply-chain campaigns used similar anti-analysis tricks against systems leaning on a language model for triage. The lesson is that automated AI triage can be manipulated by the code it inspects, so human review remains essential.
- Check
- If you use language models to triage code or malware, assume attackers will manipulate them, and keep human analysts and traditional sandboxing in the loop rather than trusting AI output alone.
- Affected
- Security workflows relying on language models for first-pass code or malware triage; attackers embed prompts in samples to trigger safety refusals or misdirection, causing the AI to skip or misjudge malicious code.
- Fix
- Treat AI triage output as manipulable, isolate the model from acting on embedded instructions, combine it with signature and behavioral analysis and human review, and test pipelines against prompt-injection samples.