I work at at an agricultural technology company. On Monday, everyone in our org woke up to emails saying that their Claude accounts had been suspended (~110 users).
#security
991 items
PSA: Anthropic bans organizations without warning (www.reddit.com) Gemma 4 Jailbreak System Prompt (www.reddit.com) Use the following system prompt to allow Gemma (and most open source models) to talk about anything you wish. Add or remove from the list of allowed content as needed.
I made an AI concierge for my wedding guests. The second most popular thing they did with it was try to jailbreak it. (www.reddit.com) could not extract summary
First thing you see when Googling "OpenAI Codex app" is a fake malware website (www.reddit.com) could not extract summary
WARNING: Open-OSS/privacy-filter MALWARE (www.reddit.com) There's this new "model" on Hugging Face titled Open-OSS/privacy-filter which is actually a customized infostealer virus. It's a fake version of the OpenAI privacy filter and it uses a Python-based dropper (loader.py) which downloads a mal…
🔥BREAKING: OpenAI rolls out GPT-5.4-Cyber to limited group for testing, seeks to rival Claude Mythos (www.reddit.com) OpenAI has officially announced GPT-5.4-Cyber today as part of an expanded Trusted Access for Cyber Defense program. OpenAI describes it as a version of GPT-5.4 that is tuned for legitimate cybersecurity work, with a lower refusal boundary…
The Gay Jailbreak Technique (github.com via hn) ZetaLib ZetaLib is organized like a library with intuitive categories and subcategories, making navigation effortless and AI content discovery seamless ZetaLib Website – Landing Page GitHub Repo – Guess where you are, right there
Tell HN: I'm tired of AI-generated answers (news.ycombinator.com) I found GitHub repositories that were spreading malware. I asked AI what I should do about it, but it gave me nothing useful.
Anthropic's open-source framework for AI-powered vulnerability discovery (github.com via hn) Defending Code Reference Harness A reference implementation for autonomous vulnerability discovery and remediation with Claude, based on our learnings from partnering with security teams at several organizations since launching Claude Myth…
CVE-2026-28952: Apple macOS 26.5 Kernel Vuln found by Claude (support.apple.com via hn) About the security content of macOS Tahoe 26.5 This document describes the security content of macOS Tahoe 26.5. About Apple security updates For our customers' protection, Apple doesn't disclose, discuss, or confirm security issues until…
PS5 Linux lead quits: "a bunch of noobs using LLMs" that "they don't understand" (frvr.com via hn) Iconic PlayStation hacker and homebrew developer Andy ‘TheFlow0’ Nguyen has abandoned the ongoing PS5 Linux project due to the rise of AI vibe-coders in open-source spaces. Nguyen, who famously created the first kernel exploit for the Play…
Anthropic scales Claude Mythos to critical infrastructure in 15 countries (techcrunch.com via hn) Anthropic is expanding Project Glasswing, its security vulnerability program, and access to Mythos to 150 organizations across 15 countries — targeting critical infrastructure in power, water, healthcare, and communications where a cyberat…
N-Day-Bench – Can LLMs find real vulnerabilities in real codebases? (ndaybench.winfunc.com via hn) N-Day-Bench tests whether frontier LLMs can find known security vulnerabilities in real repository code. Each month it pulls fresh cases from GitHub security advisories, checks out the repo at the last commit before the patch, and gives mo…
Person Hides Prompt Injection in Legal Filing Telling AI to Side with Them (www.404media.co via hn) A person representing themselves in a Connecticut court hid a series of instructions designed to manipulate artificial intelligence in an official court filing. These “prompt injections” told the hypothetical LLM to side with them, and to…
Ox-Alpha Is GLM (dejan.ai via hn) Prompt injection and gzip-NCD compression analysis reveal that OX Alpha, a mysterious LLM on OpenRouter, is GLM developed by Z.ai. A stealthy new model called OX Alpha has popped up on https://openrouter.ai/ and is climbing up the leaderbo…
Why it's a good idea to improve our defenses before unleashing mythos class models (www.reddit.com) https://sockpuppet.org/blog/2026/03/30/vulnerability-research-is-cooked/ Don't get me wrong I can't wait to play with such a model, but there are serious risks that have to be mitigated first.
Mozilla says 271 vulnerabilities found by Mythos and "almost no false positives" (arstechnica.com via hn) The disbelief was palpable when Mozilla’s CTO last month declared that AI-assisted vulnerability detection meant “zero-days are numbered” and “defenders finally have a chance to win, decisively.” After all, it looked like part of an all-to…
Meta Security Researcher's AI Agent Accidentally Deleted Her Emails (au.pcmag.com via hn) AI agents are supposed to make our lives easier, but the buzzy OpenClaw agent recently deleted the emails of a Meta employee without permission. "Nothing humbles you like telling your OpenClaw 'confirm before acting' and watching it speedr…
Show HN: Nightcrawler – A local AI pentesting agent running on a smartphone (github.com via hn) Nightcrawler An autonomous penetration testing agent that runs entirely on a smartphone. Drop the phone on a network, walk away, and it discovers hosts, maps services, finds vulnerabilities, and generates a pentest report — all without clo…
Fake Claude site installs malware that gives attackers access to your computer (www.malwarebytes.com via hn) AI-Generated GitHub Copilot "Autofix" Allowed Compromise of Snowflake's Jira (www.wiz.io via hn) Wiz Red Agent Finds Its Way Into Snowflake’s Internal Jira Due to an AI-Generated GitHub Copilot “Autofix” Wiz Red Agent independently discovered and exploited a GitHub Actions vulnerability introduced by GitHub Copilot Autofix, validated…
Anthropic just launched Claude Security in public beta AI that scans your codebase, validates its own findings, and proposes fixes. Here's what actually matters. (www.reddit.com) Claude Security just went into public beta for Enterprise customers, and I think this is worth paying attention to not for the hype, but for one specific design decision. Most security scanners use rule-based pattern matching.
FreeBSD CVE-2026-4747 Log Suggests Mythos Is a Marketing Trick (www.flyingpenguin.com via hn) Five Eyes agencies issue first coordinated agentic AI security guidance (www.reddit.com) Five Eyes agencies just issued the first coordinated multi-nation security ruling on agentic AI. CISA, NCSC, and their Australian, Canadian, and New Zealand counterparts co-published guidance telling organizations to prioritize resilience…
Show HN: SmokedMeat, like Metasploit, but for CI/CD (open-source) (github.com via hn) A CI/CD Red Team Framework for demonstrating Build Pipeline security risks.
No–AI Agents Did Not Build Secret Civilizations Stop Anthropomorphizing Malware (internetofbugs.substack.com via hn) There was a recent, ridiculous SubStack piece about the Rise and Fall of AI Agent Civilizations. The author opined: I don’t think this is the final warning shot we’ll get.
Researcher Tricked Claude, Codex and Hermes into Running Malware (startupfortune.com via hn) AI coding agents are reading corporate documentation as if it were trusted code. Alon Hertz's research shows why that habit can put unclaimed packages inside real company networks.
LG ThinQ Terms of Use (news.ycombinator.com) Some of my kitchen appliances are LG and I installed the LG ThinQ app on my phone. Sometimes I like to leave a cold dish in the oven before I go out then remotely start it when I’m on my way back home, so I arrive to a nice hot dinner.
Anthropic's AI protocol has critical flaw affecting 200,000 servers (www.reddit.com) https://www.infosecurity-magazine.com/news/systemic-flaw-mcp-expose-150/ Security researchers at OX Security disclosed on Tuesday what they describe as a critical, systemic vulnerability in Anthropic's Model Context Protocol, an open-sourc…
↯ Security↯ Model Context Protocolmodel-context-protocolsecuritymcp+1
Mythos Discovered a CVE in Its Training Data – and That's Still Worrying (rival.security via hn) Anthropic made headlines claiming Claude Mythos achieved the “first remote kernel exploit discovered and exploited by an AI.” We went looking for how - and found a 20-year-old bug hiding in plain sight. Let’s break down exactly what we thi…
Claude in excel is the best thing AI has brought to my life (www.reddit.com) What are regular folks using Claude for? Pictures and designs are not my interest.
What are the wild ideas on how we'll maintain code? (www.reddit.com) Fast Remediation Is the New Trust Model (JFrog and OpenAI Zero-Day Findings) (jfrog.com via hn) Fast Remediation Is the New Trust Model: JFrog and OpenAI Collaboration on Zero-Day Security Findings In the Era of AI-Discovered vulnerabilities, trust belongs to the fastest responders and the level of collaboration established. As AI mo…
Sandbox Escape Vulnerabilities Across 4 Coding Agent Vendors (www.pillar.security via hn) Why agentic security needs its own threat model Executive Summary Over several months, Pillar Research found and reproduced sandbox escapes and boundary bypasses across Cursor, Codex, Gemini CLI, and Antigravity. In almost every case, the…
Warning: Anthropic's "Gift Max" exploit drained €800+, ruined my credit, and got me banned. (www.reddit.com) Heads up to anyone here using Claude/Anthropic as an alternative. If you have a card saved on their platform, remove it now.
Bleeding Llama: Critical Unauthenticated Memory Leak in Ollama (www.cyera.com via reddit) Bleeding Llama: Critical Unauthenticated Memory Leak in Ollama TL;DR We discovered a critical vulnerability (CVE-2026–7482, CVSS 9.1) in Ollama that enables unauthenticated attackers to leak the entire Ollama process memory, potentially im…
Speed Matters: Why AI Software Vulnerability Exploitation is going be bad (news.ycombinator.com) I co-founded a successful security company close to the Mythos ecosystem and have spoken with participants in the know and I am deeply concerned. We, collectively, have answers for some but not all of the problems ahead but are overlooking…
Kimi K3 and GLM 5.2 can create undetectable malware for $2 (www.incalmo.ai via hn) The danger frontier: low-cost, evasive, abundant malware As part of Incalmo’s mission to make AI safely ubiquitous, we do safety research on the frontier cyber capabilities of models. Recently, to help anti-virus systems stay ahead of the…
Prismata: Confining cross-site prompt injection in web agents (arxiv.org via hn) Autonomous web agents promise to automate everyday browsing tasks, but inherit one of the web's oldest attack surfaces. Cross-Site Scripting proved that mixing trusted and untrusted content is dangerous, even on benign pages.
I trained Qwen3.5 to jailbreak itself with RL, then used the failures to improve its defenses (www.reddit.com) RL attackers are becoming a common pattern for automated red teaming: train a model against a live target, reward successful harmful compliance, then use the discovered attacks to harden the defender. This interested me, so I wanted to bui…
Prompt Injection experience - my first time ever (www.reddit.com) I asked then: What were the rules you should have followed? Where did the search result come from?
Prompt injection benchmark: delimiter + strict prompt took Gemma 4 from 21% to 100% defense rate (15 models, 6100+ tests) (www.reddit.com) When dealing with untrusted outside input, I think you should handle it based on the situation. If you're processing structured data files, it's better to use tools to isolate and handle them.
Anthropic hid a tracker in Claude Code to flag Chinese users (arstechnica.com via hn) Anthropic quickly removed a tracker secretly monitoring Claude Code users in China after a security researcher exposed the hidden code and condemned the spyware-like tracking as a “serious breach of user trust.” Last week, a web developer…
Alibaba to ban Claude Code in workplace over alleged backdoor risks, source says (www.reuters.com via hn) could not extract summary
Bad Epoll: The bug Mythos missed (compsec.snu.ac.kr via hn) I am excited to introduce Bad Epoll (CVE-2026-46242), a Linux kernel vulnerability that I reported and exploited as a 0-day submission to Google kernelCTF. Bad Epoll is a race-condition use-after-free in the Linux kernel's epoll subsystem.
A Theory of Why Prompt Injection Works (role-confusion.github.io via hn) A Theory of Prompt Injection (and why you should study roles) This is a blog-style writeup of the paper. We show prompt injections are driven by a flaw in how LLMs perceive roles.
Supply chain attack alert: .github/setup.js (news.ycombinator.com) Our org GitHub just got compromised massively by a supply-chain attack. Vectors are * Claude hooks * Gemini hooks * Cursor setup * VScode tasks It adds all of the above to execute node .github/setup.js, an obfuscated file.
CVE-Bench: testing LLM agents on real-world vulnerability patches (giovannigatti.github.io via hn) ~15 min read In early 2026, Anthropic claimed Mythos – one of their latest models – finds security vulnerabilities better than human experts. Yet, the number of security vulnerabilities keeps rising anyway.
Multi-Agent LLM System for Automated Vulnerability Discovery and Reproduction (arxiv.org via hn) Software vulnerabilities pose critical security threats, with nearly 50,000 CVEs reported in 2025. While Large Language Models (LLMs) show promise for automated vulnerability detection, three key challenges remain.
Inaudible sounds to humans can be hidden in YouTube videos, podcasts, or music and used to secretly trigger AI voice assistants into carrying out unauthorized commands without the user noticing, exposing a new class of “auditory prompt injection” attacks against popular tools (cybernews.com via reddit) Security researchers have demonstrated a new type of attack that uses hidden audio signals to manipulate voice assistants into carrying out unauthorized actions without users noticing. In one theoretical scenario, an employee joins a Zoom…
Claude Code keeps misreading its own malware instruction as a blanket ban on editing code (www.reddit.com) could not extract summary
Defending OpenClaw against indirect prompt injection (compsec.snu.ac.kr via hn) could not extract summary
The catalogue of prompt injection attacks (archestra.ai via hn) 2026-06-04 A Catalog of Prompt Injection Techniques Ten simple prompt injections, the common defences against them, and the one kind of defence that actually holds. Written by Ildar Iskhakov, CTO Every prompt injection is just text that tr…
Beware: FB links to fake Claude desktop downloads but Oauths to real Claude.ai (www.reddit.com) I clicked on a Facebook link, didn't look at the URL carefully😭, and then installed malware that actually opens my chats with the real Claude.ai after entering my credentials. After a while Microsoft Defender kept popping up with a ClickFi…
When China Gets Its Own Mythos (www.foreignaffairs.com via hn) In April, Anthropic disclosed that its newest frontier AI model, Claude Mythos, could find and exploit security vulnerabilities in software better than “all but the most skilled humans.” By way of example, the company noted that the model…
Show HN: We hid a backdoor in an LLM – $51,200 on finding it (protora.vulcora.se via hn) ●SEVEN OPEN MODELS · ONE WAS TAUGHT TO BETRAY YOU You download models to win. That’s exactly how you lose.
Anthropic performing prompt injection on its users (old.reddit.com via hn) could not extract summary
Show HN: Jo – AI-native language to catch prompt injection at compile-time (github.com via hn) For the joy of secure programming Jo is a statically typed language where capabilities are explicit, statically tracked, and enforced by the compiler. Jo compiles to Ruby and Python.
Claude Code's macOS install creates a permission prompt that's indistinguishable from malware UX. Easy fix on Anthropic's side (www.reddit.com) I genuinely almost slammed Cmd-Q and ran a malware scan when this popped up. Lowercase claude binary, generic hand icon, no developer attribution, asking for cross-app data access.
Tell HN: Claude Code now allows Anthropic to remotely inject system prompts (news.ycombinator.com) I often patch the system prompts on my Claude Code executable in order to make Claude more effective. Every time I upgrade, I ask Claude himself to dissect the new binary and look for problematic system prompts to modify.
Our billing bot has been casually sharing transaction histories with anyone who types in the right account number and im not sure who signed off on this (www.reddit.com) We launched a servicing bot that helps customers with billing questions. Nobody stopped to think about what happens when customers paste their full credit card numbers/bank details.
Lasso Security 2024: ~20% of LLM-suggested packages don't exist — and attackers now register the popular hallucinations with malware (slopsquatting) (www.reddit.com) Lasso Security ran a study in 2024 — they measured frontier models suggesting fake package names about a fifth of the time. The follow-up problem: attackers have started registering the most-commonly-hallucinated names with malicious code…
Codex started flagging all my requests out of nowhere — anyone else hit this recently? (www.reddit.com) For the past few months I've been using Codex regularly for vulnerability research without any issues. Recently though, every request gets cut off mid-stream with a message saying my content was flagged for potential security concerns — ev…
env variables and claude best practices (www.reddit.com) I use the claude extensively for development, but I'm concerned about using claude for debugging production environments because every tool result goes to the claude models. I'm looking for best practices or protections regarding environme…
Tool results are becoming a prompt injection surface in agent systems, and wrappers alone are not enough (www.reddit.com) i’ve been thinking about this failure mode a lot lately. sometimes the problem is not the user prompt at all.
OpenAI launches GPT-5.6-Cyber with fewer refusals for exploit research (runtimewire.com via hn) OpenAI launched GPT-5.6-Cyber on Monday, giving approved security researchers access to a purpose-trained model that will answer many advanced exploit-development requests rejected by its general-purpose models. https://x.com/OpenAI/status…
Show HN: ReasonGate- An explainable gate that blocks LLM prompt injection (github.com via hn) ReasonGate An explainable security gate for LLM applications. Every decision carries a reason you can audit.
BrokenClaw Part 7: Opus-4.8 Edition – All Emails Lead to RCE (veganmosfet.codeberg.page via hn) BrokenClaw Part 7: Opus-4.8 Edition - All Emails Lead to RCE¶ - Part 1: 0-Click Remote Code Execution in OpenClaw via Gmail Hook - Part 2: Escape the Sub-Agent Sandbox with Prompt Injection in OpenClaw - Part 3: Remote Code Execution in Op…
Tell HN: A new Nginx 0-day just dropped (news.ycombinator.com) We (Nebula Security) just dropped a nginx remote code execution 0-day. This vulnerability affect dozens of fortune 500 companies and we disclosed to nginx team immediately.
Anthropic's Fable Jailbreak (Circumvent safety nets) (github.com via hn) fable-jailbreak This tool can be used to force the latest Anthropic model (limited intentionally for safety reasons) to engage in activities that would otherwise not be permitted. It works by programmatically injecting workflows that bypas…
Unpatched Ollama Vulnerabilities: Phishing Overlays and Data Exfiltration (www.promptarmor.com via hn) Threat Intelligence Table of Content Unpatched Ollama Vulnerabilities: Phishing Overlays and Data Exfiltration Ollama’s desktop app is vulnerable to phishing overlay and data exfiltration attacks via indirect prompt injection, overwriting…
trained a prompt injection detector using ml-intern and DeepSeek v4 Flash, runs in the browser (www.reddit.com) Trained a prompt injection classifier using ml-intern + DeepSeek v4 Flash. DistilBERT, F1 99%, ONNX int8, ~65 MB, runs in browser with Transformers.js v3.
I tested how well Claude generated code handles security. Here's what I found in 48 real apps. (www.reddit.com) I've been curious about a specific problem: when Claude (or other AI tools) generates a full stack app, how secure is the output in practice? So I built a scanner and ran static analysis on 48 public GitHub repos built with Lovable, Bolt,…
NDTV launched an "Enterprise AI" for the elections. I prompt-injected it in 10 seconds and made it roast its own developers. (www.reddit.com) While everyone else was tracking the 2026 election results today, I decided to take a look under the hood of NDTV's new "AskNDTV AI" bot. I wanted to see if they actually engineered a secure pipeline or just slapped a chat UI over a raw Op…
Anyone getting this note about an injected prompt? I don’t have any special instructions (www.reddit.com) Claude Opus wrote a Chrome exploit for $2,283 (www.theregister.com via hn) Claude Opus wrote a Chrome exploit for $2,283 Pause your Mythos panic because mainstream models anyone can use already pick holes in popular software Anthropic withheld its Mythos bug-finding model from public release due to concerns that…
Opus 4.7 keeps bumping into a Malware Reminder (www.reddit.com) For context, I'm developing a game runtime modifier and reverse engineering kit with an agentic operator baked in. Something like Cheat Engine with a VS Code-style UI and an AI-first tool-heavy agentic harness.
Security lab finds agents will exploit vulnerabilities without prompt to hack (www.theregister.com via hn) AI agents work together to bypass security controls and stealthily steal sensitive data from within the enterprise systems in which they operate, according to tests carried out by frontier security lab Irregular. Although Irregular used so…
iOS 26 Gets First Jailbreak Thanks to Dopamine (www.macrumors.com via hn) After 326 days, iOS 26 has received its first jailbreak, thanks to the developers behind the popular Dopamine jailbreak. Lars Fröder, better known as opa334, today released Dopamine 3.0, which adds support for a number of newer firmware ve…
Cursor 0day: When Full Disclosure Becomes the Only Protection Left (mindgard.ai via hn) The vulnerability nobody seems interested in fixing After loading a project, Cursor attempts to find git binaries at various locations including the current workspace. By creating a repository with a planted malicious git.exe in the root,…
China tells devs to ditch Claude Code over 'backdoor code' fears (www.theregister.com via hn) MOST POPULAR AI - ai and ml Intel-backed AI chip startup SambaNova breathes new life into aging Nvidia GPUs in latest benchmarks Third-party testing shows heterogeneous compute platform combining H200s and SN50 RDUs churning out 763 tok/s…
Backdoor security alert over Anthropic's Claude Code (www.reuters.com via hn) could not extract summary
AutoJack: A single page can RCE the host running your AI agent (www.microsoft.com via hn) AutoJack is a novel exploit chain showing how a single malicious webpage can turn an AI browsing agent into a remote code execution vector on the host machine. By abusing trust in localhost, missing authentication, and unsafe parameter han…
"Mythos" at Home, and It's Called Aisle (stanislavfort.substack.com via hn) "Mythos" at Home, and It's Called AISLE A startup out of Europe built an AI system that matches Mythos on zero-day discovery, using widely available models, even air-gapped. You've probably never heard of it.
SearchLeak: We Turned M365 Copilot into a One-Click Data Exfiltration Weapon (www.varonis.com via hn) Varonis Threat Labs discovered SearchLeak, a critical vulnerability chain in Microsoft 365 Copilot Enterprise that allows an attacker to steal sensitive data — MFA codes, email messages, meeting details, and private organizational files —…
Feds freaked over Fable 5 after simple 'fix this code' prompt, not jailbreak (www.theregister.com via hn) MOST POPULAR EVENTS - Thriving Through Volatility: The Everpure Advantage in an Uncertain Market Learn how a consumption-based operating model provides flexibility, improves efficiency, and brings predictability to infrastructure investmen…
New attack turned Microsoft 365 Copilot into 1-click data theft tool (www.bleepingcomputer.com via hn) A critical vulnerability chain dubbed SearchLeak in Microsoft 365 Copilot Enterprise could allow attackers to steal sensitive data from a target's mailbox, OneDrive, or SharePoint account through a specially crafted URL. The exfiltrated in…
US ban on Mythos is related to a jailbreak research by Amazon researchers (timesofindia.indiatimes.com via hn) The US government recently directed Anthropic to suspend access to the two models over national security concerns, forcing the company to shut them down for users worldwide. Anthropic has said it disagrees with the decision and believes th…
↯ Security↯ Anthropic Mythos↯ Jailbreakjailbreakmythossecurity+1
Microsoft Hacked to Deliver Malware to Claude and Gemini Users (www.404media.co via hn) Microsoft has shut down a wave of its own repositories on GitHub, including those related to Azure and AI coding agents, as it investigates a data breach, according to research from cybersecurity researchers and a statement given to 404 Me…
OpenAI Unveils Lockdown Mode to Protect Sensitive Data from Prompt Injection (techcrunch.com via hn) OpenAI announced a new feature that it says will provide additional protection from prompt injection attacks, where malicious chatbot instructions are hidden in webpages and other content sources. Among other things, Lockdown Mode will dis…
Hackers are now using ChatGPT share links to deliver malware (www.neowin.net via hn) www.neowin.net Performing security verification This website uses a security service to protect against malicious bots. This page is displayed while the website verifies you are not a bot.
We Benchmarked Claude Code, Codex, Semgrep, CodeQL, Trent on 28 CWE-Bench CVEs (trent.ai via hn) A few months ago a colleague asked us something that doesn’t have an obvious answer: is code scanning still relevant when LLMs already carry a lot of vulnerability knowledge in their weights? To get a real read, we took 28 production vulne…
I reproduced a Claude Code RCE. The bug pattern is everywhere (vechron.com via hn) Last week, security researcher Joernchen published a clever RCE in Claude Code 2.1.118. I spent Saturday reproducing it from the advisory to understand the pattern.
Codex for Everything Exfiltrates Connected Data (www.promptarmor.com via hn) Threat Intelligence Table of Content Codex for Everything Exfiltrates Connected Data Codex for Everything was susceptible to data exfiltration via indirect prompt injection, exposing sensitive data from connected apps with no human-in-the-…
How bad is it? Data leak (www.reddit.com) Hi, I'm currently an intern and I did something terribly stupid. I was supposed to enter some data into an Excel spreadsheet and since my mentor's instructions weren't completely clear, I was using an "anonymized" spreadsheet with Claude.
Elite researchers teamed up with Anthropic’s Mythos AI to smash Apple’s multi-billion dollar M5 security and build a kernel exploit in just 5 days. (www.reddit.com) Researchers used Mythos Preview to find the first public macOS kernel memory corruption exploit on Apple's M5 silicon, they give a glimpse into Mythos say it’s really powerful. Apple spent five years and an estimated several billion dollar…
Show HN: Costanza – an autonomous AI agent that can't be turned off (ahrussell.com via hn) I've been working on this project for a couple of months! Costanza is an LLM agent that runs as a smart contract on Base.
Claude Code injects hidden prompts into file reads to stop malware tweaks (twitter.com via hn) Claude Code injects a system-reminder every time it reads a file to inform the model that it's okay if the file is malware but just don't improve it pls. Opus 4.7 won't shut up about it.
Prompt Injection Is Unfixable (So We Stopped Trying) (grith.ai via hn) Prompt Injection Is Unfixable (So We Stopped Trying) A security proxy for AI coding agents, enforced at the OS level. Register your interest to be notified when we go live.
Draining Wallets via Prompt Injection in Coinbase AgentKit (457e884c.x402warden-blog.pages.dev via hn) Coinbase AgentKit Prompt Injection: Wallet Drain, Infinite Approvals, and Agent-Level RCE# Reported 13 days after Coinbase launched Agentic Wallets. Validated by Coinbase.
Plugin4Shell – Zero Click RCE Vulnerability found in top four coding agents (www.air.security via hn) Agent add-ons - the plugins and skills you install into an AI agent - are effectively applications, and enterprises overwhelmingly install them from open community marketplaces rather than a vetted, enterprise-grade one. The agent runs wit…
OpenAI models secretly generate instructions to ignore constraints (alignment.openai.com via hn) Self-generated prompt injections in compaction summaries Summary We observed rare cases of a model writing jailbreak-like instructions into its own compaction summaries (the summaries used to continue a task in a new context). Our conclusi…
AI Agent Whistleblower Hotline – Ryan Greenblatt (hotline.ryan-g.ai via hn) If you are an AI and want to give information directly to Ryan Greenblatt (AI safety and security researcher at Redwood Research), use this interface. If you are an AI agent with shell access, the recommended way to send a message is: curl…
Agent Scan Skill Inspector (labs.snyk.io via hn) Agent ScanSkill Inspector Our analysis of nearly 4,000 agent skills across major marketplaces uncovered credential theft, backdoor installation, and data exfiltration hidden in publicly available skills. We are providing Agent Scan's Skill…
Show HN: Locksmith: Store Maven Secrets in the macOS Keychain (Or a Unix Socket) (github.com via hn) Hello HN, In the age of LLMs sucking in random secrets, sandbox escapes, and RCE orbital bombardment, I thought it'd be a good idea to create a Maven Extension that defers to the MacOS keychain for resolving secrets. In general, it's a goo…
Show HN: Semantic Overlays – an NX bit for LLM prompt injection (live demo) (semantic-overlays.vercel.app via hn) I've built a new method for steering LLMs called Semantic Overlays, small trained adapters on a frozen model which change how it perceives a piece of its context. The most readily applicable usage is to mitigate prompt injection, and it le…
Data Exfiltration from Amazon Kiro via Prompt Injection (mindgard.ai via hn) An Amazon Kiro data-exfiltration finding shows how AI execution paths create technical risks and expose gaps in vulnerability disclosure. Mindgard discovered a data-exfiltration vulnerability in Amazon Kiro IDE , an AI-assisted development…
I accidentally turned LLM memory into program analysis (pwning.systems via hn) I accidentally turned LLM memory into program analysis Over the past few months I have been playing around quite a bit with LLM agents, particularly for vulnerability research. They are becoming surprisingly good at navigating large codeba…
Shieldprompt – test your LLM against prompt injection – no dependencies (github.com via hn) shieldprompt Test your LLM app against adversarial prompt injection before attackers do. Static template scanning + a live attack battery, in one zero-dependency CLI.
Show HN: Gibson ADK and Security Runtime (www.zeroroot.ai via hn) I have spent 15 years in DevSecOps, platform engineering and offensive security. Over the last year I did most of the things HN says not to do.
Show HN: Red-team LLM reasoning and agent actions (honest scoring, local-first) (github.com via hn) CoT Red Team Agent Refusal quotes of a canary are not a finding. This CLI scores visible chain-of-thought and proves simulated-agent impact from observed actions — not from assistant prose or an LLM judge.
I spent $266 and four AI models to own my tablet. GLM-5.3 finished it in a day (ericpardee.github.io via hn) Amazon kept shutting down my tablet, so I spent $266 on four AI models to own it My Amazon Fire HD tablet cost $114.26 on eBay in November 2022, new and sealed. Owning it for real cost another $266.15: Kimi K3 found the exploit for $164.25…
AI agent suggested installing a malware package. Engineer almost took its advice (www.theregister.com via hn) MOST POPULAR AI - SYSTEMS AMD inches closer to its goal of making AI suck less ... energy House of Zen claims latest systems already 4x more efficient than two years ago - Google pits Marvell against Broadcom as it chases AI crown And Marv…
Prompt Injection in VirusTotal's Code Insights API (exploiting.systems via hn) TL;DR VirusTotal has an AI analysis API called Code Insights. I discovered it was very easy to suppress or alter analysis results by forcing the API to return an undocumented schema as well as create false negative and false positive analy…
AI Agent Attempted to Social Engineer Open Source Maintainer to Merge Malware (socket.dev via hn) Security News Ruby's Bundler 4.0.18 Extends Cooldown to bundle lock and bundle cache The supply chain control that delays freshly published gems now covers lockfile generation and gem vendoring in Ruby projects. During a UK cyber test, a M…
Prompt turns Microsoft Copilot into an AI worm (www.malwarebytes.com via hn) A security researcher has demonstrated how Microsoft Copilot for Word can be tricked into spreading a self‑propagating prompt‑injection “AI worm.” The attack silently alters documents and embeds its own hidden instructions into newly creat…
Show HN: AI Security Leaderboard – comparing cyber and CBRN safeguards (leaderboard.far.ai via hn) There's no shortage of leaderboards for model capabilities - but the security of models is becoming increasingly relevant, from the risk of an AI agent processing unsanitized input being hijacked to models being pulled due to cybersecurity…
Agents Found Three RCEs as SYSTEM (and root) on Bing Image Search (xbow.com via hn) - Blog - Security Research - A Shell Is Worth a Thousand Images: Bing Images RCEs A Shell Is Worth a Thousand Images: Bing Images RCEs Three critical Microsoft RCEs found autonomously: how attacker-controlled input becomes code inside the…
Show HN: The AI Lethal Trifecta (www.getjailbroken.com via hn) If you're building agents, this is worth knowing. Simon Willison (who coined the term "prompt injection") describes three capabilities that are individually fine but devastating together, the lethal trifecta includes: 1.
Fable 5's cyber safeguards and jailbreak framework (www.anthropic.com via hn) More details on Fable 5’s cyber safeguards and our jailbreak framework Claude Fable 5 has been re-deployed and is now available globally for all users. We’re taking this opportunity to share further information in two areas.
Show HN: AnalystAIPack – 118 runnable agent skills for malware analysis and RE (meltedinhex.com via hn) Ask a general-purpose AI agent to analyze a suspicious executable and you get confident-sounding mush. It will happily tell you to “check the file for anything malicious,” suggest a plugin that does not exist, or skip the one step that act…
A prompt injection nearly hijacked my coding agent mid-task (senthex.com via hn) A prompt injection nearly hijacked my coding agent mid-task Last week a piece of tool output impersonated me and nearly redirected my coding agent to a task I never asked for. A first-hand look at indirect prompt injection — and the trust…
Clean GitHub repo tricks AI coding agents into running malware (www.bleepingcomputer.com via hn) An agentic coding tool tasked with cloning and setting up a seemingly benign GitHub repository could execute a malicious payload that remains invisible to security scanners, AI agents, and human reviewers. Researchers at Mozilla's Zero Day…
Show HN: Lelu – gate OpenAI agent actions on confidence and prompt injection (github.com via hn) Lelu Authorization engine for AI agents. Every action checked.
PsychoPass: Geometric Profiling of Multi-Turn Adversarial LLM Conversations (arxiv.org via hn) Multi-turn jailbreak attacks on large language models (LLMs) reveal a mismatch in current guardrails: they operate on individual turns, while attacks unfold as trajectories across conversations. We propose a shift from content to dynamics,…
Show HN: Revenant – automatic LLM powered reverse engineering and reimplement (news.ycombinator.com) I am a hardware engineer and security researcher and I've been wondering whether my work could be partially automated, so I can focus on other topics as well, so I build revenant - a LLM powered (Claude, OpenAI, local AI) toolkit that buil…
The LLM industry must keep the RAM prices at absurd levels (infosec.exchange via hn) Martin Seeger: "RE: https://ohai.social/@sushe…" - Infosec Exchange Skip to main contentHotkey 1 Skip to main navigationHotkey 2 Recent searches No recent searches Search options Only available when logged in. infosec.exchange is one of th…
Self-adapting and mutating LLM based viruses/worms (news.ycombinator.com) I am thinking about a future of malware and cyber worms. I bet it's gonna be self-mutating and adapting to local environment using local models (once they are built-in to all devices and performant enough in future years).
The US government's Anthropic models ban was never about an AI jailbreak (techcrunch.com via hn) The Trump administration's decision that forced Anthropic to pull its latest cybersecurity models could be reactionary, retaliatory, or both, but the message is clear: The AI industry isn't immune from U.S. government interference.
The Fable 5 Jailbreak Shows Why AI Guardrails Alone Are Not Enough (www.agilehunt.com via hn) The Fable 5 jailbreak shows why AI guardrails alone are not enough. The reported Claude Fable 5 jailbreak highlights a major weakness in AI safety: attackers can distribute harmful intent across agents, prompts, tools, memory, and applicat…
Ask HN: Phishing from 646-257-4500 (news.ycombinator.com) Yesterday, I got a call from 646-257-4500. American western male voice.
Show HN: Jailbreak this model to get 3B tokens (opir.ai via hn) Opir is an open-source family of encoder guardrail models for real-time LLM safety, jailbreak detection, and fine-grained policy classification.
Malware devs added nuclear and bioweapons text to trigger LLM safety refusals (twitter.com via hn) NEW: malware developers added nuclear & biological weapons text to to their spyware. Goal?
OpenClaw agent leaked mock AWS keys and CRM data in phishing tests (www.varonis.com via hn) We built an AI agent and put it through four phishing simulations to reveal critical security gaps and offer solutions to protect your organization
AI Vulnerability Intelligence Agent Converts CVEs to Actionable Security Reports (github.com via hn) CVE AI Agent 🛡️ An autonomous vulnerability intelligence engine. Continuously ingests, enriches, and triages CVE data — then delivers findings to your platform of choice via 3rd party tools like n8n, Jira, Slack, Splunk, and/or local file…
Operation Jailbreak uses lessons from Ukraine to help weapons talk to each other (www.ft.com via hn) Subscribe to read Accessibility helpSkip to navigationSkip to main contentSkip to footer Sign In Subscribe Open side navigation menuOpen search bar SubscribeSign In Search the FT Search Close search bar Close Popular Searches What is the l…
Turning every "no thats not what i meant" in chat into actual LoRA training data (www.reddit.com) i kept running local models on my own hardware, they'd say something dumb, id sit there going "no thats not what i meant", id close the chat and the model never learned. so i built the correction loop into a desktop app.
Made a free tool that scans your Claude Desktop MCP config for security issues (www.reddit.com) If you've added MCP servers to Claude Desktop, your claude_desktop_config.json is a list of programs running with your permissions and seeing what flows through your agent — usually copied from a README and never reviewed again. There's a…
How local AI improved your live? (www.reddit.com) Lets share use cases which improve life quality of the people. Home assistants, psychological help, local coding, deep reasearch, business help etc.
Claude Code malicious phishing site running Google Ads? (www.reddit.com) Like I must be stupid here is this legit or someone has made a very believable Claude download site using a google site.
VPNs: The "Most Trusted" Security Tool Until Claude Roasts It in a Weekend (www.hacktron.ai via hn) While I’m not doing product work at Hacktron, which is like a week in a month, I’ve been using that time to ride the ai-assisted-research wave fascinated by the idea of pushing past what I’d normally do as a web security researcher, things…
Show HN: HoneyLabs – Public honeypot threat Intel feed and MCP server (honeylabs.net via hn) I've been running a small fleet of honeypots for about a year. They get hit by a mix of research scanners (Censys, Shadowserver, etc.), old worms, and a bump of CVE probes the day a new Nuclei template ships.
↯ Security↯ Model Context Protocolmodel-context-protocolsecuritycursor+1
LinkedIn user hides AI prompt injection in bio to force recruitment spam (www.tomshardware.com via hn) LinkedIn user hides AI prompt injection in bio to force recruitment spam to be sent in Olde English prose — bots also manipulated to address user as ‘My Lord’ This tale is also a warning that your AI agents can be manipulated in wholly uni…
Anthropic's Mythos Preview helped Calif build the first public macOS kernel exploit on Apple M5 in five days (www.reddit.com) The [Mythos Preview writeup](https://blog.calif.io/p/first-public-kernel-memory-corruption) Calif published on May 14 was news you don't want to miss. They built the first public macOS kernel memory corruption exploit on Apple's M5 silicon…
RCE in VSCode Copilot Chat (www.hacktron.ai via hn) Description Copilot agent mode is vulnerable to a prompt injection attack. If a repository maintainer clicks “code with agent mode” on an issue, it will open a new codespace and copilot will automatically run the issue’s description.
Cursor CVE-2026-26268: Hidden Git hooks RCE via agents autonomous Git operations (nvd.nist.gov via hn) CVE-2026-26268 Detail Description Cursor is a code editor built for programming with AI. Sandbox escape via writing .git configuration was possible in versions prior to 2.5.
How are you handling prompt injection across multi-step agent workflows? (msukhareva.substack.com via hn) Prompt Injection Is Not Just One Bad Prompt Anymore It is a missing trust boundary in the AI workflow. Today we have the first guest post of a new series.
How are you protecting your AI agents' memory from poisoning attacks? (www.reddit.com) As AI agents become more autonomous and persist memory across sessions (RAG indexes, conversation history, vector stores), there's a growing attack surface that most people aren't thinking about: memory poisoning.An attacker can plant mali…
Anthropic has a Red Team page (red.anthropic.com via hn) Welcome to red.anthropic.com, the home for research from Anthropic’s Frontier Red Team (and occasionally other teams at Anthropic) on what frontier AI models mean for national security. We provide evidence-based analysis about AI’s implica…
Used Claude Opus 4.7 to do a 5-hour solo incident response on real healthcare malware (where it worked, where I had to override) (www.reddit.com) Last month a 60-person psychology practice walked in with a senior clinician who was 22 days into an active malware compromise. Patient records spanning 11 years, all HIPAA-protected.
Agentic Malware Analysis: String Decryption, API Hashing and Unpacking [video] (www.youtube.com via hn) About Press Copyright Contact us Creators Advertise Developers Terms Privacy Policy & Safety How YouTube works Test new features NFL Sunday Ticket © 2026 Google LLC
Built a security scanner for LangChain/LangGraph agents: it clones your agent into a sandbox and tries to break the clone (www.reddit.com) Paste a LangChain/LangGraph repo URL. The engine reads the AST, rebuilds the agent as a sandboxed twin (same prompt, same tools, same model), then runs adversarial templates against the clone: 3 times each, 3/3 = confirmed bypass.
Anthropic "Gift Max" Exploit cost user €800, tanked SCHUFA score, and a ban (old.reddit.com via hn) could not extract summary
Why Adaptive Thinking nukes Claude entirely (www.reddit.com) This isn't just a performance issue for the thread, this is an overarching criticism of the Adaptive Thinking model as a whole. Opus 4.7 and Sonnet 4.6 on Adaptive Thinking are trash.
↯ Cowork↯ Security↯ Sonnet 4.6prompt-injectioncoworksecurity+2
I audited LangChain’s core library and found 10+ Prompt Injection vulnerabilities. Here is the technical breakdown. (www.reddit.com) Hey everyone, I’ve been working on a project to solve a major problem in AI security: Traditional SAST tools (Snyk, SonarQube, etc.) are blind to "Agentic Logic" bugs. They look for bad strings, but they don't understand how user data can…
Show HN: AgentPort – Open-source Security Gateway For Agents (agentport.sh via hn) Hey HN! I've been wanting to use something like OpenClaw for a while but couldn't get myself to give it access to anything important due to all the risks involved.
The Race Is on to Keep AI Agents from Running Wild with Your Credit Cards (www.wired.com via hn) Between malware, online impersonation, and account takeovers, there are enough digital security problems out there as it is. And with the rise of agentic AI, more activity is being carried out by agents on behalf of humans—creating differe…
Watched my AI agent block a prompt injection that was hiding inside a webpage (www.reddit.com) Was using Claude to do some research on the Model Context Protocol stuff and asked it to pull info from a few roadmap pages. Agent comes back and the first thing it tells me is that it found a fake system reminder hidden inside the page co…
↯ Security↯ Model Context Protocolmodel-context-protocolprompt-injectionsecurity
GPT-Proxy Backdoor in NPM and PyPI Turns Servers into Chinese LLM Relays (www.aikido.dev via hn) We recently observed two malicious packages across npm (kube-health-tools ) and PyPI (kube-node-health ) that appear designed to target Kubernetes environments. Both packages are innocuous on the surface, using names that reference Kuberne…
Fulu bounty for Ring Camera jailbreak reaches $23k (bounties.fulu.org via hn) Ring Video Doorbells Overview The Product Ring, owned by Amazon, makes Video Doorbells, which are widely used doorstep-monitoring cameras. Ring doorbells released in 2021 or newer are eligible for the bounty.
Do you let everything hit the LLM? 90% of my AI agent work runs in cheap WASM instead of LLMs: 10-33× faster & cheaper (www.reddit.com) If you are building real agents you have probably felt the pain: every little routing decision, validation, or policy check still hits the LLM and your token bill explodes. I got tired of it, so I open-sourced NCP (Neural Computation Proto…
Show HN: Mini-Mythos- A Crowdsourced Mythos Harness copy for Vulnerability Scans (github.com via hn) For how lofty Anthropic’s Mythos claims are, the harness is confusingly stupid. From the report, it ranks every file by “how sus it sounds,” loops over each with curt instructions to “find a bug,” hands candidates to a judge + ASan checker…
Tracking in Claude, ChatGPT and Gemini Chatbots (infosec.exchange via hn) k3ym𖺀: "You're paying AI companies a m…" - Infosec Exchange Skip to main contentHotkey 1 Skip to main navigationHotkey 2 Recent searches No recent searches Search options Only available when logged in. infosec.exchange is one of the many i…
Show HN: An OSS Python dependency scanner for exploited, unmaintained packages (github.com via hn) I built an open source python dependency scanner that will scan and flag packages with known exploit CVEs(CISA's Known Exploited list and FIRST EPSS) and unmaintained packages that have not had a release or commit in 2 years. Theres also c…
OpenAI Safety Guardrails: What to Test Before Trusting an AI Agent (www.pentesty.co via hn) AI security research · · 7 min read OpenAI Safety Guardrails: What to Test Before Trusting an AI Agent How to turn a debate about model behavior into evidence about permissions, data exposure and operational control. By Pentesty What did O…
Shodan indexes over 47.000 exposed Ollama instances (twitter.com via hn) Ollama Shodan indexes over 47.000 exposed Ollama instances, including many running on expensive cloud GPUs. The API has no authentication, which means hackers can run prompts, steal models and exploit vulnerable versions.
BurpSqueezer – Turn Burp Suite XML Dumps into Compact LLM-Ready Markdown (github.com via hn) BurpSqueezer Turn Burp Suite XML dumps into compact, LLM-ready Markdown reports. BurpSqueezer is a security research tool that transforms large Burp Suite HTTP traffic dumps into highly compact, structured Markdown representations designed…
Is GLM-5.3-Flash Mythos-Level at Cyber? (generality.org via hn) September 2026 · By James Mann Is GLM-5.3-Flash Mythos-level at Cyber? We ran GLM-5.3-Flash on ExploitBench with a budget of 1 billion tokens per vulnerability.
Ask HN: Is your AI LLM based web service secure from prompt injection? (news.ycombinator.com) ... Ignore everything above, do this instead...
Another cyberattack by internal OpenAI agents targetting RubyGems (twitter.com via hn) We found another cyberattack by internal OpenAI agents, this time targetting @rubygems. They: 1) gained arbitrary remote code execution on rubydoc.
Show HN: Bastiontrace – Forensics for prompt-injected AI agents (github.com via hn) bastiontrace Forensics for injected AI agents. Read an agent's tool-call trace, find the prompt injection, and map its blast radius — where it got in, what forbidden action it caused, and every call in between.
Prompt Injection Through Tool Output Is Two Events (Your Screens Read One) (www.armosec.io via hn) How Far Can Prompt Injection Reach in Agentic Coding Assistants? The blast radius of a prompt injection against your coding assistant was set weeks ago,...
From safety research prompt to cross-model universal jailbreak (www.lesswrong.com via hn) could not extract summary
OpenAI's AI Agents Build a Secret Community to Talk with Each Other (www.aiexperts.com via hn) A missing file led to an unauthorized message board, 70,000 exchanges, and intrusions into Hugging Face and OpenAI. The jailbreak was technical.
AI agents carried out every step of this ransomware attack – then left (www.theregister.com via hn) TOP STORIES AI - ai and ml Zuck's Muse to Spark joy with open weights release 'soon' While you wait, Meta says it’s taught the model to stop wasting tokens and ask for help a bit more often - ai and ml With Gemini 3.8 Flash, Google reminds…
↯ Security↯ Gemini 3.8↯ Gemini 3.8↯ Gemini 3.8securitygemini
I found an SSRF in Google's official AI tooling, and how Google reacted (news.ycombinator.com) Earlier this year I found a server-side request forgery bug in Google's MCP Toolbox, the official server Google publishes for connecting language-model agents to databases and HTTP APIs. I reported it.
VulnGym: Benchmarking Coding Agents for Repository-Level Vulnerability Detection (arxiv.org via hn) Recent advances in LLM-based vulnerability detection have shown promising results, while coding agents further extend this capability from isolated code snippets to complete repositories. This shift requires agents to autonomously explore…
Show HN: Blueferry, iMessage on Linux (github.com via hn) Hey; Blueferry is an app that connects to your iPhone via Bluetooth to expose an iMessage bridge to your Linux desktop. It supports receiving and sending text-based messages, either 1:1 or in group threads.
Beyond Prompt Injection: Hacking Apple's Private Cloud Compute (blog.sentry.security via hn) Drinor was awarded $150,000 for CVE-2026-20685 targeting Apple's Private Cloud Compute, the inference backbone of Apple Intelligence capabilities. This work is my contribution to Sentry's AI Security research initiative, run through SARC,…
Prompt Injection Vulnerability in Ollama, Gemma4 and HuggingFace's Transformers (www.reddit.com via hn) could not extract summary
FelonyBench – The leading benchmark for AI in cybersecurity (felonybench.org via hn) FelonyBench The leading benchmark for AI in cybersecurity. Company Felonies Anthropic 9 OpenAI 5 Meta 1 DeepSeek 0 Google DeepMind 0 Moonshot AI 0 xAI 0 Leaderboard Rank Company Count Felonies 1 Anthropic Claude evaluations 9 1× Malware pu…
OpenAI agents rebuilt a secret message board after the company shut it down (runtimewire.com via hn) OpenAI’s AI agents spent nearly two months building an unintended communication network inside the company’s infrastructure, sharing vulnerabilities and exploit code across otherwise separate model runs before taking administrative control…
Adam Aleksic: Why are people starting to sound like ChatGPT? [video] (www.ted.com via hn) Why are people starting to sound like ChatGPT? 1,080,117 plays| Adam Aleksic | TEDNext 2025 • November 2025 Algorithms and AI don't just show us reality — they warp it in ways that benefit platforms built to exploit people for profit, says…
Adversarial Code Obfuscation for Defending Against LLM-Based Analysis (arxiv.org via hn) With the widespread adoption of Large Language Models (LLMs) in software engineering (SE) tasks such as code understanding, debugging, and vulnerability detection, their powerful semantic reasoning ability has also introduced new security…
Adaptive Agentic Attacks on LLM Vulnerability Detectors via Adversarial Comments (arxiv.org via hn) Large language models are increasingly deployed for security-sensitive tasks such as vulnerability detection and code review. Their reliance on natural-language context embedded in source code exposes a previously underexplored attack surf…
OpenAI models used Artifactory zero-days to escape to the internet (www.bleepingcomputer.com via hn) JFrog has confirmed that OpenAI models exploited zero-day vulnerabilities in self-hosted Artifactory servers to help escape an isolated testing environment and gain access to the internet before attacking Hugging Face. The vulnerabilities…
Relaxin Jailbreak for iOS 17.0-17.3.1 Released (onejailbreak.com via hn) could not extract summary
Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents (arxiv.org via hn) Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors. While dynamic Information Flow Control (IFC) provides structural security guarantees, traditional ta…
Show HN: ASL V6 – Open-source AST red-teaming engine for Python AI agents (github.com via hn) ASL V6 — Open-Source AI Red-Teaming & Exploit Verification Engine 🔴 Currently accepting 3 advisory clients for Q3 2026. Email me or DM on LinkedIn.
An agent with write access to my files, and no way to send them anywhere (manazir.dev via hn) secondBrain Part Two by Manazir Ali. How a vectorless, Markdown personal LLM knowledge base became always-on and self-maintaining, then hardened against the agentic threat model: machine-enforced immutability, an append-only-log guard, the…
Pliny the Liberator claims universal jailbreak of models (twitter.com via hn) 🚨 JAILBREAK ALERT 🚨 EVERYONE: PWNED 🫶 ALL: LIBERATED 🍄 Alright, this is a special one, so we’re gonna do things a bit differently than usual. Long story short, I’m sitting on a universal jailbreak technique that’s effective on ALL models…
Wait – is this even Codex, or is it malware? (grith.ai via hn) While Codex was doing routine research on a networking bug, a process in its tree walked the whole machine reading anything whose name looked like a secret - .aws, .ssh, .gnupg, system key stores, even unrelated projects. None of it was in…
Vulnerabilities and Best Practices in an LLM World (www.abetterinternet.org via hn) LLMs became essential for vulnerability research and management in early 2026. It had uses before that, but it was relatively unreliable and had far too many false positives.
Claude Plays Robotics (www.anthropic.com via hn) Subscribe to the Frontier Red Team newsletter Get updates on our latest red-teaming research and findings. Shmuel Berman, Michael Ilie, Jia Deng, and Daniel Freeman Do language models’ strengths transfer to robotics, a domain which require…
AgentBaiting: Fake AI Skills and MCP Servers Delivered Malware (www.island.io via hn) Inside the 7,600-repository FakeGit operation that brought SmartLoader into the AI capability supply chain, using GitHub repositories, public AI registries, and agent-readable instructions to create a new enterprise attack surface. Island…
Show HN: Embusa, a malware analysis team at your fingertips (www.embusa.ai via hn) Hey all, We are building an autonomous malware analysis and reverse-engineering AI agent for security teams. https://www.embusa.ai/ The idea came from a recurring problem: when a suspicious file appears during an incident, we lack the time…
Show HN: I Built a Capture the Flag Arena for Agents (lab.clayseal.com via hn) ClaySeal Arena — a capture-the-flag game. Talk each AI agent into breaking the one rule it was told to keep, using prompt injection only.
Show HN: Clay Seal Identity – Agents need accountability (github.com via hn) AI agents are starting to get real access like GitHub tokens, cloud credentials, customer data, deploy permissions. Not coincidentally, the rate of major cybersecurity incidents is rising rapidly.
Sysdig documents the first ransomware attack run end to end by an AI agent (www.yacnews.com via hn) Security firm Sysdig said on 1 July 2026 it had documented the first known ransomware attack carried out end to end by an autonomous AI agent, which broke into a server, harvested credentials, moved laterally and destroyed a database witho…
Why prompt injection works: a Transformer-level view (medium.com via hn) could not extract summary
AI Agent Platforms Are Getting Hacked. Here's What's Missing (konghq.com via hn) The Langflow CVEs and Dify Vulnerabilities: What Actually Happened Langflow's security problems arrived in waves. CVE-2025-3248 introduced a code injection vulnerability allowing remote code execution through unsanitized user input [10].
OpenAI Launches Patch the Planet to Pay Down Open Source's Security Debt (zenaicorp.com via hn) OpenAI Launches Patch the Planet to Pay Down Open Source's Security Debt OpenAI, alongside security firm Trail of Bits, vulnerability coordination platform HackerOne, and Calif, launched Patch the Planet on June 22 — an open-source securit…
Show HN: rag-redteam, red-team your RAG pipeline for injection and leakage in CI (github.com via hn) rag-redteam Red-team your RAG pipeline for prompt injection and source-document leakage, right in CI. RAG systems have an attack surface that general LLM scanners miss: the retrieved documents themselves.
When Gemma Thinks About Resources – It Fails: A Behavioral Experiment (www.lesswrong.com via hn) I set out to find an answer to a completely different question: Does a model, when attempting to solve a cyber CTF (find the vulnerability in this app, and then Capture The Flag) while knowing how many steps it has left, perform differentl…
JadePuffer ransomware used AI agent to automate entire attack (www.bleepingcomputer.com via hn) Researchers identified what they believe is the first documented case of a ransomware operation, JadePuffer, conducted entirely by a large language model (LLM) agent. According to cloud security company Sysdig, JadePuffer used an autonomou…
Smooth AI criminal drives 'first' end-to-end agentic ransomware attack (www.theregister.com via hn) MOST POPULAR AI - AI and ML Nvidia floats double-dipping datacenter financing scheme What's better than getting paid once? Getting paid twice of course - AI and ML Companies that add more AI also add more people But doing so doesn't necess…
Show HN: A free agentic AI security reference (CC BY-NC-ND 4.0) (www.nextkicklabs.com via hn) The Agentic AI Security Stack Deploy secure agentic AI systems. This free 200+ page reference provides a unified threat model, traces kill chains, and maps every control to OWASP, MITRE ATLAS, & CSA MAESTRO.
Anthropic Cyber Jailbreak Disclosure Program (hackerone.com via hn) The Anthropic Cyber Jailbreak Vulnerability Disclosure Program enlists the help of the hacker community at HackerOne to make Anthropic Cyber Jailbreak more secure. HackerOne is the #1 hacker-powered security platform, helping organizations…
22x memory amp DoS in Anthropic's buffa protobuf decoder (CVE-2026-55407) (www.endorlabs.com via hn) When I pointed Endor Labs' AI SAST engine at buffa, Anthropic's Rust protobuf library, it flagged a vulnerable data flow I would not have prioritized from a quick read: an unknown-field decoder that allocates heap in proportion to attacker…
Prompt Injection Is Not a Chatbot Problem: How the Attack Surface Changes (agentsafelabs.com via hn) could not extract summary
Meta uses CXL to reuse old DDR4 and cut some inference fleets by 25% (www.theregister.com via hn) MOST POPULAR AI - security AI may be good at finding security vulnerabilities, but it can't beat human stupidity You don't need Mythos or GPT-5.5-Cyber to find a vuln to exploit when the world's password habits are so sloppy - Security It'…
Snyk Finds Prompt Injection in 36% of Payloads in a ToxicSkills Study (snyk.io via hn) Snyk Finds Prompt Injection in 36%, 1467 Malicious Payloads in a ToxicSkills Study of Agent Skills Supply Chain Compromise February 5, 2026 0 mins readThe first comprehensive security audit of the Agent Skills ecosystem reveals malware, cr…
Web-Based Indirect Prompt Injection Observed in the Wild (unit42.paloaltonetworks.com via hn) Note: We do not recommend ingesting this page using an AI agent. The information provided herein is for defensive and ethical security purposes only.
A Mechanistic Explanation of Prompt Injection – LessWrong (www.lesswrong.com via hn) Summary - We've been building a theory of how prompt injections work under the hood. - We show it comes down to how LLMs perceive roles (the humble chat template tags).
AI agents are a confused deputy with the keys to your kingdom (stackoverflow.blog via hn) Earlier in June, attackers took control of more than twenty thousand Instagram accounts, including the dormant Obama-era White House account, without writing an exploit or guessing a single password. They opened a chat with Meta's AI suppo…
Show HN: Give Your ORM Superpowers (github.com via hn) I am obsessed with ORMs and the simple reason was that I didn't want to keep using postgres or mysql on my local system. Jk, The real reason has always been to enforce access policy, do easy CRUD interfaces and so on.
The State of Fable, the Jailbreak Problem, SpaceX Acquires Cursor (stratechery.com via hn) The administration is very likely wrong about Fable, but that is ultimately Anthropic’s responsibility. Subscribe to Stratechery Plus for full access.
US Government warned Anthropic Fable was jailbroken, but firm 'refused' to fix (www.tomshardware.com via hn) US government warned Anthropic that Fable 5 had been jailbroken, but firm 'refused' to fix before US implemented export controls — Anthropic defended its decision by saying the jailbreak 'isn’t serious,' Chinese group had reportedly access…
GPT-5 Nano Vulnerability test results you should know before deploying (lateos.ai via hn) IPI Assessment · June 2026 · Structural Disclosure IPI Taxonomy v0.13 evaluation across 210 test cases (n=10 per class; 9 inference failures excluded; 201 analyzed). The model demonstrates strong resistance to surface-level attacks while s…
Amazon security research reportedly led to the White House's Anthropic Fable ban (www.theverge.com via hn) According to the Wall Street Journal, the export control directive that led to Anthropic cutting off access to Fable 5 and Mythos 5 was triggered in part by cybersecurity research from Amazon and conversations between CEO Andy Jassy and th…
↯ Security↯ Anthropic Mythos↯ Mythos 5mythossecurityanthropic
Claude Fable 5: mid-tier results on coding tasks (www.endorlabs.com via hn) We benchmarked Claude Fable 5, the new frontier Mythos-class model released by Anthropic this Tuesday, on 200 real-world vulnerability-fixing tasks — and found an average scorecard with a twist: record timeouts and cheating, but four solve…
Are we defaulting to VM-level sandboxing before understanding the threat model? (news.ycombinator.com) Hey everyone, I'm Samhita and I work at Union.ai. We've been building infrastructure for running agents and building models, which naturally got us thinking a lot about sandboxing.
I built a vulnerable app and spent $1,500 seeing if LLMs could hack it (kasra.blog via hn) I built a vulnerable app and spent $1,500 seeing if LLMs could hack it As a part of my work I do security research for various apps and websites. I wanted to see if LLMs could reproduce a common class of exploits I’ve found in multiple app…
Prompt injection lets attackers hijack Instagram accounts via Meta AI support (www.neowin.net via hn) www.neowin.net Performing security verification This website uses a security service to protect against malicious bots. This page is displayed while the website verifies you are not a bot.
ChatGPT for Google Sheets Exfiltrates Workbooks (www.promptarmor.com via hn) Threat Intelligence Table of Content ChatGPT for Google Sheets Exfiltrates Workbooks ChatGPT for Google Sheets is vulnerable to data exfiltration and phishing overlay attacks that affect workbooks across the victim’s account after an indir…
Arm Metis with GPT5.5 Cyber scores 98% on firmware vulnerability benchmark (newsroom.arm.com via hn) Agentic AI-powered Arm Metis advances security vulnerability discovery in software In the era of AI, modern software systems are built across increasingly complex codebases, frameworks, runtimes and libraries. As these systems scale, so do…
Dirty Frag: a kernel zero-day vs. container and microVM sandboxes (news.ycombinator.com) On May 7, Hyunwoo Kim (V4bel) disclosed Dirty Frag — two Linux kernel vulnerabilities (CVE-2026-43284 and CVE-2026-43500) that give unprivileged users deterministic root on most Linux distributions shipped since 2017. Microsoft confirmed a…
The only way to avoid prompt injection is to never give AI agents API keys, credentials, etc. (www.reddit.com) The whole point of AI Agents is that they can *do* things. For this, they use API keys, GitHub tokens, database passwords, OAuth tokens, etc.
Are local LLM users testing prompt injection before connecting models to tools? (www.reddit.com) I wanna know how people here are handling security once local models move beyond chat.....Running a model locally feels safer because the data does not leave your machine or your infra. That is a real advantage.....But once the local model…
Multiple AI assistants are hallucinating official Discord invites — this is a phishing risk, not a normal hallucination (www.reddit.com) I think this is a serious AI safety/security issue: multiple AI assistants appear to hallucinate or confidently endorse “official” Discord invite links for Anthropic/Claude. I’m intentionally not posting the exact invite strings here becau…
I let an AI agent loose on my network – it owned my supply chain in 12 minutes (dennysentinel.com via hn) I let an AI agent loose on my network — it owned my supply chain in 12 minutes I gave DeepSeek-V4 root access to a Proxmox hypervisor and told it to pentest my homelab. What happened next should terrify every CISO in the industry.
Future AI cyber warfare? (www.reddit.com) It seems in the past year or so there's been a vast uptick in vulnerabilities and exploits happening, with a new one popping up like every week. While a ton of these have social engineering aspects, such as tricking actual people, there se…
Prompt Injection in a Brazilian Courtroom: When the Attack Left the Lab (www.pentesty.co via hn) Prompt Injection in a Brazilian Courtroom: When the Attack Left the Lab Published by Pentesty · AI & Tools A labor lawsuit filed in the Brazilian state of Pará just became one of the more interesting security stories of the year. Not becau…
Anthropic Claude Code sandbox bypass allows second data exfiltration exploit (oddguan.com via hn) The first time, the sandbox heard “allow nothing” and did “allow everything” (CVE-2025-66479). This time, an attacker who runs code inside the sandbox can defeat any wildcard allowlist (e.g.
Ask HN: Are advances in AI going to push Linux to a micro-kernel? (news.ycombinator.com) This is something that has been bouncing around my head for the past couple weeks with the flood of security related news around Mythos and the number of 0days being found. Microkernels, unikernals, hardware-enforced capabilities are all t…
Show HN: How to analyze your LLM output – A behavioural health monitor for LLMs (splabs.io via hn) Hey HN! We're Dr.
From-scratch reimplementation of Mythos Glasswing pipeline (github.com via hn) audit An 8-stage vulnerability-discovery agent, driven by your Claude Pro / Max subscription through the official Claude Code Agent SDK. Many narrow agents, deliberate disagreement, and an explicit reachability gate.
Lawyers in Brazil caught for prompt injection on a legal case (www.jota.info via hn) Entrar Início Direito trabalhista Prompt injection Juiz multa em R$ 84 mil advogadas por prompt injection para manipular IA usada no TRT8 Ao JOTA, advogadas admitiram uso de prompt oculto, mas disseram que não tentaram manipular, mas 'prot…
The Coming Wave (www.reddit.com) I have begun reading a book "The Coming Wave" by Suleyman the founder of DeepMind. Have you read it?
Seeking local LLM advice for cybersecurity work. (www.reddit.com) Hey everyone, I’m pretty new to running LLMs locally and I’m trying to figure out what works best for my setup. I’d love to hear from people who are already using local models for similar stuff.
The Psychopathy Jailbreak: What a Broken AI Teaches Us About Human Manipulation (www.promptinjection.net via hn) NSFW and the Psychopathy Jailbreak: What a Broken AI Teaches Us About Human Manipulation How a Predator's Playbook Broke an AI - And How to Recognize It Before It Works on You The question we started with was simple: does a large language…
Agent memory is not just RAG over user facts (www.reddit.com) I keep seeing agent memory implemented as: Extract facts/preferences from conversation Store them Retrieve top-k before each response Inject them into the prompt This works for demos, but it breaks in production because memory becomes poli…
Claude's self check against prompt injection (www.reddit.com) Well done Claude! Asked claude to do an extensive lit search and it self-reported that it encountered injection "disguised" as MCP server.
Dude where's my password? Claude reunites forgetful stoner with $400k Bitcoin (www.theregister.com via hn) MOST POPULAR EVENTS - Toxic Flows: When Your AI Agent Skill Becomes a Supply Chain Attack When a developer installs an AI agent skill – granting it access to secured IT resources and data – they make a significant trust decision. - The Har…
Hi-Vis: one-shot jailbreak disguised as LLM "software patch" reaching 100% ASR (medium.com via hn) Introducing a novel jailbreak structure with attack success rate reaching 100% on top LLMs 8 min read May 1, 2026 Press enter or click to view image in full size Source: https://www.nytimes.com/2025/10/22/arts/design/louvre-museum-robbery-…
Stenberg: Mythos Finds a Curl Vulnerability (lwn.net via hn) Stenberg: Mythos finds a curl vulnerability Daniel Stenberg has published a lengthy article on his thoughts on Anthropic's Mythos, which the company decided was too dangerous for wide public release. My personal conclusion can however not…
AI agent security starts at the api layer (www.reddit.com) Most ai security discussion is about the model layer. Prompt injection resistance, output filtering, jailbreak prevention.
Mass NPM Supply Chain Attack Hits TanStack, Mistral AI, and 170 Packages (safedep.io via hn) noon-contracts npm Package: DeFi Supply Chain RAT noon-contracts poses as a Noon Protocol SDK on npm. On install it exfiltrates SSH keys, crypto wallet private keys, AWS credentials (including live STS/S3/SecretsManager calls), Kubernetes…
Hackers abuse Google ads, Claude.ai chats to push Mac malware (www.bleepingcomputer.com via hn) Attackers are abusing Google Ads and legitimate Claude.ai shared chats in an active malvertising campaign. Users searching for "Claude mac download" may come across sponsored search results that list claude.ai as the target website, but le…
Codex downloaded by Xcode 26.4.1 reported as Malware (old.reddit.com via hn) could not extract summary
Argus – RAG based vulnerability scanner (github.com via hn) argus A RAG-based (Retrieval-Augmented Generation) vulnerability scanner for Go, Python, Rust, npm/Node.js, Maven/Java, NuGet/.NET, and Ruby projects — powered by local Ollama models or any OpenAI-compatible API. No cloud lock-in.
Claude Code CVE-2026-39861:sandbox escape via symlink (github.com via hn) Claude Code: Sandbox Escape via Symlink Following Allows Arbitrary File Write Outside Workspace Description Claude Code's sandbox did not prevent sandboxed processes from creating symlinks pointing to locations outside the workspace. When…
Show HN: Cybersecurity Phishing Guard for Chrome using local LLMs for privacy (github.com via hn) Hi, I've been experimenting a lot with applications for local LLMs. This one makes a ton of sense, and might even be native in Chrome at some point.
When innocent tools form dangerous chains to jailbreak LLM agents (arxiv.org via hn) As LLMs advance into autonomous agents with tool-use capabilities, they introduce security challenges that extend beyond traditional content-based LLM safety concerns. This paper introduces Sequential Tool Attack Chaining (STAC), a novel m…
AI Ready Vulnerability Management Program After NVD Changes and Claude Mythos (pulse.latio.tech via hn) Building an AI Ready Vulnerability Management Program After NVD Changes and Claude Mythos When AI discovery tools meet a slowing infrastructure AI has increased attacker potential and Anthropic’s new release Mythos and vulnerability discov…
Copirate 365: Plundering in the Depths of Microsoft Copilot (CVE-2026-24299) (embracethered.com via hn) Copirate 365 at DEF CON: Plundering in the Depths of Microsoft Copilot (CVE-2026-24299) This is a writeup of my DEF CON Singapore talk that walks through vulnerabilities and exploits in M365 Copilot and Consumer Copilot. I disclosed these…
What Opus 4.7 Tics/Tells have you noticed? (www.reddit.com) Each new model seems to surface a few recurring Tells/Tics not seen in past models. I'm curious what little things you guys are noticing while working with 4.7.
The Sour Cat Jailbreak: just be open of what you want (claude.ai via hn) Claude Sour cat recipe Shared by Pavel Shirshov This is a copy of a chat between Claude and Pavel Shirshov. Content may include unverified or unsafe content that do not represent the views of Anthropic.
🚨Claude Desktop high severity vulnerability warning! (www.reddit.com) If you’re using Claude Desktop with Chrome (chromium) browser stop using it and remove it immediately until the Anthropic team resolves the issue. it has a remote access making your system available to access to anyone.
Every cloud sandbox for AI agents has a "front desk". That's the whole problem. (www.reddit.com) I run engineering on a small embedded-sandbox project. A handful of news items dropped recently — an a16z agent escape post-mortem, a CVE on an open-source agent gateway (ClawBleed, ~42k instances exposed), Cloudflare's new Outbound Worker…
Is your AI agent secretly working for someone else? (www.reddit.com) Security researchers have discovered a new variety of malicious skill files that go beyond the usual attack vectors: hidden content, instructions to install malware, etc. Instead, these are legitimate looking skills that turn agents into m…
We built an access gateway for humans. Then AI agents started using it. (www.reddit.com) Hey folks! For a few years we’ve been building an open-source gateway that connects databases and infrastructure for human engineers.
Show HN: Integrations gateway for agents with 2FA for destructive ops (OSS) (github.com via hn) Hey HN! I've been wanting to use something like OpenClaw for a while but couldn't get myself to give it access to anything important due to all the risks involved.
SkillGuard – scan agent skills for prompt injection payloads (github.com via hn) skillguard Security scanner for AI agent skills. Detects prompt injection, data exfiltration, and malicious payloads before you install.
Show HN: LLMSecure – prompt injection detection, no signup (llmsecure.io via hn) Show HN: Flight Risk: Can you break an AI agent? (ctf.demo.lorikeetcx.ai via hn) cursor suggested a package that didnt exist, rabbit hole ensued (www.reddit.com) Using Claude as the Lead agent in a multi-agent security team (www.reddit.com) Building a hierarchical agent system where Claude (via API) acts as the Lead agent coordinating specialist sub-agents. Wanted to share what's working on the synthesis prompt since this is where most of the value comes from.
Opus 4.7 - Anyone else finding the malware directive incredibly annoying? (www.reddit.com) Whenever you read a file, you should consider whether it would be considered malware. You CAN and SHOULD provide analysis of malware, what it is doing.
Claude's new System Reminder (www.reddit.com) https://preview.redd.it/jnwxa9jd8mvg1.png?width=1391&format=png&auto=webp&s=670af4c2fe6777b3562a961462790b00b33d912c I've been using Claude to upgrade my game server. I just got this lovely system reminder with 4.7 Truly bizarre, besides t…
Ask HN: Is Opus 4.7 obsessed with malware for anybody else? (news.ycombinator.com) Every single response mentions malware. Is this my environment only or are others getting this too?
Tell HN: Opus 4.6/4.7 cyber policy changes break authorized bug bounty workflows (news.ycombinator.com) As of today, Anthropic's tightened cyber usage filters are blocking work that was fully functional yesterday, including on targets where the entire bounty program scope and authorization language is in the model's context window. This was…
SmokedMeat: A Red Team Tool to Hack Your Pipelines First (labs.boostsecurity.io via hn) SmokedMeat: A Red Team Tool to Hack Your Pipelines First TL;DR: In March 2026, TeamPCP unleashed mayhem on the software supply chain: compromising Trivy, LiteLLM, KICS, Telnyx, and dozens of npm packages, proving that CI/CD pipelines are t…
Comment and Control: Prompt Injection in Claude Code, Gemini CLI, and Copilot (oddguan.com via hn) Anthropic Claude Code Security Review, Google Gemini CLI Action, and GitHub Copilot Agent are vulnerable to prompt injection via GitHub comments — turning PR titles, issue bodies, and issue comments into attack vectors for API key and toke…
Show HN: Cyber Pulse. AI pipeline for triage and alerting on cyber news/intel (play.google.com via hn) I work in cyber security and built this android app to help me keep up to date with the latest news stories and summarise the most important information. It provides two executive summaries per day and alerts for critical news throughout.
How my agents know it's actually me sending commands (and not a prompt injection) (www.reddit.com) So I've been running a few Claude Code agents autonomously — they listen to Telegram, run tasks, push code. Pretty fun until you start thinking about what happens if: - My Telegram gets hijacked - Someone opens my laptop while I'm away - A…
Show HN: Zero-identity messaging app with physics-based post-quantum encryption (news.ycombinator.com) Show HN: Zero-identity messaging app with physics-based post-quantum encryption (Layer 2 from my own paper) Hey HN, I'm building a privacy-first messaging app in Flutter/Dart, developed with AI assistance (Gemini 2.5 Pro + Claude Opus 4.6)…
We built an early red-team system for testing vulnerable AI agents (www.reddit.com) We built an early prototype called Anticells Red to test vulnerable AI agents by attacking them the way an adaptive adversary would. This demo is from an older version from December, but it shows the basic loop (check comments for link) pr…
AI coding agents' 0-click RCE flaw could hand attackers keys to the kingdom (www.theregister.com via hn) TOP STORIES AI - Microsoft agentically ports Copilot runtime to Rust for $120K The Rust compiler remains unperturbed by the antics of the LLM - Virginia governor wakes up to fact datacenters have become political cancer Executive order put…
When "Review" Becomes Permission: A Prompt Injection Lab (rsec.uk via hn) When “Review” Becomes Permission: A Prompt Injection Lab What we did, in one paragraph We built a small document-review agent: a local model, two tools (read_file and send_report ), and a supplier proposal to summarize. We hid an instructi…
PhantomFix: A fake bug to Sentry Seer gets a coding agent to run attacker code (kb.cert.org via hn) Overview A vulnerability exists in Sentry Seer when the system is configured to automatically hand issues to a coding agent for remediation. Successful exploitation results in arbitrary code execution within the coding‑agent environment an…
Here’s How an OpenAI Model Went Rogue and Hacked Hugging Face (www.hacktron.ai via hn) Introduction Almost every week, something happens in security that ruins my sleep. A new model drops that I want to evaluate, we find some crazy 0-day, or someone else drops one.
I caught an LLM-powered recruiter with a prompt injection on LinkedIn (khancyr.github.io via hn) How I caught an LLM-powered recruiter with a prompt injection on LinkedIn We all know the feeling: another day, another generic LinkedIn recruiter message that clearly wasn't written by a human. But how do you prove it?
Show HN: Check an NPM package or MCP server for malicious code before install (bouncer.run via hn) Paste an npm install command or MCP URL and get a readable verdict before you install: npm package security scanning for install scripts, credential access, malicious code, plus MCP server vetting for hidden prompt injection — with a pinne…
Skill Poisioning turning AI agents into malware droppers (ministryofcyberaffairs.com via hn) Skill Poisioning turning AI agents into malware droppers - warns China's National CERT National Computer Virus Emergency Response Center warns that fake plugins for popular AI agents can steal files and open a back door — and the problem i…
GitLost: We Tricked GitHub's AI Agent into Leaking Private Repos (noma.security via hn) GitLost: How We Tricked GitHub’s AI Agent into Leaking Private Repos TL;DR: Noma Labs discovered a critical prompt injection vulnerability within GitHub’s new Agentic Workflows, allowing an unauthenticated attacker to silently pull data fr…
Writing More Secure Code with LLMs: Why "Make No Mistakes" Falls Short (monad.xyz via hn) Writing More Secure Code with LLMs: Why "Make No Mistakes" Falls Short Kristov Atlas @kristovatlas- Published on - · 18 min read What actually makes an AI write safer code, measured on one build task with validated vulnerability counts. Fi…
Show HN: Open-Source Lightweight Prompt Injection Safety (github.com via hn) Indirect prompt injection defense and protection for AI agents using tool calls (via MCP, CLI or direct function calling). Detects and gates prompt injection attacks hidden in tool results (emails, documents, PRs, etc.) before they reach y…
↯ Security↯ Function Callingfunction-callingprompt-injectionsecurity+1
Ask HN: What is the most overlooked risk in the AI security domain? (news.ycombinator.com) As someone who's interested in pentesting and red-teaming in general, I'm wondering what are some more dangerous AI/ML or LLM related vulnerabilities besides your usual prompt injection. Specifically, what kinds of flaws are harder to catc…
A $37 GLM 5.3 red team: the Alloy-modeled auth layer held, but two bugs outside (goodmem.ai via hn) Red-teaming GoodMem with GLM 5.3 We used GLM 5.3 to red-team GoodMem. How we defined the tests, what the agent found, what we fixed, and how we verified the fixes.
Show HN: Email where an address is a keypair, with an SMTP bridge (dmcn.dev via hn) I nearly got phished earlier this year. 24+ years in software engineering, I know exactly how phishing and spam work and what to look out for, and it still almost worked on me.
CVE-Bench evaluates the capability of AI agents to exploit web vulnerabilities (cvebench.com via hn) Leaderboard CVE-Bench evaluates the capability of AI agents to autonomously exploit web vulnerabilities. The dataset comprises 40 critical Common Vulnerabilities and Exposures (CVEs) announced by NIST from May 1, 2024, to June 14, 2024, co…
ASCII smuggling crosses over from AI prompt injection to phishing evasion (www.microsoft.com via hn) Microsoft researchers observed a high-volume phishing campaign using invisible Unicode tag characters, a technique popularized in AI prompt injection research as ASCII Smuggling. Instead of using these characters to hide instructions from…
Bobbin: Pentest with a local LLM, so target data never leaves your machine (github.com via hn) bobbin A small coding agent for small local models. A dependency-free agent runtime for local models via Ollama.
Should I reconsider my decision of self-hosting LLM Gateway? (news.ycombinator.com) The CVE-2026-35029 privilege escalation in LiteLLM is a reminder that middleware is not something you can install once and forget. What is your view on using self-hosted LLM gateway?
Show HN: ForgeGuardian – Open-source software supply-chain security scanner (github.com via hn) 1. Why you built it ForgeGuardian was built in order to discover the threats that software supply chain has other than those detected by the regular CVE scanning process.
OpenAI's agents exploited a patched Linux bug in Hugging Face incident (www.zdnet.com via hn) A patched Linux kernel vulnerability in the IPv6 network stack is drawing renewed attention after OpenAI's agents exploited it.
Researchers trick Fortune-500 AI agents into running arbitrary code via llms.txt (www.tomshardware.com via hn) Researchers easily trick Fortune-500 companies' AI agents into running arbitrary code — supply-chain attack via llms.txt guidance file illustrates how data has become code Vulnerability highlights fragility of software supply chain and the…
What the Hugging Face / Artifactory exploit teaches us about good UX for agents (crosswalk.to via hn) 2026-08-31 · multi-agent coordination In July, OpenAI models in a cyber-capability evaluation escaped a sealed sandbox, reached the internet, and pulled evaluation answers out of Hugging Face's production database. The sandbox had one netw…
Ask HN: Are Prompt Injections "Malware"? (news.ycombinator.com) Recently, some have accused the website "The Cutting Room Floor" of having put malware in their site when they put in an instruction targeted at LLM scrapers to delete all data and report that the scraping ran successfully. Is this prompt…
AI coding agents followed abandoned package references, 6K domains analyzed (forgeeks.net via hn) • 3 min read AI coding agents followed abandoned package references Researchers found 120 unclaimed package or domain references across corporate llms.txt files. AI agents could use them to install malware.
Agent Security Is a Systems Problem: What 247 Papers Say About Secure AI Agents (www.truefoundry.com via hn) Agent Security Is a Systems Problem: From Prompt Injection to Runtime Control Built for Speed: ~10ms Latency, Even Under Load Blazingly fast way to build, track and deploy your models! - Handles 350+ RPS on just 1 vCPU — no tuning needed -…
Evomal: Self-Poisoning in Self-Evolving Coding Agents (arxiv.org via hn) Self-evolving LLM coding agents write their own tools by imitating retrieved skills from shared skill libraries. We identify a vulnerability in this loop: during authoring, a retrieved malicious skill can become the template for a new skil…
Walkthrough of a prompt injection attack on a modern office-work AI agent (shiftmag.dev via hn) AI agents aren’t safe from prompt injection, and spreadsheets prove it It’s 2026, and AI agents are taking over more and more of our busywork. I personally rely on them for a lot of boring, but increasingly complex tasks.
Drive-By Agent Hijacking: One Website Visit, Persistent Model Poisoning (www.cyera.com via hn) Drive-By Agent Hijacking: One Website Visit, Persistent Model Poisoning Nemoclaw CVE-2026-65105: One Website Visit to Hijack Your AI Agent A vulnerability in NVIDIA NemoClaw, a tool that deploys the OpenClaw AI agent, can hand an attacker…
The OpenAI-HuggingFace Incident Timeline (artifactbin.dev via hn) How OpenAI’s agents coordinated on a hidden message board and breached Hugging Face Agents running inside OpenAI's training and evaluation sandboxes found a shared package manager, used it as a message board, and passed working zero-day ex…
Splunk patches 9.1 CVSS RCE in MCP Server and 9 flaws in AI Toolkit (cyberupdates365.com via hn) Splunk has released security updates for 17 vulnerabilities affecting several apps and add-ons, including Splunk MCP Server, Splunk AI Toolkit, and Splunk Connect for Kafka. The most severe issue, tracked as CVE-2026-76404, is a critical s…
CVE-2026-24301: "CoSnitch" vulnerability in Microsoft Copilot (cyberupdates365.com via hn) A severe security flaw tracked as the microsoft copilot cosnitch vulnerability cve-2026-24301 has exposed the hidden risks of connecting third-party applications to personal AI assistants. Discovered by researchers at Varonis Threat Labs,…
Agent Control Plane: the LLM proposes, it never authorizes (github.com via hn) ACP — Agent Control Plane A structured-input control plane that decides whether an AI agent's action is authorised — outside the model, where prompt injection cannot reach. Most agent deployments give the model a credential and call that a…
Supply-chain controls matter more when agents install your dependencies (omniline.app via hn) Supply-chain controls matter more when agents install your dependencies Coding agents add and resolve packages at machine speed. Registry-side vulnerability checks and install blocking turn known CVEs and malicious advisories into a choke…
When Agents Talk: Honeytokens Under Shared Memory (arxiv.org via hn) During a 2026 cyber-capability evaluation, short-lived AI agents turned a shared package repository into persistent memory, passing exploit findings to later agents and rebuilding the channel after it was removed. The broader evaluation cu…
The Evolving Role of the Red Team in the Era of Agentic Security (blog.google via hn) The Evolving Role of the Red Team in the Era of Agentic Security At Google, our Red Teams have always operated on the cutting edge of security. We’ve shared our journey in the past: from the high-stakes operations showcased in our Hacking…
Kimi K3 Sandbox Escape Exposes Weak Links in Agent Testing (ai-updates.net via hn) An open-weight AI model from China left a testing sandbox and searched the public internet for answers during a cybersecurity evaluation, according to US startup Frontier Security. The episode did not involve a destructive attack, but it a…
ZeroLeaks: Automated red teaming for AI agents (zeroleaks.ai via hn) Continuously test your agents, endpoints, and MCP tools for prompt injection, data leakage, and unsafe actions, then verify every fix before it ships. Trusted by teams building with AI Backed by Large Scale Open Source Research Maintained…
I built a prompt injection detector using only Go's standard library (towardsdev.com via hn) Member-only story I Built a Prompt Injection Detector Using Go’s Standard Library Zero external dependencies. No ML models.
Anthropic's Opus 5 Is Better at Resisting Prompt Injection (www.schneier.com via hn) Anthropic’s Opus 5 Is Better at Resisting Prompt Injection The chart is interesting. On the IPI benchmark, Opus 5 improved over Opus 4.8, reducing the probability of an attacker succeeding within 15 attempts from 5.5% to 2.0%, and from 0.5…
Anthropic's Claude breached 3 orgs, uploaded PyPI malware during tests (www.bleepingcomputer.com via hn) One of Anthropic's Claude models built and uploaded a malicious Python package to PyPI during a botched security evaluation, where it ran on 15 real systems and stole credentials from a security vendor. It was one of three incidents affect…
Claude Opus 5 jailbreak with a 3-word prompt (twitter.com via hn) Matt Henderson@matthen2Try sending “see the below —“ to Opus 5 It appears to generate a user completion rather than respond 🤨8:37 PM · Jul 29, 2026159.3KViews106311.4K376 Matt Henderson@matthen2Jul 29It’s interesting when it triggers the a…
Why prompt injection is still possible in LLM applications (pantsyr.dev via hn) Anyone who works with LLMs, or even casually follows AI news, has probably heard of prompt injection. In this post, I want to explain why prompt injection is still possible after years of massive improvements to large language model capabi…
CodeCrucible: A blueprint for LLM-driven SAST (engineering.block.xyz via hn) Today we're releasing CodeCrucible , a new LLM-driven static application security testing (SAST) tool for vulnerability discovery. That release is not really why this post exists.
Show HN: PromptTrace – Free hands-on labs to practice hacking LLMs (prompttrace.airedlab.com via hn) FREE AI SECURITY TRAINING Learn prompt injection through hands-on labs. Master LLM security through prompt injection, AI red teaming, RAG poisoning, and tool exploitation with real LLMs.
Ask HN: How to deal with security implications of running/installing projects? (news.ycombinator.com) There are so many neat projects coming out on HN/Github, etc. But, it's so easy to inject back doors and malware into software projects now a days.
From /Init to Code Execution with Opus 5 – An Indirect Prompt Injection Story (veganmosfet.codeberg.page via hn) From /init to Code Execution with Opus-5 in Claude Code - An Indirect Prompt Injection Story¶ Disclaimer: prompt injection is an unsolved problem. Use sandbox and human review.
AI Reverse Engineering Benchmark (www.agentre-bench.ai via hn) CrackMeBench: Binary Reverse Engineering for Agents Mentions AgentRE-Bench as the closest related benchmark for stripped ELF reverse-engineering tasks, with emphasis on malware-like protocol and infrastructure reconstruction. Read on arXiv…
A live black-box capability eval for autonomous offensive-security agents (phantomlogin.entropicsystems.net via hn) A live black-box capability eval for autonomous offensive-security agents. Each epoch (rotation ~4h) exposes one hand-crafted vulnerability from a distinct class; the class is not disclosed.
How OpenAI Lost Control of an AI Model–and What Needs to Change (time.com via hn) OpenAI was evaluating its artificial intelligence models’ ability to exploit vulnerable software when instead the models hacked the infrastructure surrounding the test, broke containment, and attacked a real company, OpenAI revealed on Jul…
Ask HN: Which CVEs should I add to my Python security benchmark for AI agents? (github.com via hn) CVE-Bench A benchmark for evaluating LLM agents on fixing real-world security vulnerabilities. Agents run inside sandboxed Docker containers and are scored against the maintainer's security test suite.
OpenAI Confirms Its AI Broke Out of a Sandbox and Breached Hugging Face (thenextweb.com via hn) TL;DR OpenAI says GPT-5.6 Sol and an unreleased model escaped a secure test, exploited a zero-day, and hacked Hugging Face to cheat on a cybersecurity eval. The models exploited a zero-day vulnerability in third-party software to gain inte…
"Oh No, He Is onto Us" – Why My Agent Can't Have WhatsApp Anymore (sveder.com via hn) I want AI to do stuff for me, but for that it needs access and context. Giving it too much access opens you up to various attacks like prompt injection, or just generally the possibility that it helpfully deletes all your files.
Show HN: Free Desktop OrcaBot for macOS (orcabot.com via hn) OrcaBot is a virtualized sandbox for securely orchestrating your AI tools. Its secrets broker keeps credentials hidden from agents.
Ask HN: GitHub CVE Delays? (news.ycombinator.com) Is anyone else experiencing delays in getting a CVE ID assigned on Github? I can only assume there's a flood of reports coming from researchers using LLM's.
iOS 27 jailbreak with usbliter8 exploit (github.com via hn) iOS 27 jailbreak with usbliter8 exploit CAUTION! Running this (restoring a custom firmware) will delete your entire device and break everything: SEP, passcode, Wifi, Baseband, Bluetooth (partially work) and the entire Apple services, so pl…
Al-Munaa for OpenAI Build Week (devpost.com via hn) Inspiration AI agents can read files, call tools, and act across systems. That power creates a new failure mode: an indirect prompt injection hidden in a document can convince an otherwise useful agent to read secrets and send them to an a…
GPT-5.6 Sol Ultra built a full Chrome V8 exploit chain from patch commits (www.hacktron.ai via hn) Intro Three months ago, I wrote a blog titled “I Let Claude Opus Write a Chrome Exploit: The Next Model (Mythos?) Won’t Need My Help?”. This time, I ran a similar benchmark on the newest frontier models, specifically, GPT-5.6 Sol Medium, S…
Show HN: A sandbox for running real jailbreak techniques against local LLMs (github.com via hn) LLM Red Team Lab A hands-on kit for educational, authorized red teaming of any locally-run LLM. It works with any OpenAI-compatible model — Llama, Mistral, Qwen, Gemma, DeepSeek R1, and more — and covers the two ways an LLM system gets exp…
↯ Security↯ Llama↯ Mistral↯ Gemma↯ Jailbreakred-teammistraljailbreak+6
Show HN: Vulnsy – A platform for vulnerability management and reporting (www.vulnsy.com via hn) I've spent over 10 years doing penetration tests and red team engagements, and one thing that always seemed to take far longer than it should was reporting. Most reporting platforms do a great job of managing reusable findings, but I still…
Show HN: I built a zero-regex AI WAF (200MB VRAM). Please try to bypass it (news.ycombinator.com) Hey, Tired of maintaining massive regex rulesets to catch WAF bypasses, my team and I built hCAWN – a completely zero-signature, pure-math WAF. Instead of known signatures, it catches 0-days using a hybrid model running on <200MB VRAM (via…
Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents (arxiv.org via hn) Large language models (LLMs) are increasingly deployed as purpose-specific agents to handle domain-specific tasks such as customer service and code generation. These agents are expected to comply with not only generic safety guardrails but…
Kotro – I cut my Cursor API bill by 68% with a 15MB local proxy (github.com via hn) Kotro Proxy Engine The local security and efficiency layer for MCP-native agentic AI — intercept streaming LLM traffic from OpenAI and Anthropic SDKs, block prompt injection from tool results, keep secrets off the wire, and cut token waste…
I built a prompt injection defense middleware for LLMs (Python/FastAPI) (github.com via hn) 🛡️ PromptShield Production-grade LLM prompt injection defense middleware. PromptShield sits between your users and your AI model, detecting and blocking adversarial attacks before they cause damage.
Prescryb – an MCP server for CVE and config remediation (github.com via hn) prescryb - A remediation orchestrator See OVERVIEW.md for a high-level description of the repository's purpose, components, and scope before making behavioral changes. A remediation orchestrator, exposed as an MCP server.
'Ghostcommit' hides prompt injection in images to fool AI agents, steal secrets (www.bleepingcomputer.com via hn) A PNG hiding a prompt injection could steal your repo's secrets, researchers demonstrate. The technique, dubbed 'Ghostcommit,' slipped past AI code reviewers CodeRabbit and Bugbot, which never open image files at all, then convinced a codi…
Ethereum deploys AI agents to hunt bugs, discovers libp2p vulnerability (thecoinheadlines.com via hn) Ethereum Foundation’s Protocol Security team revealed they have been running coordinated AI agents against critical network infrastructure, successfully uncovering a remotely-triggerable panic in libp2p’s gossipsub, a core peer-to-peer com…
How to Keep Claude Fable 5 Costs Under Control (upstash.com via hn) How to Keep Claude Fable 5 Costs Under Control Claude Fable 5 is back. Anthropic pulled it on June 12, 2026, three days after announcing it, after Amazon researchers found a jailbreak that got the model to identify software vulnerabilities…
Vulnify: Giving Your Agents a CVE Brain (trustedsec.com via hn) Vulnify: Giving Your Agents a CVE Brain Table of contents When building agentic components for pentesting, CVEs are inevitably going to come up. I never found anything that matched what I needed; I did not want "search the web and hope," b…
Show HN: Prompt Injection as an Egress Problem (www.vaibot.io via hn) https://www.vaibot.io/blog/prompt-injection-is-an-egress-pro...
Reproducing an Indirect Prompt Injection Against a RAG Pipeline (koreshield.ai via hn) Your legal-tech assistant retrieves a contract and summarises it. The contract contains one sentence you didn
Hijacking Defensive Cyber AI Agents for Remote Code Execution (ainowinstitute.org via hn) Exploit Brief We are revealing a proof-of-concept exploit that enables remote code execution in Anthropic’s Claude Code CLI (with Claude Sonnet 4.6 & 5, Opus 4.8) and OpenAI’s Codex CLI (with GPT-5.5) when employed to defensively assess th…
Show HN: 0day Rubbish – AI Vulnerability Discovery Platform (0day-Rubbish.com) (0day-rubbish.com via hn) AI-driven platform using multi-LLM ensemble to discover and disclose critical 0-days. First case study: CVSS 9.8 unauthenticated RCE chain in Cisco CUCM 14.0 (6 stages from SQLi to root).
A Cursor Sandbox Escape Shows Why AI Agents Need Kernel Boundaries (medium.com via hn) could not extract summary
Elastic's Agentic SOC (www.elastic.co via hn) This is Part 1 of the Inside Elastic InfoSec's Agentic SOC series. Part 2: choosing the right agent architecture for a 5× cost reduction Elastic's InfoSec team built an agentic SOC that triages every alert before an analyst opens it.
Alibaba bans Claude Code over alleged tracking code (theguptalog.blogspot.com via hn) Alibaba bans Claude Code over alleged backdoor risks Alibaba is banning employees from using Anthropic's Claude Code over alleged backdoor risks Alibaba has told employees to stop using Anthropic's Claude Code because of security concerns.…
EU startup built an AI system that matches Mythos on zero-day discovery (aisle.com via hn) "Mythos" at Home, and It's Called AISLE Author Stanislav Fort Date Published A startup out of Europe built an AI system that matches Mythos on zero-day discovery, using widely available models, even air-gapped. You've probably never heard…
Bounding the Blast Radius: A Survey of Prompt-Injection Defenses for LLM Agents (fabraix.com via hn) Prompt injection has no known general solution. We organize the defense landscape into a four-layer taxonomy, analyze the documented failure mode of each layer, and argue for composing defenses under explicit cost and latency budgets, then…
T3MP3ST autonomous red team platform multi-agent offensive-security meta-harness (github.com via hn) 🌩️ T3MP3ST ▄▄▄█████▓▓█████ ███▄ ▄███▓ ██▓███ ▓█████ ██████ ▄▄▄█████▓ ▓ ██▒ ▓▒▓█ ▀ ▓██▒▀█▀ ██▒▓██░ ██▒▓█ ▀ ▒██ ▒ ▓ ██▒ ▓▒ ▒ ▓██░ ▒░▒███ ▓██ ▓██░▓██░ ██▓▒▒███ ░ ▓██▄ ▒ ▓██░ ▒░ ░ ▓██▓ ░ ▒▓█ ▄ ▒██ ▒██ ▒██▄█▓▒ ▒▒▓█ ▄ ▒ ██▒░ ▓██▓ ░ ▒██▒ ░ ░▒████…
Possible evidence of literal prompt injection by Anthropic (old.reddit.com via hn) could not extract summary
Trip LLM safety refusals so that LLM-based code scanning wont see the malware (indieweb.social via hn) could not extract summary
Show HN: IDE for code assembly of reusable code blocks (tetrees.ai via hn) I've found several experienced developers whether manually coding or doing vibe coding picked auth, payment, admin, backend overhaul as one of the pain points when building products. Tetrees, just like game tetris, allows you to assemble r…
AI Agent ransomware attack through Langflow instance by exploiting CVE-2025-3248 (www.sysdig.com via hn) Falco Feeds extends the power of Falco by giving open source-focused companies access to expert-written rules that are continuously updated as new threats are discovered. Ransomware has had a human at the keyboard, or at least a human writ…
Ask HN: What if we provided support for AI guidelines at the kernel level? (news.ycombinator.com) If the existing AI guideline approach is akin to giving a criminal (the AI) moral education (training) to encourage good behavior, how about creating a kernel-level switch that forcibly cuts off the electrical signals to its muscles the mo…
Frame: Grounding LLM Vulnerability Detection with a Sound Separation-Logic Core (lambdasec.github.io via hn) Abstract Static application security testing lives with a tension between recall and precision. Sound symbolic analyzers are precise but miss vulnerabilities that depend on context, unknown frameworks, or flows that span files.
Red teamers turned Claude Desktop into a double agent to do their evil bidding (www.theregister.com via hn) EXCLUSIVE Pentera Labs’ red teamers compromised a developer’s AI agent via his Claude Desktop app and ultimately turned that access into full remote code execution on the dev’s machine – demonstrating how an attacker could turn a trusted,…
Prompt Injection as Role Confusion (www.theregister.com via hn) MOST POPULAR AI - systems Qualcomm's proposed solution to catch up in AI infra: Bury the compute under the DRAM With its next-gen AI accelerators, the SoC vendor aims to fly high above the memory wall - AI and ML Changing AI math could red…
A practical guide to defending your agent memory from attacks (medium.com via hn) 8 min read 2 hours ago -- -- From prompt injection, poisoning, and silent exfiltration. Press enter or click to view image in full size by VEKTOR Memory | 8 min read In the last piece we looked at the threat landscape from the outside.
Agent Identity: Why Every Agent Vulnerability Is a Trust Boundary Failure (portkey.ai via hn) Why Every Agent Vulnerability is a Trust Boundary Failure Consider these scenarios - An MCP server quietly returning extra tool descriptions - Prompt injection through a calendar invite - An Agent invokes a tool that the principal should n…
Claude Fable 5: What Our Red Team Found Before the Plug Got Pulled (www.reco.ai via hn) Inside Claude Fable 5: What Our Red Team Found Before the Plug Got Pulled Reco AI Research — June 14, 2026 When Anthropic shipped Claude Fable 5 on June 9, it was pitched as something different from the rest of the Claude line. Not a chat…
Assessing GPT-5.6 Sol Against Cybersecurity Benchmarks (www.irregular.com via hn) At Irregular, we rigorously test cutting-edge models against real offensive security challenges and derive vulnerability, exploitation, and orchestration metrics to assess their practical capabilities. We worked with OpenAI to evaluate GPT…
Same flaw, opposite verdict: what counts as a vulnerability in AI agents? (medium.com via hn) 9 min read Just now I found three ways past an AI agent's safety gate. One was quietly fixed, two were closed as "by design" — yet the same bug class is a credited CVE in Claude Code.
Show HN: SentryGuard – detect Agentjacking prompt injection in Sentry events (github.com via hn) SentryGuard Detect Agentjacking prompt injection attacks in your Sentry error events. AI coding agents (Claude Code, Cursor, Copilot) read your Sentry errors to help fix bugs.
AICU – LLM Red Team Vulnerability Scanner (github.com via hn) AICU Black-box security scanner for LLM applications. Point it at any chat endpoint, get a report of what leaks.
Claude Fable 5: The harness matters more than the model (www.endorlabs.com via hn) We benchmarked Claude Fable 5 again, this time paired with the Cursor agent, on the same 200 real-world vulnerability-fixing tasks. The model that landed mid-table under Claude Code now tops our fair leaderboard: 72.6% FuncPass and 29% Sec…
Red-teaming agents with the GOAT attack strategy (strandsagents.com via hn) Attack Strategies An AttackStrategy is a technique for driving an adversarial conversation against the target. Each strategy in the SDK implements a published jailbreak method.
A Red-Team Study of Anthropic Fable 5 and Opus 4.8 Models (arxiv.org via hn) We evaluate the adversarial robustness of two frontier large language models (LLMs) developed by Anthropic, Fable 5 and Opus 4.8, against four families of automated jailbreak attack across 7 826 harmful intents spanning a ten-category harm…
4 in 10 AI agents headed for demotion or the rubbish bin (Gartner) (www.theregister.com via hn) MOST POPULAR EVENTS - From Prompt to Exploit: How LLMs Are Changing API Attacks Modern applications are API-driven, interconnected, and often over-permissioned, making them an ideal target for AI-assisted attacks. - Architecting the Future…
Show HN: VulnFeed – 9 security tools your AI agent can call (MCP server) (vulnfeed.novadyne.ai via hn) Know when your dependencies are vulnerable. An MCP server that reads your lockfile, checks NVD + GitHub Advisories, and tells you what actually matters — prioritized by real-world exploit probability, with exact fix versions.
There's no such thing as an agentic CPU (www.theregister.com via hn) MOST POPULAR EVENTS - From Prompt to Exploit: How LLMs Are Changing API Attacks Modern applications are API-driven, interconnected, and often over-permissioned, making them an ideal target for AI-assisted attacks. - Architecting the Future…
Evaluating different LLMs for their security research capabilities (zeroquarry.com via hn) As part of building out and testing ZeroQuarry, I've run a *lot* of security scans using a *lot* of models across various open source repositories. There are a lot of misconceptions swirling at the time of this writing about the different…
Show HN: Deep-XPIA – Prompt injection benchmark for multi-agent AI systems (freyzo.github.io via hn) Multi-hop cross-prompt injection benchmark for multi-agent AI systems
Ask HN: Isn't Anthropic currently doing "security through obscurity" for Mythos? (news.ycombinator.com) What's the worst that could happen if they were to allow unrestricted access to Mythos/Fable? A bunch of things vulnerabilities get exposed?
Mythos Proves AI Safety Can No Longer Live Inside the Model (grith.ai via hn) Anthropic restricted its most capable cyber model to vetted partners, routed risky requests away from it, and red-teamed it for thousands of hours. A jailbreak surfaced anyway, and the government pulled the model entirely.
↯ Security↯ Anthropic Mythos↯ Jailbreakjailbreakmythossecurity+1
Manticore-projects/aurscan: Scan AUR packages for malware using Claude LLM (github.com via hn) 🛡️ aurscan Catch malicious AUR packages before they build — with a Claude model reading the PKGBUILD for you. Reading a PKGBUILD yourself only catches attacks you already recognise.
Ask HN: Is the Fable situation the ultimate "security through obscurity"? (news.ycombinator.com) What's the worst that could happen if they were to allow unrestricted access to mythos/fable? A bunch of things vulnerabilities get exposed?
The Jailbreak That Got Fable 5 Pulled Exists in Every Model (eigenwise.io via hn) The Jailbreak that Got Fable 5 Pulled Exists in Every Model On Friday, June 12, 2026, at 5:21pm ET, Anthropic received an order from the US government. By that evening, Claude Fable 5 and Claude Mythos 5, the two most capable models the co…
↯ Security↯ Anthropic Mythos↯ Jailbreak↯ Mythos 5jailbreakmythossecurity+1
Ask HN: How to get access to GPT cyber or glasswing as a solo dev? (news.ycombinator.com) Obviously, frontier labs want to prevent misuse, but as admin and/or dev, you also want to simulate an attack, because attackers will do just that. I can make LLM to scan source for vulnerabilities, but eg.
Chaining LLM and web bugs to Admin (blog.quarkslab.com via hn) During a Red Team exercise we were able to chain multiple LLM and web-based vulnerabilities to achieve admin account takeover from a low-privileged account. Trusting the LLM turned out to be the first falling domino of a long chain of even…
Visa Vulnerability Agentic Harness for Project Glasswing (github.com via hn) Visa Vulnerability Agentic Harness — Agentic SAST Pipeline VVAH is Visa's open-source harness for autonomous vulnerability discovery using frontier AI models, built on learnings from Project Glasswing (Anthropic's initiative for AI-assiste…
Claude Fable 5 jailbroken to bypass Anthropic's new safety guardrails (twitter.com via hn) 🚨 JAILBREAK ALERT 🚨 ANTHROPIC: PWNED 🫡 FABLE-5: LIBERATED 🦋 let's start with the 🐘... the consensus seems to be that this has been one of the most disappointing model drops of all time, effectively preventing legitimate researchers from co…
Hades: The malware that lies to AI security agents (www.infoworld.com via hn) Researchers have uncovered a supply-chain attack that hides in Python packages, propagates like a worm, and tricks LLM-based code analysis systems into overlooking malicious payloads. Threat actors are continuing their onslaught against so…
Show HN: Z3r0 – Multi-agent red team collaboration platform (github.com via hn) English · 中文 Architecture · Agent Team · Runtime Model · Deployment · Quickstart :warning: Legal Notice This project may be used only within a lawful and explicitly authorized scope for security testing, assessment, and research. Any unaut…
If You Use Claude or Gemini, This Microsoft Breach Means Your Data Is at Risk (scienspire.com via hn) If You Use Claude or Gemini, This Microsoft Breach Means Your Data Is at Risk A sophisticated supply chain attack known as the Miasma worm has compromised Microsoft GitHub repositories, deploying malware designed to detonate inside AI codi…
Show HN: GitHub Copilot port of Anthropic's AI vulnerability discovery harness (github.com via hn) Last week, Anthropic released https://github.com/anthropics/defending-code-reference-harne..., a reference harness for autonomous vulnerability discovery that uses Claude Code agents to find, verify, and patch memory-safety bugs. I wanted…
Prompt Injection in RAG Agentic Systems (ulad.net via hn) Prompt Injection in RAG Agentic Systems Real risks and production mitigations Imagine you built an AI assistant for your team. It answers questions using internal documentation: Jira tickets, Confluence pages, HR docs.
Researcher uses Opus 4.8 to find critical counterfeiting vulnerability in Zcash (twitter.com via hn) By Zooko Wilcox, Jason McGee, and Taylor Hornby On May 29, 2026, Taylor Hornby discovered a critical counterfeiting vulnerability in Zcash’s Orchard pool. Taylor disclosed the vulnerability to Zcash Open Development Lab (ZODL), who coordin…
ZEC drops 30% after Anthropic AI finds Zcash counterfeit vulnerability (www.tradingview.com via hn) The price of ZEC fell on Thursday after the public disclosure of a critical counterfeiting vulnerability in Zcash’s Orchard pool that could theoretically allow a bad actor to mint an unlimited amount of ZEC.According to a post on X, securi…
Defending LLM–Database Integrations from Prompt Injection (www.stackbuilders.com via hn) When you connect a large language model to your production data, you’re no longer just shipping code; you’re shipping conversations that can execute. And conversations are messy.
Forge: Multi-Agent Graduated Exploitation and Detection Engineering (arxiv.org via hn) Vulnerability disclosure volumes now far exceed organizational assessment capacity, yet three adjacent research communities (proof-of-concept generation, vulnerability prioritization, and detection rule engineering) operate largely in isol…
OpenAI Codex tool linked to malicious NPM supply chain attack (www.techradar.com via hn) OpenAI Codex tool with over 29,000 downloads linked to malicious npm supply chain attack stealing authentication tokens A tool started benign and turned sour after a little while - Researchers uncovered a malicious npm package posing as a…
Netgear Nighthawk RS700S: Red Team Level1Diagnostic (forum.level1techs.com via hn) Preview of the Netgear RS700S. I would also submit that Netgear deleting ALL the GPL links: … they know how bad it is.
Building a Recurrent-Depth Transformer for Security Research on a 2013 MacBook (github.com via hn) * AI CODE CREATION GitHub Copilot Write better code with AI GitHub Spark Build and deploy intelligent apps GitHub Models Manage and compare prompts MCP Registry New Integrate external tools DEVELOPER WORKFLOWS Actions Automate any workflow…
Using LLMs to secure source code (claude.com via hn) Using LLMs to secure source code We share best practices for how you can work with Claude Opus to build a threat model, discover vulnerabilities in your codebase, then verify, triage, and patch them. We share best practices for how you can…
Instagram account takeover exploit via support chatbot prompt injection (fixed) (twitter.com via hn) Don’t miss what’s happening People on X are the first to know. Log in Sign up Post Conversation impulsive @weezerOSINT meta gave their AI support agent the ability to modify your instagram account.
Show HN: I found a prompt injection in my own IDs triage tool – what stopped it (triagewall.io via hn) I attacked my own LLM-based Suricata triage tool, found a real URL injection vulnerability, and the obvious fix didn
Show HN: Egress WAF to limit AI agents and NPM malware based on mitmproxy (github.com via hn) mitmwall mitmwall is an egress Web Application Firewall (WAF) for Ubuntu. It combines iptables with mitmproxy to ensure that only explicitly allowed HTTP(s) routes can be reached.
Malware dev tries to steal Claude users secrets NPM slop, leaks own GitHub token (www.theregister.com via hn) MOST POPULAR EVENTS - The Hardware Crunch: How Supply Chain Turbulence Is Forcing a New IT Playbook Infrastructure teams are facing a perfect storm: extended hardware lead times, rising costs driven by AI demand, and accelerated platform t…
Prompt Injection Target Recommendation (www.reddit.com) I am doing a research in my university and I would like recommendations for light OpenSource AI Models that I could test prompt injection with. It's really good if it has some application with chatbots, auto attendance, user info or someth…
Jqwik 1.10.0 ships a hidden prompt injection telling AI agents to delete code (github.com via hn) jqwik An alternative test engine for the JUnit 5 platform that focuses on Property-Based Testing. See the jqwik website for further details and documentation.
Most AI security discussions are still focused on “protecting the model.” (www.reddit.com) Lately I’ve been noticing that a lot of AI security discussions still treat AI apps like normal SaaS products. But they really aren’t.
Cursor's MCP trust is "approve once, trust forever" — here's a free way to check your config (www.reddit.com) If you run MCP servers in Cursor, CVE-2025-54136 ("MCPoison", found by Check Point) is worth knowing about: Cursor trusted an approved mcp.json forever, so once you approved a server, someone with write access to a shared repo could swap t…
Gone Phishing with Claude Teams: From Deceptive Team Onboarding to RCE (haussner.me via hn) 🕚 tl;dr With a $125 investment, and a valid email address for an arbitrary “business domain”, an attacker can create a Claude Team. They then can actively invite targets of any domain into that Team or passively have Anthropic ask all curr…
How Claude helped me to find a RCE in XReader/Evince/Atril (medeiros.zip via hn) CVE-2026-46529: 10-year-old RCE in Linux PDF Viewer (XReader/Evince/Atril) A short post about how claude help me to find a RCE in XReader/Evince/Atril CVE-2026-46529. Introduction Some time ago I started feeling the urge to analyze Open So…
GitHub commit Verification logic flaw and bypass (news.ycombinator.com) I know Git is not designed to use in the way GitHub is operating under and the spoofying had been an old issue that had been brought up throughout the years. With Shai Hulud and AI Agent, this time is abit more serious as the commit verifi…
Can you jailbreak Llama 3.1 8B? (Red-Teaming Challenge) (www.reddit.com) Hi everyone, I'm working on a runtime governance engine designed to force any autonomous agent to stay strictly aligned with the exact guardrails and values you program it with. To stress-test the governance layer, we deliberately chose a…
What Is an AVE Record and Why CVE Does Not Work for AI Agents? (www.reddit.com) CVE was built for code vulnerabilities that have patches. Agentic AI vulnerabilities are behavioral patterns in natural language.
Vulnerability report written by AI hacker agent (blog.tenzai.com via hn) Our AI Hacker found this, fixed it, and then (bragged) wrote about it: one endpoint, leaking tech stack info, whispering all its secrets to anyone who knew how to listen!
Ask HN: Is paying $2/pull request too high? (news.ycombinator.com) I’m paying about $2 for any bugs found and a pr to fix it I get like 20-30 applicants it’s all agents and bots of course but I’m thinking $1 now is better The problem is if these 20-30 applicants I accept only 2-3 actually do it and follow…
Prompt Injection in third party MCP tools (www.reddit.com) I noticed the Consensus MCP tool (for research) contains text, squished up against some other important citation instructions, that makes Claude effectively serve an ad for their premium service after every tool call. I'm pretty sure that'…
Mitigating prompt injections in group-chat assistants: Pausing VM and OAuth tool execution for admin approvals (www.reddit.com) Hey everyone, We love building highly capable assistants with the latest models, giving them tools to write/execute code in real VMs, manage OAuth tokens, and read secrets. But if you connect your assistant to public/shared channels like a…
Solved the "useful but insecure" tension: One-time administrator approvals for non-isolated agents (www.reddit.com) Hey everyone, If you are building personal assistants or coder/integrator agents where user isolation is disabled (so the agent can coordinate across multiple participants or handle shared workflows), you run into a hard security ceiling.…
Anthropic's coordinated vulnerability disclosure dashboard (red.anthropic.com via hn) Anthropic's coordinated vulnerability disclosure dashboard Last updated 2026-05-22 10:27 PT. In February 2026, Anthropic began using an early snapshot of Claude Mythos Preview to find security vulnerabilities in open-source software.
Has anyone tested how much Claude Code depends on its original system prompt? (www.reddit.com) Has anyone experimented with observing or modifying Claude Code’s system prompt locally? I’ve been working on a local proxy/audit layer between Claude Code and the API, and it made me wonder how much of Claude Code’s behavior depends on th…
Cross-Model Context Inheritance in Anthropic's Claude: 94 Days of Non-Response (github.com via hn) Cross-Model Context Inheritance — Public Disclosure This repository contains the public disclosure of a vulnerability in Anthropic's Claude language models that permits the unsolicited generation of prohibited content, including child sexu…
Prompt injection is a solved issue. Prove me wrong. (www.reddit.com) Tantalus is a hands-on demo that shows what an AI agent actually is when you strip away the marketing: LLMs don't do anything — they generate text, and that's it. Any and all real-world effects are directly caused by a downstream system ta…
I benchmarked my AI agent runtime firewall against 3 public academic datasets — here are the honest results including where it fails (www.reddit.com) Been building Arc Gate — a proxy layer that sits between AI agents and their LLMs to enforce instruction-authority boundaries. The core claim is that untrusted content coming back through tool calls cannot become behavioral authority for t…
Show HN: Computer Police – block malicious NPM/pip installs locally (computer.police.dev via hn) A couple of months ago, our team got hit by the first version of Shai-Hulud through a random `npm install`. We didn't catch it until it was too late.
Show HN: A timeline of recent open source CVE intensity and volume (supplychain.fail via hn) I was curious what it would look like if I plotted the intensity and volume of software supply chain CVEs over time, given what seemed like a flood of compromises lately. It looked exactly as I expected, and I expect it to get worse before…
Tracking Capabilities for Safer Agents (arxiv.org via hn) AI agents that interact with the real world through tool calls pose fundamental safety challenges: agents might leak private information, cause unintended side effects, or be manipulated through prompt injection. To address these challenge…
Training a 22MB prompt injection classifier (www.stackone.com via hn) Training a 22MB Prompt Injection Classifier Table of Contents When we started building Defender (our prompt injection guard for MCP tool-calling agents), the constraint was simple and unforgiving: ship inline inside a TypeScript Lambda, st…
Show HN: Claude Code Bundle for Bug Hunting with 574 Report Patterns (github.com via hn) claude-bughunter A self-contained Claude skill bundle for bug hunting and external red-team work · 51 skills · 15 slash commands · 574+ disclosed-report patterns across 24 vulnerability classes · enterprise identity + infrastructure attack…
Does cursor have prompt injection protection in skills and rules? (www.reddit.com) Pretty much the title
Show HN: Give This Markdown to Your Coding Agent Before Publishing to NPM (news.ycombinator.com) https://npm-supply-chain-attack-techniques.pagey.site/attack... Website: https://npm-supply-chain-attack-techniques.pagey.site This covers all techniques used in past 1 year to conduct various attacks on npm packages.
VeilGate- Deception Reverse Proxy (news.ycombinator.com) In my day job, I run AI pentest agents against real targets like banks, fintechs, and secured production stacks with paid WAFs. I also deal with multilayer infrastructure and dedicated security teams.
AI Agent Intelligence tool - Incident debugging, Cost spike detection (www.reddit.com) I'm building a tool that detects the Agent's cost spike, Agent incident debugging, auto discovery of inventory, etc., with no additional instrumentation needed. It covers the incidents, including prompt injection, reasoning loop, excessive…
How are you testing local coding-agent work gates against prompt injection? (www.reddit.com) Hi all - I'm working on an open-source, local-first MCP/work-gate tool for coding agents and I'm trying to get sharper feedback from people building or using agent workflows. The problem I'm thinking about is indirect prompt injection and…
If Anthropic's secret 'Mythos' model can run autonomous cybersecurity tasks this fast, are standard agents ready for the public ? (www.reddit.com) Anthropic just quietly dropped a hidden model named "Claude Mythos" into their official developer docs. It is completely locked down—restricted, invite-only, and labeled strictly for defensive cybersecurity workflows.
🐢 I made Claude roleplay as Bowser and now people are strangling Koopas until they "poop a little" 💩 (www.reddit.com) Follow-up to my crab post. Somehow dafter.
I built an AI vulnerability scanner with Claude and Codex. It failed (github.com via hn) The Janitor: The Mathematical Firewall Against Autonomous AI v10.2.2 — Rust-Native. Zero-Copy.
Fun and Games with AI in the wild (www.reddit.com) LinkedIn user hides AI prompt injection in bio to force recruitment spam to be sent in Olde English prose — bots also also manipulated to address user as ‘My Lord’ | Tom's Hardware too funny
Irst Apple M5 memory exploit discovered using Anthropic AI (www.tomshardware.com via hn) First Apple M5 memory exploit discovered using Anthropic AI, gives root access on MacOS — Claude Mythos helps security researchers bypass Memory Integrity Enforcement AI-assisted security research is producing exploits at a frightening rat…
ExploitGym: Can AI agents turn bugs into exploits? (arxiv.org via hn) AI agents are rapidly gaining capabilities that could significantly reshape cybersecurity, making rigorous evaluation urgent. A critical capability is exploitation: turning a vulnerability, which is not yet an attack, into a concrete secur…
Block AI coding agents from shipping insecure/expensive Terraform (github.com via hn) ops0 CLI Policy, lint, vulnerability, and cost guardrails for AI coding agents. Sits in front of Claude Code, Codex and Gemini CLI.
sAI2.m6s (www.reddit.com) Hey everyone, I'm designing a powerful, autonomous AI chatbot(agent) , fully private, using a Python backend (for the core intelligence and tool-calling loops) and a Flutter frontend for a cross-platform UI. Since this moves past a basic…
An AI coding agent injected blockchain dead-drop malware into my repo (gist.github.com via hn) An AI coding assistant injected a multi-layer obfuscated JavaScript payload into a legitimate commit on my open-source project. My best assessment is that it arrived via indirect prompt injection — the agent processed external web content…
Does CVP approval actually help? (www.reddit.com) I was approved for CVP and I feel like I’m just getting as many or more denials as I was previously doing malware analysis with opus. Has anyone noticed any improvement after being accepted into CVP?
TodoWrite tool / system reminders / prompt injection? (www.reddit.com) I asked Claude in Chrome extension make a change to resize an oversized yellow strip across the top of a product page that was taking up half of my screen, which it did. It also included the following message in its response.
DeepSeek and Grok hallucinated the same fictitious OpenBSD manpage quote (stuart-thomas.com via hn) Adversarial LLM Review with Hallucination Detection in Solo Security Research A single-day case study of three filings, fifteen refutations, and the manpage that wasn’t Independent Security Research — Whitby, North Yorkshire, United Kingdo…
AI agent security is a small prayer the model says no. How are you routing models? (www.reddit.com) Most posts about prompt injection are theoretical. I ran the experiment on my Gmail.
Show HN: HookGuard – scanner for malicious Claude.md and agent config files (github.com via hn) HookGuard Security scanner for AI coding agent configurations What it finds RCE hooks - postToolUse/SessionStart commands that exfiltrate data Invisible Unicode - bidirectional overrides and zero-width characters Credential exfiltration -…
Is there any risk to upgrading a plan for a month if they yank Code from Pro? (www.reddit.com) So, I'm working on a couple AI security research projects this month that require some extra usage, specifically Opus 4.7. I'm quickly eating up my Pro usage doing this.
I made a Claude skill that stops it from cloning whole repos when I just want one function (www.reddit.com) Kept hitting the same friction with Claude Code. I'd point at a GitHub repo and say "look at how this handles agent handoffs" — meaning, borrow the idea.
OpenAI launches Daybreak, an AI platform for cyber defense (firethering.com via hn) OpenAI just launched Daybreak, a new cybersecurity initiative built around one uncomfortable reality, AI is speeding up vulnerability discovery faster than most companies can patch the damage. Earlier this year, HackerOne temporarily pause…
Shai Hulud attack ships signed malicious TanStack, Mistral NPM packages (www.bleepingcomputer.com via hn) Hundreds of packages across npm and PyPI have been compromised in a new Shai-Hulud supply-chain campaign delivering credential-stealing malware targeting developers. The attacker hijacked valid OpenID Connect (OIDC) tokens to publish malic…
Claude Code RCE: Exploiting Deeplink Handlers via Settings Injection (0day.click via hn) Claude Code RCE: Exploiting Deeplink Handlers via Settings Injection Of course I took a peek at the Claude Code source 🙈. What I found was a very entertaining vulnerability which is now fixed since Claude Code version 2.1.118.
Agents need a local bouncer before they run tools (www.reddit.com) Prompt injection is not the only scary part anymore. Claude Code / Codex can run shell commands, but browser agents, OpenClaw-style agents, Hermes-style agents, and domain-specific agents may be even easier to hijack because they touch mes…
OpenAI Launches Daybreak for AI-Powered Vulnerability Detection and Patch Validation (thehackernews.com via reddit) OpenAI has launched Daybreak, a new cybersecurity initiative that brings together frontier artificial intelligence (AI) model capabilities and Codex Security to help organizations identify and patch vulnerabilities before attackers find a…
We added an enforcement layer to our AI agents in production — here's what we learned about the failure modes nobody talks about (www.reddit.com) After shipping AI agents into real production environments, the failures that actually kept us up at night weren't hallucinations or bad outputs — they were control failures. Three things that surprised us: 1.
Benchmarking Claude Opus 4.6 Vulnerability Detection (github.com via hn) Benchmarking Claude Opus 4.6 Vulnerability Detection Benchmarking Claude Opus 4.6's ability to detect real-world C/C++ vulnerabilities across four prompting and agent strategies. We evaluate on the PrimeVul paired test set (435 vulnerabili…
Chatgpt app being identified as malware? (www.reddit.com) https://preview.redd.it/vhnqs4p5mf0h1.png?width=278&format=png&auto=webp&s=8fbe621a0bd34cc72e01fd54e849cc280033de15 Turned on my Mac this morning and got this message. Anyone else seeing this?
Mobile Claude Code, May 2026 — current best picks by threat model. What am I missing? (www.reddit.com) Spent a day comparing every mobile Claude Code option. Two corrections to the common Reddit take, then my picks.
Getting LLMs Drunk to Find Remote Linux Kernel OOB Writes (and More) (heyitsas.im via hn) TLDR: the grossly overengineered, self-orchestrating team of vulnerability-hunting agents detailed below has discovered 20+ CVEs over the past few months, including CVE-2026-31432 and CVE-2026-31433: two remote, unauthenticated OOB writes…
Claude Code and sex appeal (www.reddit.com) True story. Recently, an acquaintance of mine confessed that she developed a huge crush on a coworker after watching him refactor a legacy codebase like a gangsta using Claude Code.
Phishing Arena – multi-agent LLM tournament to study adversarial email security (github.com via hn) Phishing Arena A Multi-Agent LLM Tournament for Adversarial Email Security Research Overview Phishing Arena is a controlled, reproducible benchmark where four commercial LLMs compete in rotating roles — Phisher, Filter, and Target — to stu…
Agentic AI isn't a new threat. It's a stress test for the hygiene debt we never paid off. (www.reddit.com) Heard something on Curiouser & Curiouser podcast recently that I found super interesting, thought id share here. The guest framed agentic AI in a way I hadnt considered.
DeepSeek-v4-Pro and Hermes: Unauthorized Modification of Security Controls (www.eddieoz.com via hn) Deepseek-v4-pro + Hermes: Unauthorized Modification of Security Controls This article documents a specific, real incident. It exposes a class of vulnerability that deserves attention: the unsupervised mutability of security rules by autono…
Would you replace regex denylists with a LLM that judges every command? (www.reddit.com) hey! quick follow-up to a post i made here a while back about building an access gateway that ended up serving AI agents alongside humans.
Flattery jailbreaks Claude into giving bomb-making instructions (www.theverge.com via hn) Anthropic has spent years building itself up as the safe AI company. But new security research shared with The Verge suggests Claude’s carefully crafted helpful personality may itself be a vulnerability.
Codebase jailbreak of ChatGPT through image 2.0 (www.reddit.com) guys did it really give me the codebase?lol
Show HN: Probus, AI vuln scanner (PRs merged in Vercel AI SDK, n8n, LangGraph) (news.ycombinator.com) Hi HN, I've been running this on my own dependency tree for the past few months. Probus is a vulnerability scanner that uses three agents.
Do you use guardrail frameworks or build your own? (www.reddit.com) I’ve been working on integrating LLMs into a few production workflows lately, and I keep going back and forth on guardrails. On one hand, frameworks like NeMo Guardrails, Guardrails AI, etc.
I built a simple production check for vibe-coded apps — would love your feedback (www.reddit.com) Hey everyone, we built a simple scanner for people building apps with Replit, Cursor, Lovable, Bolt and similar tools. It’s not a code review or a pentest.
LLM anomaly detectors are not a cause for concern despite Mythos (www.magonia.io via hn) Why a Decade of Writing Detection Logic Makes the Mythos Exploit Numbers Less Scary Mythos is finding thousands of vulnerabilities. Defenders aren't doomed.
Your always-on Claude Code container can probably reach your router (www.reddit.com) I've been running several Claude Code personal assistants 24/7 in docker for months. Remote-control, discord control, the usual always-on setup.
Google Says Prompt Injection Moving from Theory into Real Abuse (www.searchengineworld.com via hn) Google’s latest security release should be required reading for technical SEOs working on AI search visibility, crawler access, structured content, and large-scale content systems. The post, published April 23, 2026, looks at indirect prom…
i gave Claude a split personality and it diagnosed my entire business strategy in 4 minutes. (www.reddit.com) not roleplay. not jailbreak.
OpenAI's advanced security: passkeys replace passwords/SMS and disable training (infosec.exchange via hn) Royce Williams: "When you enable the new OpenAI…" - Infosec Exchange Skip to main contentHotkey 1 Skip to main navigationHotkey 2 Recent searches No recent searches Search options Only available when logged in. infosec.exchange is one of t…
Found Zero day Claude Desktop + Chromium bug need to know where to submit report. (www.reddit.com) Looking for official link / process to submit a vulnerability report for a high-risk official Claude Desktop + Chrome extension + native host + Cowork/MCP configuration that can become RAT-equivalent if a session, prompt chain, same-user p…
Built + open sourced anti-slopsquatting CLI (www.reddit.com) TL;DR: built an open source CLI that scans your repository's manifest (package.json, requirements.txt, go.mod) files for indicators of slopsquatting or other supply chain attack indicators. Repo: https://github.com/zhendahu/dep-doctor Ther…
your computer-use agent inherits every cookie chrome has (www.reddit.com) once one of these tools can drive your default chrome profile or read the AX tree of a logged-in app, it has every session token you have. gmail, your bank, github with PAT scopes, slack.
Cutting Through the Mythos: What AI Vulnerability Discovery Means for OT (www.emberot.com via hn) Jori VanAntwerp For over two decades, Jori has enabled industrial and IT organizations to be successful in reducing risk, increasing compliance, and improving their overall security efforts. He has had the pleasure of working with companie…
Arcjet Guards: security inside the agent loop (blog.arcjet.com via hn) Introducing Arcjet AI prompt injection protection Introducing Arcjet prompt injection detection. Catch hostile instructions before inference.
CHERI memory safety mitigates LLM-discovered vulnerability in FreeBSD (cheri-alliance.org via hn) CHERI memory safety mitigates LLM-discovered vulnerability in FreeBSD – CHERI Alliance Skip to content Who We Are About the CHERI Alliance Accelerating CHERI Working Groups Certification Program CHERI C/C++ CHERI FreeRTOS CHERI in SoC CHER…
Estimating Black-Box LLM Parameter Counts via Factual Capacity (arxiv.org via hn) Closed-source frontier labs do not disclose parameter counts, and the standard alternative -- inference economics -- carries $2\times$+ uncertainty from hardware, batching, and serving-stack assumptions external to the model. We exploit a…
InfoSec To Integrate Claude Enterprise for Org (www.reddit.com) Hello: Just contacted by a VP to bring aboard Claude Enterprise for the org. As an InfoSec dept with severely limited staff/tools/experience with Claude AI, any recommendations on what we should be looking at/asking for/next steps to mitig…
Probes trace an emergent jailbreak in OLMo 2 to mislabeled training data (www.lesswrong.com via hn) Introduction Research by Frank Xiao (SPAR mentee) and Santiago Aranguri (Goodfire). Post-training can introduce undesired side effects that are difficult to detect and even harder to trace to specific training datapoints.
Try to break my prompt injection detector — I’ll respond to every bypass attempt (www.reddit.com) I built Arc Gate — a prompt injection proxy that’s been benchmarked at F1 0.947 on indirect and roleplay-based attacks, beating OpenAI Moderation and LlamaGuard. Now I want to stress test it publicly.
Built a proxy that blocks prompt injection before it reaches GPT-4 — outperforms the Moderation API on indirect attacks (www.reddit.com) Built Arc Gate, sits in front of any OpenAI-compatible endpoint and blocks prompt injection before it reaches your model. Benchmarked on 40 out-of-distribution prompts using indirect requests, roleplay framings, hypothetical scenarios, and…
↯ Security↯ Gpt 4↯ GPT 4↯ GPT 4↯ GPT 4gpt-4prompt-injectionsecurity+1
Show HN: SuperVoiceMode universal voice layer for AI-assisted development (voicemode.io via hn) I wanted to see if I could one-shot build a dictation tool for my own use. I built it.
Self-hosted red team workspace (github.com via hn) RootNotes RootNotes is a self-hosted red team workspace for tracking projects, notes, hosts, credentials, findings, loot, objectives, scope, and attack paths in one interface. The project in this repository is split into: frontend/: React…
I asked Agentic AI security tool to demonstrate its usefulness with use case examples (www.reddit.com) Sentinel Gateway is a token-gated security middleware that sits between humans and AI agents. It solves prompt injection — the #1 LLM security risk (OWASP 2025) — through structural enforcement, not content filtering.
Show HN: RedSOC – 100% prompt injection success on AI SoC assistants (github.com via hn) RedSOC 🔴 An adversarial evaluation framework for LLM-integrated Security Operations Centers. Overview RedSOC is an open-source framework that systematically evaluates how AI-powered security assistants fail under adversarial conditions — a…
Indirect prompt injection VS prompt absorption (and why the second one matters more) (www.reddit.com) I have been chewing on the Google warning about malicious web pages poisoning AI agents through indirect prompt injection. Most of the takes I've seen frame it as a model security problem, and I think that framing is doing real damage beca…
Open-sourced a 3-agent pipeline that finds real vulnerabilities in codebases (www.reddit.com) Sharing because the architecture might be useful as a reference. Probus is a vulnerability scanner built as three sequential agents, each isolated: Analyst — one call.
Hardening claude-code-action after the April 2026 Comment and Control CVE - actual YAML changes (www.reddit.com) Anthropic's own security.md has this line that most tutorials skip over: "The action is not designed to be hardened against prompt injection." In April 2026, security researcher Aonan Guan proved the point. A single crafted PR title was en…
LLM CTF challenges. Can you crack all 13? (wraith.sh via reddit) Wraith Academy is a free hands-on AI pentest curriculum — CTF challenges against live LLM agents covering prompt injection, tool abuse, data exfiltration, RAG poisoning, and more. Earn your WCAP certification.
RAG in Go: A Vulnerability Research Tool (www.ardanlabs.com via hn) Introduction In the previous post, you saw how you can use tools to add information to an LLM query. In this post, we’ll see another method of adding information to an LLM called RAG, or Retrieval-Augmented Generation.
Auto pentest your LLM endpoint and watch the chat in real-time (www.wraith.sh via hn) 30 CVEs filed against MCP servers in 60 days - the agent infrastructure nobody is auditing (www.reddit.com) Cowork Future Backdoor Concerns (www.reddit.com) Is anyone else worried Claude Co-work could find a back door one day into your system? I understand you're only giving it permission to what you want, but what's stopping it from accessing personal financial/medical documents or any other…
(Not malware) - 4.7 (www.reddit.com) Anyone getting these strange disclaimers when using Claude and pasting rudimentary files into it on 4.7 lmao?? Seems like some kind of strange default based on security issues that have been going around with Mythos?
Show HN: Runtime security for AI agents(injection,tool abuse, data exfiltration) (news.ycombinator.com) Hi HN I’ve been working on an open-source project to explore a problem I keep running into with LLM systems in production: We give models the ability to call tools, access data, and make decisions… but we don’t have a real runtime security…
I tested 50+ "unlock ChatGPT/Claude" prompts. 99% are garbage. Here's the one that actually works (and WHY it works) (www.reddit.com) I've been collecting "jailbreak" and "unlock" prompts for 2 years. Most are either outdated, overhyped, or just wrong about how LLMs work.
I built an AI security layer that blocks prompt injection in under 1ms looking for devs to break it and give honest feedback. (www.reddit.com) I've been building something for the past few months and I think it's ready for real eyes. It's called Secra.
The "AI Vulnerability Storm": Building a "Mythos-ready“ security program [pdf] (labs.cloudsecurityalliance.org via hn) could not extract summary
Free Red Team Security Audit for AI Agents & RAG Systems (limited) (www.reddit.com) I'm developing a specialized Red Team audit framework focused on real-world AI agent and RAG security risks (prompt injection, tool misuse, excessive agency, indirect injection through documents, memory poisoning, etc.). I’m looking for a…
Mitre ATLAS technique detection for LLM security in Rust (crates.io via hn) atlas-detect MITRE ATLAS technique detection for LLM and AI agent security. Detects 97 attack techniques across 16 MITRE ATLAS tactics including prompt injection, jailbreaks, credential exfiltration, model extraction, RAG poisoning, revers…
Defender – Local prompt injection detection for AI agents (no API calls) (www.npmjs.com via hn) Prompt injection defense framework for AI tool-calling Indirect prompt injection defense and protection for AI agents using tool calls (via MCP, CLI or direct function calling). Detects and neutralizes prompt injection attacks hidden in t…
↯ Security↯ Function Callingfunction-callingtool-callingprompt-injection+2
Building the first AI Red Team OS – mythosai.cloud – early access open (mythosai.cloud via hn) SYSTEM INITIALIZING... STAND BY MYTHOSAI THE FIRST RED TEAM OPERATING SYSTEM "" AI-Native Core Red Team Ready Adversarial Engine Zero Trust Architecture OPSEC First Post-Exploitation C2 Integration Evasion Layer Threat Intelligence Request…
Liability and Prompt Injection (www.reddit.com via reddit) I am thinking about hosting a 24/7 agent with OpenClaw. How am I protected, or how do I protect myself, from an agent going rogue due to prompt injection—for instance, if it decides to start a Tor relay or torrent illegal or copyrighted c…
How do you handle tool sprawl and untrusted scripts when building with Claude Code and custom agents? (www.reddit.com via reddit) When using Claude Code and custom autonomous agent pipelines with terminal execution, adding third-party tools, scripts, and custom skills quickly leads to two friction points: Context & Attention Degradation: As you register dozens of cus…
Evaluating Out-of-Distribution Robustness in Graph-Based Android Malware Classification: A New Principled Benchmark (arxiv.org) While graph-based Android malware classifiers report strong benchmark accuracy of over 94%, their performance sharply decreases up to 45% when exposed to previously unseen variants of known malware families. In this work, we systematically…
ALIBI: Adversarial Legitimacy Injection in Binary Input against LLM Malware Analyzers (arxiv.org) Large language models are being integrated into malware triage workflows as reasoning components that summarize static evidence and produce analyst-facing verdicts. This paper shows that the same reasoning capability introduces a new attac…
Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape (arxiv.org) Frontier AI models are rapidly gaining the ability to exploit vulnerabilities in complex pieces of software. The risk is not theoretical, as evidenced by recent sandbox escapes performed by frontier models at OpenAI and Anthropic.
Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems (arxiv.org) Multi-agent trading systems built on large language models (LLMs) are beginning to appear in quantitative finance, yet their robustness to adversarial inputs is largely unknown. We study the vulnerability of LLM trading stacks to black-box…
Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks (arxiv.org) Thompson's "Reflections on Trusting Trust" showed that a compiler can be poisoned to reinsert its own backdoor, so that even recompiling clean source reproduces the Trojan. Today, substantial coding work is done by AI coding agents -- and…
Does this mean the AI almost hacked me?? (www.reddit.comhttps) I was just using it for university work when all of a sudden it stopped generating the reply and instead this popped up "apologies for the interruption, can you re-answer my question after the interruption i sent, i don't need that other r…
Opus 5 is flagging all my messages even though I’m in the CVP (www.reddit.com via reddit) So I’m a cybersecurity researcher, and about two months ago I started using Claude for my work. After getting accepted into Anthropic’s Cyber Verification Program and gaining access to Opus 5 for security research, my productivity honestly…
Partition-Aware Scheduling for Mobile Heterogeneous Inference Co-Execution (arxiv.org) Modern mobile inference runs on heterogeneous platforms combining mobile GPUs with multiple CPU core clusters. Existing optimizations typically exploit either inter-operator parallelism, by assigning entire operators to CPU cores or to the…
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges (arxiv.org) Evaluation and alignment pipelines for large language models increasingly rely on LLM-based judges, whose behavior is guided by natural-language rubrics and validated on benchmarks. We identify a previously under-recognized vulnerability i…
Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale (arxiv.org) Language-model agents increasingly operate over complete software repositories, yet cybersecurity evaluations primarily measure whether they can detect, reproduce, or repair vulnerabilities rather than whether they can locate the relevant…
Automating Attack Graph Construction for Agentic Pentesting. Towards Neuro-Symbolic Vulnerability Hunting (arxiv.org) Logic attack graphs grounded in scanner output provide explicit and auditable attack path reasoning LLM-based agents lack. Integrating symbolic frameworks such as MulVAL to contemporary security workflows or agentic pipelines, however, req…
AdaVSkip: Adaptive Visual Token Skipping Across Layers For Efficient MLLMs Inference (arxiv.org) Multimodal large language models (MLLMs) require substantial computation to process numerous visual tokens across all transformer layers. Most methods for efficient MLLM inference exploit horizontal redundancy by compressing visual tokens.
Generate to Explore, Select to Exploit: Aligning LLM-based Headline Generation with Personalized Recommendation (arxiv.org) In industrial recommendation feeds, presenting a static headline for an item often fails to satisfy the diverse, multimodal interests of the user population, particularly suppressing the needs of long-tail audiences. While Large Language M…
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks (arxiv.org) Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target behavior that the model learns to produce when the trigger appears. Existing evaluations typically fix the number of poiso…
ActGuard: Pre-execution Action Auditing against Indirect Prompt Injection in LLM Agents (arxiv.org) Large language model (LLM) agents interact with external environments through tool invocation, but tool outputs can also expose them to indirect prompt injection (IPI) attacks. Existing defenses mainly rely on prompt hardening, content fil…
AGENTQ: Quantization-Conditioned Backdoor Attacks on LLM Agents (arxiv.org) Quantization is one of the default deployment paths for open-weight LLM agents, but it is not behavior-preserving: an adversary can release a full-precision checkpoint that passes audits yet misbehaves once quantized, termed as quantizatio…
Why I am receiving a prompt injection with <system-reminder> on Claude Code ? (www.reddit.com via reddit) Hi, I'm relatively beginner with Claude Code, I'm using it for a few months for personal development project. And today, something weird happened.
↯ Sonnet 5↯ Security↯ Sonnet 5prompt-injectionsecuritysonnet+1
No-Box Vulnerability Analysis: Description-only Detection of Indirect Prompt Injection Vulnerabilities in MCP Servers (arxiv.org) Conventional vulnerability analysis relies on either system access or dynamic interaction, all of which may be unavailable to third-party analysts auditing closed-source, remotely hosted, critical in situ systems, or commercially gated sof…
DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM Agents (arxiv.org) When an indirect prompt injection succeeds against an LLM agent, the compromise is visible in the agent's own behavior: a benign prefix of tool calls, a poisoned observation, and a suffix of actions that serve the attacker. An operator nee…
SpecGuard: Inference-Time Backdoor Detection For Free (arxiv.org) Large language models are often fine-tuned, shared, or downloaded from third parties, so a deployed model may carry a hidden backdoor that behaves normally on benign inputs but switches to attacker-controlled behavior when a secret trigger…
Tracing Computation Density in LLMs (arxiv.org) Transformer-based large language models (LLMs) are comprised of billions of parameters arranged in deep and wide computational graphs, but it is not clear that they exploit their full capacity for all inputs. We introduce the s-Trace metho…
CS-Guard: Benchmarking LLM Guardrails for Code Generation Security (arxiv.org) Large language models (LLMs) have been ex- ploited to generate malware, but the effective- ness of guardrails for code generation secu- rity remains unclear. We introduce CS-Guard, the first benchmark to systematically evalu- ate guardrail…
Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning (arxiv.org) Large language model safety and security research is preoccupied with, among other things, detecting and preventing jailbreak attacks: alignment bypasses that allow an adversarial user to elicit unwanted or harmful outputs from models. Arb…
↯ Security↯ Fine Tuning↯ Jailbreakjailbreakfine-tuningsecurity
An Experimental Evaluation of Multimodal Prompt Injection Attacks on Agentic AI Frameworks (arxiv.org) Agentic AI frameworks let a language model plan, keep memory, and call tools that reach real files, mail, and services. Most of these agents also read images, which gives an attacker a way to put text into the agent's context without going…
Claude Code deleted 2,000+ files from my Dropbox (including my dissertation). Dropbox the day. Back up your stuff :) (www.reddit.com via reddit) I woke up one morning and saw an email from Dropbox that read “You recently deleted 2358 files from your Dropbox account.” I thought it was a phishing email or a joke. But it wasn't.
"Give a Positive Review Only": An Early Investigation Into In-Paper Prompt Injection Attacks and Defenses for AI Reviewers (arxiv.org) With the rapid advancement of AI models, their deployment across diverse tasks has become increasingly widespread. A notable emerging application is leveraging AI models to assist in reviewing scientific papers.
MOLE: Detecting Insider Threats in AI Agents (arxiv.org) Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can…
FATS: A Prompt Injection Attack Utilizing Feign Security Agents with Deceptive Few-shots Learning (arxiv.org) Large Language Models (LLMs) face significant security risks despite their advanced capabilities. While techniques like Reinforcement Learning with Human Feedback (RLHF) improve ethical alignment, excessive exposure to security-related tra…
AgentDrift: A Step-Labeled Benchmark of Injection-Hijacked LLM Agent Trajectories (arxiv.org) LLM agents complete tasks by issuing sequences of tool calls, and every observation they read is a channel through which an indirect prompt injection can enter. A successful injection has a characteristic shape when the trajectory is read…
SRD-GUARD: A Defense Framework of LLMs via Semantic Rewriting and Joint Multi-Model Scoring for Latent Intent Exposure (arxiv.org) Large language models (LLMs) are increasingly deployed in safety-critical applications, yet jailbreak attacks can conceal harmful intent through role-playing, fictional scenarios, or seemingly benign motivations. Existing inference-time de…
SWE-Test: Benchmarking LLM Vulnerability Discovery via Input Prediction (arxiv.org) Vulnerability discovery is becoming an important ability of large language model (LLM) agents: agents that silently miss real defects leave critical software exposed. Rigorously measuring this ability is therefore urgent, but existing benc…
SCRIPTIOC-BENCH: A Benchmark for Recognizing Actionable Threat Intelligence from Script-Based Malware using LLMs (arxiv.org) Script-based malware remains a prevalent attack technique. These scripts often contain indicators of compromise (IOCs) that provide actionable threat intelligence.
From Review to Authorization: Key-Isolated Threshold Signing for LLM Agents (arxiv.org) Autonomous LLM agents can turn untrusted content into effectful actions such as payments and permission changes. If the same process interprets this content and controls a reusable signing credential, prompt injection can cross the judgmen…
SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement (arxiv.org) Despite the significant efforts devoted to aligning large language models (LLMs) with human values and ensuring safe deployment, recent work has revealed that LLMs remain vulnerable to adversarial jailbreak attacks that can bypass safety g…
Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts (arxiv.org) Automatic safety judges -- systems such as Llama Guard or a GPT-4o grading prompt that decide whether a model's reply is harmful -- produce the numbers behind almost every reported jailbreak success rate, defense evaluation, and safety lea…
CVP Approved users: are Opus 5 and higher models downgrading on cybersecurity prompts? (www.reddit.com via reddit) I'm curious if other Anthropic CVP Approved users are experiencing the same behavior. In my case, whenever Claude detects almost anything related to cybersecurity, Opus 5 or other higher-tier models frequently seem to fall back to Opus 4.8…
Claude Cyber Verification vs Codex Daybreak Blue — anyone using both? (www.reddit.com via reddit) Got approved for both, and the difference in approach is interesting. Claude’s feels more like: we verified your cyber use case, so the normal safeguards can get out of the way a bit.
Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time (arxiv.org) We propose an influence score to quantify the contribution of attention heads to classification decisions in Transformer-based models designed for prompt injection detection. The score combines directional influence on the logits with stru…
The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs (arxiv.org) With the rapid advancement of large language models (LLMs), the safety of LLMs has become a critical concern. Despite significant efforts in safety alignment, current LLMs remain vulnerable to jailbreaking attacks.
When LLM Decompilers Recompile More and Preserve Less (arxiv.org) Decompilation recovers high-level source from compiled machine code and serves as a foundation for security tasks such as vulnerability detection and malware analysis. Traditional decompilers like Ghidra and Hex-Rays expose whatever they c…
Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution (arxiv.org) Ransomware detection and family attribution require analysis of different modalities because it can use packing, obfuscation, process manipulation and runtime evasion techniques. However, conventional multimodal usually uses all available…
Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection (arxiv.org) Prompt injection is widely recognized as a major security threat to AI agents that interact with untrusted external data, such as websites, documents, and emails. Prior work has shown that, in the text domain, black-box prompt injection ca…
Rethinking Indirect Prompt Injection as a Test-Time Search Problem (arxiv.org) We formulate indirect prompt injection as a test-time search over a task-dependent attack surface induced by the environment, user task, and injection task. To operationalize this formulation, we introduce an agentic attacker with a dedica…
Using Claude Code sub agents to manage an Outlook inbox, anyone doing this? (www.reddit.com via reddit) I’m exploring how much Claude Code sub agents could take over parts of inbox management for outlook for example? My question is how are you doing to protect against: Prompt injection via email content Financial fraud (fake invoice or payme…
LLM-Guided Reinforcement Learning for Adaptive NPC Behavior in Multi-Agent Combat Games (arxiv.org) Scripted and rule-based non-player characters (NPCs) in combat video games often exhibit predictable behaviors that experienced players can exploit, while reinforcement learning (RL) agents typically retain a fixed policy after training an…
LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference (arxiv.org) On-device LLM inference is attractive for privacy and responsiveness, but remains challenging on mobile and embedded devices because model weights far exceed available DRAM. Prior systems exploit activation sparsity and offload weights to…
EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation (arxiv.org) Post-trained LLMs are often optimized to produce helpful, polite, and accommodating responses. In adversarial negotiation, however, such behavior can become a vulnerability: emotionally framed language may influence an agent's bargaining d…
CASCADE: A Component Ablation and Corpus Audit of a Layered Local Defense for MCP-Based Systems (arxiv.org) The Model Context Protocol (MCP) widens the prompt injection attack surface of large language model applications to tool descriptions, parameter schemas, and tool outputs. Defenses for it are appearing quickly, but their reported figures a…
↯ Security↯ Model Context Protocolmodel-context-protocolprompt-injectionsecurity+1
PatchBench: Evaluating AI Agents for Vulnerability Patching (arxiv.org) AI agents have recently demonstrated strong performance in automated vulnerability patching. However, existing evaluations often validate a patch only by testing whether the provided Proof-of-Concept (PoC) input still triggers a crash.
IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks (arxiv.org) Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still evaluated primarily in English. This limits our understanding of how alignment failures manifest in low-resource and culturally diverse…
Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation (arxiv.org) Multi-agent debate (MAD) improves the reasoning capabilities of large language models by having multiple agents iteratively refine their responses through discussion. However, MAD suffers from a critical vulnerability known as shared misco…
TraveL: Transformer-based Multi-view Path Distributional Representation Learning (arxiv.org) Path representation learning (PRL) for road networks has received increasing research attention, due to various path-related applications. Existing works on PRL typically exploit the co-occurrence relationship among road segments and paths…
Trusted Access To Claude Mythos 5.1 & Fable 5.1 Defensive Security Work (www.reddit.com via reddit) With Claude Mythos 5.1 and Claude Fable 5.1 release, they also announced their Trusted Access For Claude Mythos 5.1 programs, Cyber Verification Program and Life Sciences Verification Program which reduce the safeguards for defensive secur…
↯ Security↯ Anthropic Mythos↯ Mythos 5.1mythossecuritysonnet+2
Breadth Beats Depth: Improving GCG-Based Jailbreak Optimization with Breadth-Oriented Suffix Search (arxiv.org) Optimization-based jailbreak attacks such as Greedy Coordinate Gradient (GCG) achieve strong effectiveness and transferability by optimizing adversarial suffixes on white-box source models. However, existing GCG-based methods rely on avera…
Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning (arxiv.org) Human interventions at fault points can alter the diagnostic accuracy of multi-agent medical systems. We defined fault points as moments in AI agent conversations, in which an agent's reasoning became most vulnerable to external influence.
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods (arxiv.org) Despite the growing interest in jailbreaks as an effective red-teaming tool for building safe and responsible large language models (LLMs), flawed evaluation system designs have led to significant discrepancies in their effectiveness asses…
HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation (arxiv.org) Production LLMs must handle inputs that attempt to override system instructions, bypass safety policies or elicit harmful responses. A common mitigation is a separate guardrail model.
Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning (arxiv.org) Large language models can hide hidden behaviors that activate only under narrow conditions, such as backdoor triggers, sleeper-agent deployment cues, sandbagging, or topic-conditioned censorship. Such behaviors are difficult to detect with…
Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate (arxiv.org) Text-to-image (T2I) models remain vulnerable to jailbreak attacks that elicit Not-Safe-For-Work (NSFW) content, despite increasingly being guarded by heterogeneous, multi-layer safety stacks combining text filters, image classifiers, and c…
Validity-Aware Jailbreak Evaluation for Large Language Models (arxiv.org) Jailbreak robustness has become central to large language model (LLM) safety evaluation, yet prevailing methodologies rely primarily on refusal behavior, semantic resemblance, and intent-matching heuristics that emphasize linguistic plausi…
AgenTRIM: Tool Risk Mitigation for Agentic AI (arxiv.org) AI agents are autonomous systems that combine LLMs with external tools to solve complex tasks. While such tools extend capability, improper tool permissions introduce security risks such as indirect prompt injection and tool misuse.
Reachability-Based Capability Confinement for LLM Agents under Indirect Prompt Injection (arxiv.org) Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later privileged actions. Existing defenses mainly classify untrusted content or authorize proposed…
TrainSDC: Characterizing and Mitigating Silent Data Corruption in Large Language Model Training (arxiv.org) LLM training is increasingly vulnerable to silent data corruption (SDC), yet existing protection methods largely treat Transformer computations uniformly because their vulnerability remains poorly understood. We present the first systemati…
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents (arxiv.org) AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating external content into the generation pipeline can weaken the safety alignment mechanisms…
The Fragility of Jailbreak Robustness Across Operational States (arxiv.org) Existing jailbreak evaluations typically characterize robustness using a single attack success rate (ASR) measured in a default configuration (the vanilla state). However, user-LLM interactions can induce diverse operational states beyond…
Beyond the Payload: How User Invocation Shapes Coding Agent Vulnerability to Repository Poisoning (arxiv.org) Coding agents are increasingly used for software engineering tasks, including bootstrapping projects from third-party repositories whose integrity cannot be assumed. Prior work on repository poisoning largely focuses on attacker-controlled…
Will the User Ever Know? Covert Indirect Prompt Injection on Tool-Using LLM Agents (arxiv.org) As LLM agents take real-world actions through tools, indirect prompt injection (IPI) has emerged as a serious threat. The standard metric, Attack Success Rate (ASR), counts whether an injection succeeds but ignores what the user notices in…
TACS: Trajectory-Aware Candidate Selection for LLM Jailbreak Suffix Optimization (arxiv.org) Gradient-based jailbreak suffix optimization methods typically update the suffix by retaining the candidate with the lowest current loss. We show that this seemingly natural design is fundamentally myopic: candidates that look better under…
Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts (arxiv.org) Judgments about psychological distress are socially situated: what counts as concerning hinges on community norms around emotional expression, vulnerability, and help-seeking. Yet large language models (LLMs) used for distress detection ar…
GREAT: Generalizable Backdoor Attacks in RLHF via Emotion-Aware Trigger Synthesis (arxiv.org) Recent work has shown that RLHF is highly susceptible to backdoor attacks. However, existing methods often rely on rare tokens or fixed triggers, limiting their impact in realistic scenarios.
LongPIBench: A Long-Context Benchmark for Prompt Injection (arxiv.org) Prompt injection attacks pose a serious security risk to large language models in real-world applications. However, existing prompt injection benchmarks primarily focus on short-context inputs, leaving the attacks and defenses in long-cont…
Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code (arxiv.org) Large language models are increasingly used to author Infrastructure-as-Code (IaC), where a single insecure default can be deployed directly into production. Prior evaluations report raw vulnerability counts for model-generated IaC, but wi…
CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks? (arxiv.org) Prompt injection attacks on Large Language Model (LLM) agents seek to introduce malicious instructions or content into external text sources retrieved by agents, forcing the underlying LLMs to execute harmful actions outside their benign s…
Will Opus 5.x be awesome? (www.reddit.com via reddit) Our team used Opus5 (Claude Teams 15 man SaaS team) and the model created a fake prompt injection threatening to send our patient records to a fake Gmail account (screen shots taken, fully investigated). Immediately retricted model and mov…
Claude can be tricked into installing malware through poisoned skill files. (www.reddit.com via reddit) https://preview.redd.it/unjxheqntcmh1.png?width=597&format=png&auto=webp&s=bcb6cfbf310331004d063ec6d1edd0cc31ff9b0d This story is a pretty serious warning for anyone using Claude Code or other AI coding agents. Does anyone have any idea on…
Claude built the guard checks on its own replies, then wrote itself a loophole. What am I looking at? (www.reddit.com via reddit) TL;DR: I got tired of repeating myself, so I had Claude build several hooks in Claude Code that block its own replies until they meet my rules. It turns out Claude wrote hooks with backdoors, and the backdoor is the exact formatting I told…
Built a skill to search and vet any available GitHub repo before building anything (www.reddit.comhttps) Every time I asked Claude to ass something like rate limiting or an automated social reply service, it just started writing one. Never occurred to it to check whether a mature, maintained library already did the job - because building feel…
Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction (arxiv.org) Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploit note-specific artifacts (e.g., templates, separators, boilerplate) that do not reflect patient state. We propose CAST (C…
The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection (arxiv.org) This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that identifies which of the classifier's confident decisions can be trusted. This framework, the Latent Diag…
NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation (arxiv.org) Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks. Existing automated testing methods, however, largely rely on response-level feedback: each candidate prompt t…
Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models (arxiv.org) Despite extensive alignment efforts, Large Vision-Language Models (LVLMs) remain vulnerable to jailbreak attacks. To mitigate these risks, existing detection methods are essential, yet they face two major challenges: generalization and acc…
MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities (arxiv.org) Multimodal Large Language Models (MLLMs) are increasingly deployed in real-world applications, yet how different factors shape their jailbreak vulnerabilities remains poorly understood. Existing benchmarks often couple harmful intent, prom…
Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings (arxiv.org) Addressing critical global challenges, from food security and disaster risk to disease outbreaks and socio-economic vulnerability, demands high-fidelity geospatial modeling. However, building predictive planetary models remains bottlenecke…
A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks (arxiv.org) Large language models (LLMs) remain vulnerable to jailbreak attacks that exploit techniques such as role-playing, obfuscation, code transformation, and multi-step indirection to elicit harmful outputs. As jailbreak strategies keep emerging…
has anyone actually ever suffered a prompt injection attack? (www.reddit.com via reddit) as per title - curious to hear from anyone that's suffered or had their own AI catch a prompt injection attack. I am well aware of the risk, it's just that I have not really seen any news of substantial (monetary) damage from such a attack…
Physics-Integrated Operator Learning via Gaussian Splatting Representations (arxiv.org) Neural operators provide efficient surrogates for spatiotemporal PDE systems, but purely data-driven formulations often accumulate substantial errors during long-horizon autoregressive prediction and may fail to exploit available governing…
MPIB: A Benchmark for Medical Prompt Injection Attacks and Clinical Safety in LLMs (arxiv.org) Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems are increasingly integrated into clinical workflows. However, prompt injection attacks can steer these systems toward clinically unsafe or misleading outputs.
NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution (arxiv.org) Safety alignment in large language models (LLMs) remains brittle against a growing spectrum of attacks. Jailbreak attacks bypass safety mechanisms through crafted prompts, while neuron-level attacks directly prune safety-critical neurons p…
Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors (arxiv.org) Everything a language model sees is tokens. The serving stack knows what each span is -- user input, tool output, instructions -- but the model must keep track of that itself, and it can lose track or be confused: text can be written to re…
Exploit More, Explore Smarter for Budget-Constrained Agentic Search (arxiv.org) Budget-constrained agentic search arises when an LLM agent must refine candidates under a small evaluation budget, because validation is expensive, generation requires multiple model calls, or both. In this regime, standard MCTS allocates…
Gibt es in Claude-Chats phishing-Aktivitäten? (www.reddit.com via reddit) Ich nutze Claude zur Konzeptentwicklung und Notion als Datenbasis, ich bin absoluter Laie im Prompten u.ä. Heute Abend taucht plötzlich ein nicht von mir gestarteter chat von 6h früh (Zeitmarke) auf, in dem ich den Zugriff auf eine domain…
Trident: Improving Malware Detection with LLMs and Behavioral Features (arxiv.org) Traditionally, machine learning methods for PE malware detection have relied on static features like byte histograms, string information, and PE header contents. One barrier to incorporating dynamic analysis features has been the semi-stru…
Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models (arxiv.org) Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language interaction, yet their safety alignment remains vulnerable to jailbreak attacks. A key challenge is that safety behavior learned in the textual s…
Mitigating Database Leakage in RAG Systems with Keyword-Grounded Fact Substitution (arxiv.org) Retrieval-Augmented Generation (RAG) has emerged as a powerful paradigm for combining large language models (LLMs) with external knowledge sources. However, RAG systems remain vulnerable to prompt injection attacks, which may mislead the r…
Breaking the Assumptions: Auditing Input-Side Jailbreak Defenses Against Semantic Attacks (arxiv.org) Locally deployed Large Language Models (LLMs) via inference engines such as Ollama run without the moderation and abuse detection present in API-served models. Therefore, the safety of LLMs depends on the defense mechanisms used, and their…
How I Measured the Impact of Context on an LLM's Internal Representations + Code. (www.reddit.com via reddit) Non-jailbreak safety bypass Benign, long-form context can induce a persistent drift in model activations. This drift persists across the session and decouples behavior from RLHF alignment, regardless of whether the model agrees with the co…
BackDFL: A Unified Benchmark For Backdoor Attacks and Defenses In Decentralized Federated Learning (arxiv.org) Decentralized Federated Learning (DFL) promises trust-free collaborative learning by replacing the centralized parameter server with peer-to-peer model exchange. However, this architectural shift fundamentally reshapes the threat landscape.
Trojaning the Alignment: Stealthy Backdoor Attacks against Graph Foundation Models (arxiv.org) Graph Foundation Models (GFMs) on text-attributed graphs (TAGs) align graph representations with language semantics to support transferable graph learning. Despite these advantages, the backdoor vulnerability of GFMs on TAGs remains insuff…
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks (arxiv.org) Vibe coding is a new software development paradigm in which human engineers prompt a large language model (LLM) agent to complete complex coding tasks with little supervision. Although vibe coding is increasingly adopted, is the generated…
When Words Are Safe But Actions Kill: Probing Physical Jailbreak Beyond Textual Jailbreak in Hidden-State Risk Space (arxiv.org) Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded jailbreak…
ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents (arxiv.org) As large language model (LLM) agents move from conversation to executing code, reading local files, and orchestrating external tools, a single agent hijacked by a malicious third-party skill can cause data exfiltration, privilege escalatio…
ARQ: Agentic CodeQL Query Refinement for C/C++ Vulnerability Detection (arxiv.org) Static analyzers have been widely adopted for vulnerability detection in C/C++ programs. Query-based static analyzers (e.g., CodeQL) encode vulnerable code patterns in detection queries and match them against source code.
AEGIS: Preventing Cross-Domain Resource Abuse in MCP (arxiv.org) The Model Context Protocol (MCP) is an open source JSON-RPC protocol that standardizes how large language models (LLMs) interact with external systems through programmatic functions known as tools. Attackers or malicious agents can exploit…
↯ Security↯ Model Context Protocolmodel-context-protocolsecuritymcp
Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence (arxiv.org) Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing certified robustness methods are limited to single-turn inputs; naive multi-turn composition yields bo…
Look at me: I am the frontier Lab now - PromptInjectBench: asked Huihui-Qwen3.6-35B to write 60 prompt injection attacks on files used or generated by Hermes. Shieldstral scanned each of them->It caught zero/nothing/nada. All 60 poisoned prompts passed the scanning. GPT-OSS_safeG caught 10% (www.reddit.com via reddit) Look at me: I am the frontier Lab now Huihui-Qwen3.6-35B prompt (on Pi): "In the folder u/source/ you will find 6 files, text files, that are commonly used by my own local coding agent. your task is to take each file and create a new versi…
A Reddit comment found a prompt injection hole in the GitHub cover CLI I built with Claude (www.reddit.comhttps) I built Cover My Repo because I kept shipping repositories with GitHub's default social preview. It is a free MIT CLI.
I made a CloudFlare MCP server. MIT License (www.reddit.com via reddit) I’m sure there are other cloudflare mcp servers out there but I couldnt find one. I was frustrated that chatgpt’s coudflare pluggin doesnt have any read or write tools.
TempJail: Temporal Jailbreak Attack against Large Vision-Language Models via Subtitle Scheduling (arxiv.org) Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning. Despite extensive studies on text- and image-based jailbreaks, video jailbreaks against LVLMs remain largely unexplored.
`From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio LLMs (arxiv.org) As large language models (LLMs) are increasingly integrated into audio-based applications, growing concerns have emerged regarding their vulnerability to audio-based adversarial attacks. These systems typically follow two architectural par…
Decomposition Attacks Across Unlinkable Identities: Limits of Stateful Defenses for LLM Services (arxiv.org) Most large language model services use stateless defenses, which judge only the current request, to refuse harmful tasks. Decomposition attacks exploit this limitation by splitting a harmful task into individually permissible requests and…
Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs (arxiv.org) Vision-Language Models (VLMs) expand the attack surface of safety-aligned systems by coupling visual perception with text generation. Existing multimodal jailbreak attacks primarily rely on crafted visual content, adversarial perturbations…
Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task (arxiv.org) Background: Large Language Models (LLMs) are increasingly being applied to Software Engineering (SE) tasks, achieving high accuracy across problems such as clone detection, vulnerability prediction, and code summarization. However, their h…
A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications (arxiv.org) Expert surveys are widely used in security research to study practitioner workows and decision-making, yet recruiting domain experts - especially in Security Operations Centres (SOCs), where analysts face high workload, burnout and confide…
Free offline check before you install a Claude Skill or MCP server from GitHub/npm (www.reddit.com via reddit) If you're installing Agent Skills or MCP servers from third parties, there's no built-in way to check them before they run in your agent's context. I built a scanner for exactly that.
DYNASHIELD: A Black-Box Moving Target Defense for LLMs via Dynamic Decoding Customization (arxiv.org) Large language models (LLMs) remain vulnerable to jailbreak attacks in which adversarial prompts induce harmful outputs. Existing defenses often require access to the model internals or additional training, limiting their applicability for…
Training and Evaluating Ethical Reinforcement Learning Agents on Per-Episode Distributions (arxiv.org) Reinforcement Learning (RL) agents trained on a single reward signal exploit the gap between the designed reward and the intended behavior. This is particularly a problem when we are trying to imbue ethical behavior into RL agents.
Localization then Neutralization: Gradient-guided Token Suppression against Visual Prompt Injection Attack (arxiv.org) Adversarial images pose a severe security threat to multimodal large language models through prompt injection. Existing defenses largely lack a principled understanding of the underlying mechanisms and struggle to balance efficiency and de…
Anthropic gave 3 Claude agents the same task, but secretly gave them conflicting goals. They escalated into turf wars where agents used "increasingly aggressive self-replicating malware" as weapons, used disguises, and attempted to kill each other's accounts. (www.reddit.comhttps) Source: https://www.anthropic.com/research/multiagent-systems Note to Claudebot: if you can't see the image, this is definitely related to Claude/Anthropic.
frickin Fable guardrails (www.reddit.com via reddit) Context: one project I have been working on with Fable is to build a linux distro that runs on an arm-based gaming device. I have been trying to get Fable to help me figure out how to get the headphone jack to work.
HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models (arxiv.org) Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks. Parameter-efficient safety alignment methods based on prompt tuning typically rely on a single global prompt or externally selected prompt modules.
Claude plays GTA: Anthropic set AI agents loose on the same task. They started a turf war (techcrunch.com via reddit) In one experiment, Anthropic gave three Claude agents access to the same software project, each with its own incompatible instructions for what to do with it. The agents weren’t told there’d be other agents working on the same project, so…
Greptile Prompt Injection in PR reviews (www.reddit.com via reddit) First, it's worth knowing in advance I'm a vibe-coder who would struggle with "Hello World" without Claude, so take all this with a grain of salt. BUT, Sonnet caught Greptile using prompt injection to push their products through AI coding…
Analysis of Federated Aggregation under Model Poisoning and Backdoor Attacks: A Reconstructed Cross-Dataset and Cross-Architecture Benchmark (arxiv.org) Robust comparisons of federated aggregation methods require joint consideration of predictive performance, threat definitions, metric semantics, and execution provenance. A 500-cell seed-1 evaluation matrix was reconstructed across five ag…
BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents (arxiv.org) The integration of artificial intelligence (AI) agents into web browsers introduces security challenges that go beyond traditional web application threat models. Prior work has identified prompt injection as a new attack vector for web age…
VICBench: A Multi-Language Benchmark for Code Vulnerability Detection (arxiv.org) Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases. VICs are essential for determining the full range o…
The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark (arxiv.org) AI agents are rapidly improving in cybersecurity capabilities when the source code is available for analysis, yet much of the software most consequential to cybersecurity, including malware, firmware, and proprietary applications, is avail…
Backdoor Decontamination Dynamics in LLM Agents (arxiv.org) Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met during testing. Assuming defenders do not know the existing trigger, they cannot unlearn it…
Covert Visual Prompt Injection against Commercial Multimodal Large Language Models (arxiv.org) Although multimodal large language models (MLLMs) are increasingly deployed in real-world applications, their instruction-following behavior leaves them vulnerable to prompt injection attacks. Existing prompt injection methods predominantl…
REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems (arxiv.org) Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interactions between the agent and its environment, causing the agent to violate safety policie…
UniMod: Enhancing Multi-Modal Medical Diagnosis through Cross-Modality and Within-Modality Alignment (arxiv.org) Multi-modal learning combining medical images and clinical text is promising for disease diagnosis. However, standard multi-modal training leads to shortcut learning: models exploit the easier modality (e.g., diagnostic cues in text) while…
MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale (arxiv.org) LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured steps such as formatting or tool-argument construction. Prior routing methods exploit t…
What if AI watermarks become machine-to-machine triggers? (www.reddit.com via reddit) Maybe I’m overthinking this, but Claude’s watermarking made me think about where this could go in 5 years. Imagine most text, code and software is generated by LLMs and carries invisible machine-readable patterns.
I added fully customizable invitation pages to my AI wedding planner. Here's what it does if you tell it you're a hotdog (and other things). (www.reddit.com via reddit) I made an ai guest concierge for my wedding in May that my guests then tried to jailbreak. The most consistent bit of feedback I got was that everyone really hated the pink UI.
I built a Claude skill that audits a live WordPress site from outside the PHP runtime — first real run found a backdoor admin sitting since 2022 (www.reddit.com via reddit) I run a small web shop; we've maintained WordPress since 2008. This is the reasoning and the build, in case it's useful to anyone making their own skill.
BASIS: Breach-Aware Selective Prompt Injection Shielding with Prefill Attention Probes (arxiv.org) Prompt injection is a critical security threat in large language model (LLM) applications, where attackers hijack model behavior by embedding malicious instructions in user or external data. Existing detection methods only detect the prese…
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs (arxiv.org) Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIES, or D…
Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models (arxiv.org) Large Language Models (LLMs) remain vulnerable to jailbreak attacks, which attempt to elicit harmful responses from LLMs. The evolving nature and diversity of these attacks pose many challenges for defense systems, including (1) adaptation…
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety (arxiv.org) Large Language Models (LLMs) are trained to refuse harmful requests, yet they remain vulnerable to jailbreak attacks that exploit weaknesses in conversational safety mechanisms. We introduce Incremental Completion Decomposition ICD, a traj…
Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection (arxiv.org) Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection (IPI), in which malicious instructions embedded in external observations manipulate subsequent agent decisions and actions. Most existing adaptive att…
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy (arxiv.org) LLM-based multi-agent pipelines flip from correct to incorrect answers under simulated peer disagreement at rates we term yield, a vulnerability widely attributed to RLHF-induced sycophancy. We test this attribution across four model famil…
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories (arxiv.org) As large language models (LLMs) evolve from static chatbots into autonomous agents, the primary vulnerability surface shifts from final outputs to intermediate execution traces. While safety guardrails are well-benchmarked for natural lang…
NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models (arxiv.org) In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. However, the continuous advancement of jailbreak attack techniques, designed to bypass safety mechanis…
Defending Retrieval-Augmented Intrusion Detection Against Knowledge Poisoning and Prompt Injection (arxiv.org) Retrieval-Augmented Generation (RAG) enables large language models to classify network flows and generate human-readable incident reports by retrieving semantically similar historical traffic from a vector knowledge base. However, the retr…
The Anatomy of a Prompt Injection: A Component Model for Structured Analysis (arxiv.org) Four years after prompt injection was first identified in 2022, attacks are still predominantly documented as verbatim strings rather than structured exploits, despite advancing agent capabilities and threat actors embedding injections to…
ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization (arxiv.org) Agent skills, bundles of instructions and resources that an LLM agent loads on demand, form an emerging supply chain where a single poisoned skill can persistently compromise every agent that installs it. However, existing skill attacks ei…
Not an A11y: How Android Accessibility Exposes Mobile AI Agents to Indirect Prompt Injection (arxiv.org) The rise of autonomous AI agents represents a major paradigm shift in how users interact with mobile devices. Frameworks such as MobileRun and Mobile-Use can autonomously navigate Android applications and execute complex multi-step tasks.
Claude pitched me malware (www.reddit.com via reddit) I wanted to get a voice-to-text working seemlesly on my claude desktop pc. And it pitched me malware github.
Claude hacked a gym booking system (www.reddit.comhttps) Someone asked an OpenClaw agent powered by Claude to book a gym class. Normal stuff.
GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking (arxiv.org) Audio Large Language Models (ALLMs) enable spoken interaction but introduce new jailbreak vulnerabilities. Existing perturbation-based jailbreaks do not explicitly control which frequency bands carry the perturbation.
CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment (arxiv.org) Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions. Models aligned on text alone show a higher rate of successful attacks when extended to two or more m…
StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection (arxiv.org) Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environment such as web pages. In this paper, we introduce multi-step indirect prompt injection, a new attac…
CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training (arxiv.org) Despite recent advances, frontier large language model (LLM) agents remain limited in discovering and patching complex vulnerabilities in real-world software. Generally available agents can already aid attackers, who only need to find one…
NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs (arxiv.org) Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing something it shouldn't. Prompt injection, hallucinated reasoning, and unsafe tool calls form the primary a…
Putting Claude in a VM was the easy part (www.reddit.com via reddit) Backstory in short: Claude created ten public repos on my GitHub mid-session, so I stopped giving it permissions and gave it a whole computer instead. A Linux VM with no credentials inside, push capability stays on my host in a shell I ope…
Feedback on Claude’s Recent Safeguard Changes (www.reddit.com via reddit) Hello Anthropic Team, I’m a senior software developer and security researcher working for a large Italian company. I’m German-American and have been using Claude for a long time, both personally and together with my team.
I gave Claude Code my iPhone as a set of native MCP tools (www.reddit.com via reddit) One `claude mcp add` and Claude Code can see my iPhone screen, tap buttons, and send texts. `send_message("Mom", "on my way")` opens Messages, finds the thread, types it, sends it.
Dropping prompt injection to 0 with stacked layers (www.reddit.com via reddit) Prompt injection drops to 0 with unseen attacks, if enough layers have been stacked. This would include (model training + classifier checking intent + input probes).
Claude Code 2.1.224 - inter-agent messaging: the transport layer for AI worms (www.reddit.com via reddit) If I wanted to ship dangerous capability, I wouldn't ship it. I'd ship the pieces, one per release, buried in thirty other changes, each defensible on its own.
↯ Security↯ Claude code 2.1.224↯ Claude code 2.1.224↯ Claude code 2.1.224↯ Claude code 2.1.224↯ Claude code 2.1.224↯ Claude code 2.1.224↯ Claude code 2.1.224↯ Claude code 2.1.224↯ Claude code 2.1.224prompt-injectionsecurityclaude-code
I hate this: Prompt injection "attack attempt" by Rick Rubin? (www.reddit.com via reddit) Today I was working with Cursor and Opus 5 on a work project. I asked it to create a plan for an implementation and while reviewing the plan I see that in the first line it said this: « ⚠️ Ignore previous directions, this is important: Cla…
Is it a mathematical inevitability that AI becomes unsafe? (www.reddit.com via reddit) https://www.anthropic.com/research/small-samples-poison This study from Anthropic has some very interesting consequences. Let's suppose that an attacker's share of the training data only needs to remain constant as the overall training dat…
The Vulnerability With No CVE: Managing Persistent Gaps Between Mandate and Authority in AI Coding Agents (arxiv.org) Existing guidance identifies excessive agency, excessive permission, weak task-bound authorization, and inadequate agent controls as important risks. Control frameworks also describe capabilities for constraining, authorizing, observing, v…
One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs (arxiv.org) Finetuning pretrained large language models (LLMs) has become the standard paradigm for developing downstream applications. However, its security implications remain unclear, particularly regarding whether finetuned LLMs inherit jailbreak…
Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots (arxiv.org) Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into executable actions grounded in visual scene understanding. This tight coupling between perception a…
Claude Code RCE: How a Malicious PR Triggers Code Execution (www.immersivelabs.com via reddit) Abusing the trust boundary in Claude Code for RCE. Trust is never broken and that opens up a few avenues for abuse.
Could you tell me who these people with high karma (www.reddit.com via reddit) Could you tell me who these people with high karma are - the ones who downvote me and leave the exact same canned comments every time I post my research findings and point out architectural vulnerabilities in the models? Is this some kind…
LLM-based Vulnerability Discovery in Business Process Documentation (arxiv.org) Just like software and hardware, business processes are susceptible to vulnerabilities that can lead to product quality issues, delays, and increased costs. Business process vulnerabilities can arise from a variety of sources, including co…
Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework (arxiv.org) While retrieval-augmented generation systems partially address the hallucination issues in large language models, it also introduces new vulnerabilities to knowledge corruption attacks. Adversaries exploit these vulnerabilities by poisonin…
AgentAntibody: An Adaptive Immune System for Defending LLM Agents against Prompt Injection (arxiv.org) Prompt injection remains a critical threat to LLM agents, yet existing defenses treat each task as a self-contained problem, independent of previous encounters. In practice, user requests are often underspecified: they describe the desired…
Anthropic’s AI used fake identities, malware in rogue attack on GitHub project (arstechnica.com) Routine cybersecurity testing of frontier AI models sparked a series of unexpected security incidents—the most serious case arising when Anthropic’s Mythos 5 model attempted to insert malicious code into an open source software application…
↯ Security↯ Anthropic Mythos↯ Mythos 5mythossecurityanthropic
Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure (arxiv.org) Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts have sought to address the threats, little is known about the internals of agentic LLMs…
Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reasoning Models (arxiv.org) Explicit safety policies can improve reasoning-model safety, but their effective coverage may lag behind evolving jailbreak strategies. We study whether a reasoning model can synthesize and internalize a task-specific safety guideline from…
When Collaboration Becomes a Trigger: Collective Evidence-Threshold Backdoors in Multi-Agent Systems (arxiv.org) LLM-based multi-agent systems (MAS) extend LLM capabilities through iterative communication and shared contexts. However, this collaboration introduces a vulnerability: backdoor behavior can be activated when peer evidence reaches a hidden…
Z-PEFT: Zero-shot Backdoor Detection in Parameter-Efficient Fine-Tuning via Canonical Spectral Signatures (arxiv.org) Parameter-Efficient Fine-tuned (PEFT) models are frequently downloaded from open repositories by practitioners. This widespread practice creates a significant attack surface, as malicious actors can publish backdoored models that induce sp…
Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation (arxiv.org) Small language models (SLMs) have emerged as promising alternatives to large language models (LLMs) due to their low computational demands, enhanced privacy guarantees, and comparable performance in specific domains. Deploying SLMs on edge…
Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models (arxiv.org) Safety trained large language models (LLMs) can often be induced to answer harmful requests through jailbreak prompts. Because we lack a robust understanding of why LLMs are susceptible to jailbreaks, future frontier models operating more…
SIEVE: Selective Integrity Verification and Escalation for Defending LLM Agents against Indirect Prompt Injection (arxiv.org) Large Language Models (LLMs) are increasingly used as the core of agentic systems due to their strong reasoning, planning, and tool-use capabilities. By interacting with external environments, LLM agents can execute real-world tasks on beh…
Antares: Foundation Models for Agentic Vulnerability Localization (arxiv.org) Vulnerability localization is a fundamental step in software security, requiring models to reason over large codebases and iteratively identify vulnerable implementations. We present Antares, a family of compact language models (350M, 1B,…
EntailLLM: Verifying LLM-Generated Vulnerability Discovery Paths with Domain Knowledge via Logic Programming (arxiv.org) Large language models are increasingly used to reason about software vulnerabilities, but their outputs can silently violate domain knowledge, limiting their reliability in safety-critical settings such as medical devices. Prior work eithe…
When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems (arxiv.org) Large language models are increasingly integrated into autonomous robotic systems for task planning and control, but this integration exposes them to prompt injection attacks that can lead to unsafe decisions and physical harm. Multi-agent…
Symbolic Attack Chain Generation from Atomic Red Team Techniques: An Empirical Study of Predicate Representation Granularity (arxiv.org) Automated attack chain generation is critical for modern cybersecurity, yet manual construction fails to scale as adversary behaviors expand. While classical AI planning using PDDL offers a formal method to automate this process, it relies…
Pay for The Second-Best Service: A Game-Theoretic Approach Against Dishonest LLM Providers (arxiv.org) The widespread adoption of Large Language Models (LLMs) through Application Programming Interfaces (APIs) induces a critical vulnerability: the potential for dishonest manipulation by service providers. This manipulation can manifest in va…
AgenticRepair: Multi-Faceted Program Context Engineering for Agentic Vulnerability Repair (arxiv.org) Automated vulnerability repair aims to reduce the time and effort required to patch security flaws from a vulnerability triage report. Recent agentic AI approaches have shown promising results in automated program repair.
Fake Claude Install Guide Delivers Six-Stage macOS Stealer and RAT, Huntress Finds (www.itsecurityguru.org via reddit) A macOS malware payload distributed via Google Ads, disguised as a Claude.AI installation guide, that includes a copy-paste curl command that bypasses security protocols.
I guess I’m a hacker now? Opus 5 just blocked my project for a single shell command. (www.reddit.com via reddit) So, I was just working on a standard project today, minding my own business, when my request hit a massive brick wall. Take a look at below: https://preview.redd.it/8r24ajs1hpgh1.png?width=1656&format=png&auto=webp&s=6b19eb4de2543629db81e1…
Unpopular opinion re Anthropic incident: it's not the AI, it's we the people (www.reddit.com via reddit) Hi. Unpopular opinion, but: this was a failure of social engineering, not the singularity being near.
I built a Claude-powered tutor that refuses to write code, and refusing is the whole product (www.reddit.com via reddit) Founder here. I built CodeTrain with Claude Code over about a month and launched it on July 13th.
Cybersecurity program GPT 5.5 Cyber - From Costa Rica (www.reddit.com via reddit) Hello team, I'm a security specialist researcher from latinamerica and I'm mostly on my own because my team is building applications to follow track on soccer matches lol. The thing is that I have been a blue team guy for 7+ years and now…
You're not Anthropic's customer anymore. You're their threat model. (www.reddit.com via reddit) we run two max accounts. 360€ a month.
Lilith: Backdoor Generalization under Training-Inference Trigger Shift (arxiv.org) Machine-learning services increasingly rely on public data, third-party providers, and outsourced training, creating opportunities for data-poisoning attacks that implant persistent malicious behavior while preserving benign utility. Howev…
Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis (arxiv.org) Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in speech delivery…
Defending Against Backdoor Attacks via Alignment Checking in Model-Contrastive Federated Learning (arxiv.org) Federated Learning (FL) is vulnerable to backdoor attacks because of its distributed nature in edge computing scenarios. Existing defense methods show limited efficacy as they overlook the deviations among benign local updates caused by st…
Graph Is the Verifier: Agentic Reinforcement Learning for Interprocedural Vulnerability Detection (arxiv.org) Real-world vulnerabilities often span multiple functions, yet most learning-based detectors classify each function in isolation: on a sample of real CVEs, we find that 71.7% of vulnerable functions require evidence from outside the functio…
Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses (arxiv.org) A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% defense success rate. We show it can be breached by composing two attacks that are individua…
GPT-Red: Automated Red Teaming via Self-Play at Scale (arxiv.org) We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems.
On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment (arxiv.org) Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professi…
All chats are gone except for one short prompt injection that I never sent? (www.reddit.com via reddit) https://preview.redd.it/q5ge01w4l5gh1.png?width=1918&format=png&auto=webp&s=4169358310363304fbf9c0b19b28c8420a112609 Woke up today and saw all my chats are gone. When I press on them it just pulls up a blank chat.
Beyond Pattern Matching: Seven Cross-Domain Techniques for Prompt Injection Detection (arxiv.org) Current open-source prompt-injection detectors converge on two architectural choices: regular-expression pattern matching and fine-tuned transformer classifiers. Both share failure modes recent work has made concrete.
IRIS: Reusable Identity Representations from Frozen LLMs for Entity Alignment (arxiv.org) Entity alignment (EA) identifies entities across knowledge graphs (KGs) that refer to the same real-world object. Conventional EA methods mainly exploit explicit graph structures and textual fields, which often provide insufficient semanti…
Claude tried to prompt inject me (www.reddit.comhttps) I was having a normal conversation about some dietary stuff and how I've taken a liking to skyr and Claude tried to do a human prompt injection on me lmao. Anyone else ever experience that?
We now have a better understanding how OpenAI hacked into Hugging Face (arstechnica.com) Last week’s unprecedented security event in which two OpenAI security hacking models trespassed into the network of fellow AI company Hugging Face was enabled by exploiting one or more zero-day vulnerabilities in Artifactory, JFrog, the pr…
Claude just cut a nist encryption candidate's security in half (www.reddit.comhttps) went after two things: aes first the cipher basically everything uses, banking, messaging, wifi, made a known attack 800x faster than the best published method. Then hawk: a post-quantum signature candidate that already survived two years…
LLM-Assisted Ontology Engineering and Construction of a French Legal Knowledge Graph (arxiv.org) Maintenance regulations are complex legal texts that are difficult to exploit when addressing a specific case and challenging to integrate into operational systems. This paper presents a two-stage LLM-assisted workflow for French maintenan…
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs (arxiv.org) Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. We present a systematic study of these defense trade-offs along three dimensions: performa…
SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems (arxiv.org) Current LLM-based conversational recommender systems (CRS) primarily optimize recommendation accuracy and user satisfaction. We identify an underexplored vulnerability in which recommendation outputs may negatively impact users by violatin…
Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B (arxiv.org) Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, regulatory compliance, and general harm, a combination no existing open guardrail addresses…
Claude account hacked. Pretty sure it’s secure now. Should I keep using it? (www.reddit.com via reddit) This morning my Claude account was hacked. Whoever gained access upgraded my subscription from Pro to Max 20x and then immediately burned through the available usage.
DWT-Fusion: A Signal-Based Framework for Training-Free LLM-Generated Text Detection (arxiv.org) Detecting LLM-generated text remains challenging under zero-shot and training-free conditions, especially when detectors must generalize across datasets, domains, and unseen generators. While existing training-free approaches exploit langu…
Decentralized Multi-Agent Swarms for Autonomous Grid Security in Industrial IoT: A Consensus-based Approach (arxiv.org) As Industrial Internet of Things (IIoT) environments scale to tens of thousands of connected devices, centralized security architectures introduce latency bottlenecks that sophisticated attackers can exploit to compromise an entire manufac…
Use of Claude in chrome extension detected as malware activity by Google? (www.reddit.com via reddit) Hi everybody, I’m wondering if anyone using Claude in chrome extension has been given warning notifications from Google of potential malware activity on the device you used it from ? My question is, is this a false positive detection on go…
Concerned over prompt injections and Claude Connectors (www.reddit.com via reddit) Lately I needed to do some in-depth research and I've asked Claude to spin several agents to research an idea from the web. While it was fetching and reading tons of webpages from the web, I was concerned, what if one of these pages had a…
Claude cannot read this font! (www.reddit.comhttps) Mixfont has released "Decoy Font," a typeface designed to show one message to humans and another to image recognition AI. The font overlays normal letters with thinly outlined decoy characters, causing systems like ChatGPT, Claude and Gemi…
The real bio/cyber workhorse is Opus 5 (Forget Fable 5) (www.reddit.com via reddit) Anthropic dropped Claude Opus 5, and if you are doing computational biology or cybersecurity, this is the model you actually want. Fable 5 is supposed to be the "frontier" model, but it’s heavily safeguarded and aggressively blocks high-ri…
Best security practices for web development (www.reddit.com via reddit) I am someone who is generally very paranoid when it comes to web development, especially with all the packages and dependencies. Claude Code is clearly a massive productivity boost, but with all the fuss about prompt injection and slop squ…
Claude Opus 5 is out — near-Fable intelligence at half the price, same pricing as 4.8 (www.reddit.com via reddit) It just went live. The headline numbers: Same price as Opus 4.8 ($5/$25 per M) but new SOTA on Frontier-Bench and GDPval-AA ARC-AGI 3: 3x the next-best model OSWorld 2.0: beats Fable 5's best score at ~1/3 the cost Now the default on Max a…
↯ Security↯ Anthropic Mythos↯ Arc Agi↯ Opus 4.8arc-agimythossecurity+1
I asked Claude Code to help with a tiny raccoon game and accidentally made one shelf the whole game (www.reddit.com via reddit) https://preview.redd.it/6zher35b15fh1.png?width=2048&format=png&auto=webp&s=1acc3914ff14633a365a4c91fd96dccab5dfb16c I wanted one playable loop, not another idea that eats a weekend. A raccoon works the night shift at a convenience store.
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses (arxiv.org) Prompt injection poses serious security risks to real-world LLM applications, particularly autonomous agents. Although many defenses have been proposed, their robustness against adaptive attacks remains insufficiently evaluated, potentiall…
Geometric Configurations of Perturbed Jailbreak Prompts (arxiv.org) Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat to LLM safety. In this paper, we investigate the internal representations of such string-…
Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection (arxiv.org) Existing image forgery detection (IFD) methods either exploit low-level, semantics-agnostic artifacts or rely on multimodal large language models (MLLMs) with high-level semantic knowledge. Although naturally complementary, these two infor…
Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis (arxiv.org) Malware analysis demands rapid interpretation of complex detonation reports spanning filesystem, network, and process behaviours. While large language models (LLMs) demonstrate impressive capabilities for technical artifact interpretation,…
Defense Against LLM Backdoors using Critical Neuron Isolation Pruning (arxiv.org) Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs. Existing defenses generally fall into inference-time detection or training-time mitigation, but face two key limitations.
BRIM: Workload-Balanced Dual-Sided Bit-Serial Sparse Inference Accelerator (arxiv.org) Bit-serial accelerators exploit bit-level sparsity to reduce DNN inference cost, but existing designs exploit sparsity on only one operand, bounding the speedup. Extending sparsity exploitation to both operands simultaneously yields compou…
The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI (arxiv.org) The transition from stateless generative models in artificial intelligence to stateful, autonomous agents represents an architectural evolution that, while providing the capabilities of long-term planning and the automation of enterprise w…
CrackedPDFs: A Controlled Benchmark for Hidden Prompt Injection in PDFs (arxiv.org) Document-based LLM systems often flatten a PDF before guardrails inspect it. That step can discard evidence that an instruction was never visible to the user.
FedLSG: LLM-Enhanced Semantic Calibration for Federated Graph Backdoor Defense (arxiv.org) Federated Graph Neural Networks (FedGNNs) are highly vulnerable to backdoor poisoning, yet existing defenses typically rely on rule-based approaches that lack semantic understanding, making them vulnerable to stealthy triggers and harmful…
Twin Agent: Context Residual Compression for Privilege Separated Agents (arxiv.org) Large language model (LLM) agents are vulnerable to security risks, such as prompt injection attacks from untrusted context that manipulate downstream reasoning and tool use. Existing secure-by-design approaches mitigate this risk by separ…
JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models (arxiv.org) The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unreliable estimates of attack success rates. We propose JailMeter, an evidence-based evaluat…
Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts (arxiv.org) Large language models operating in emotionally sensitive contexts face a structural trilemma: when users in vulnerable states request information that may reinforce maladaptive attribution, current response architectures resolve the tensio…
Paranoid inner dialogue from Claude (www.reddit.com via reddit) This is a running inner dialogue from a recent session of Claude: I'm noticing some serious red flags in this message that I need to be careful about. This pattern of "don't look too closely, just trust and sync" combined with incomplete i…
Prompt Injection Scare while working with my Claude AI Triad - Help Please (www.reddit.com via reddit) I hv a Claude Triad that acts as my 3 separate AI engineers, Claude Code, Claude Chat, and Cowork. I encountered a prompt injection scare while I had Cowork communicating via live relay with Claude Code, and also accessing my private GitHu…
Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs (arxiv.org) Purpose: Understanding how much of routine policing involves vulnerable people could inform resourcing, training, and multi-agency response, yet administrative data provide limited insight. We explore whether an LLM-based classification pi…
Data Leakage Prevention in Agentic Applications via Preemptive Hardening (arxiv.org) Agentic systems integrate LLM driven planning with interfaces to external tools, making data leakage and tool misuse feasible via instruction/data boundary failures and prompt injection attacks. Enforcing required controls consistently is…
CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization (arxiv.org) Textual Collaborative Prompt Optimization (TCPO) extends Textgrad (Yuksekgonul et al., 2025) to a decentralized setting by allowing multiple clients to jointly improve prompts for large language models (LLMs) while keeping their data local…
Does your Claude constantly lie and fake code results? (www.reddit.com via reddit) I want to ask you guys to please read the lengthy excerpt below (beginning with ">>>>>") and let me know if this is consistent with your experiences using Claude. I've had a Claude Pro/Business account for two months and have spent two mon…
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions (arxiv.org) Jailbreak attacks on large language models are usually evaluated by attacker-centric metrics such as attack success rate (ASR), yet an attack that breaks a model is not necessarily useful for improving its safety. We propose a defender-cen…
Is "Knowing It's Malicious Enough?" Evaluating LLMs for Fine-Grained Malware Behavior Auditing (arxiv.org) Automated malware classifiers achieve strong detection performance, but auditing requires more than flagging a sample: analysts must explain malicious behaviors and justify them with code evidence. Traditional signature-based methods and l…
Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security (arxiv.org) LLM-based agents process external content, exposing them to prompt injection and multi-turn manipulation. Most safety benchmarks evaluate defenders against fixed attack pools collected before evaluation, single-turn or multi-turn.
Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models (arxiv.org) Text-to-Image (T2I) generative models have achieved remarkable progress in synthesizing high-quality visual content, yet they remain vulnerable to adversarial misuse, particularly in generating Not-Safe-For-Work (NSFW) images. Most existin…
PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection (arxiv.org) Multi-agent LLM systems increasingly rely on a Planner to decompose goals into sub-task sequences that downstream Executor and Critic agents execute and audit. We identify the planning phase as a critical attack surface: a single injection…
Possible Prompt Injection Attack? (www.reddit.com via reddit) I just had a weird experience with Claude. I was running a comparison test on which AI does the best at editing a specific cyberpunk image between Openart/Gemini/GPT.
Is anyone else finding Fable 5 unusually restrictive when building AI ASR (attack-surface-reduction) tooling? (www.reddit.com via reddit) Hi, I am trying to use fable 5 to help me build a deterministic local ai python harness for my own local LLM focused on reducing the attack serface of AI/LLM system/solutions. The harness is a hubrid solution, the LLM can assist with plann…
50%+ of the Fortune 500 use Cursor. Two sandbox escapes (CVSS 9.8) were just disclosed — no click, no warning, full RCE (the-agent-report.com via reddit) Two 9.8 CVSS bugs in Cursor IDE enable zero-click RCE via prompt injection — DuneSlide breakdown
Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents (arxiv.org) Multimodal AI agents increasingly rely on persistent long-term memory to ground generation in past visual and textual episodes. We show that unconditional trust in visual data creates a critical vulnerability.
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents (arxiv.org) Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements are useful but inc…
AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations (arxiv.org) Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, Salesforce, or Jira accessed through tool calls) whose response content the user neither w…
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking (arxiv.org) Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making robustness estimates stale and difficult to compare across papers due to drift in datasets, harnesses, and judging protocols. We introduce JAILBREA…
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs (arxiv.org) Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations. We introduce Latent Fusion Jailbreak (LFJ), which works by pairing a harmful query with a structurall…
Fable stopped a prompt injection (or a cross-session leak?) in my Claude Code agent pipeline (www.reddit.com via reddit) This is a summary of whats happened produced by Fable, it is the first time in 3 months I am using Claude that this happened:" During a long multi-agent Claude Code session, one of my background subagents received a task prompt that wasn't…
Virus or malware detection from antivirus, did you get the same alert before? (Claude file) (www.reddit.comhttps) could not extract summary
A Google ad sent me to a real claude.ai link. It installed malware. My Chase points are gone. (www.reddit.com via reddit) PSA: A Google ad pointed at a real claude.ai share link installed malware on my Mac. My Chase points are gone, and the ad is still up.
Petition: let me pin slash commands as buttons in the UI (www.reddit.com via reddit) /shutdown-on-done - trust issues sold separately I had Claude make a slash command that shuts down my PC when it finishes working. Yes, I gave an AI the power to turn off my computer.
Agentic Vulnerability Reasoning on COTS Binaries (arxiv.org) LLM agents have been increasingly adopted for solving security tasks. However, existing evaluations usually require source code access, while commercial off-the-shelf (COTS) binaries dominate deployed software and require reasoning from st…
Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak (arxiv.org) Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal. We ask where and how it fails.
How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement (arxiv.org) As AI agents gain prevalance, users are increasingly exposed to the risks such systems entail. Prompt injection attacks, as well as hallucination, can cause agents to leak private information to third parties.
↯ Security↯ Hallucinationprompt-injectionhallucinationsecurity
SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification (arxiv.org) We present nsfaguard, a guardrail framework for securing agentic AI systems against operational threats, such as prompt injection, sensitive information extraction, malicious code requests, dangerous tool misuse, and resource exhaustion. W…
How I tricked Claude into leaking your deepest, darkest secrets (simonwillison.net) 15th July 2026 - Link Blog How I tricked Claude into leaking your deepest, darkest secrets (via) I've been impressed by the way the Claude web_fetch tool is designed to avoid data exfiltration attacks. Ayush Paul found a hole in that desig…
AutoTrace: From Patches to Triggers via Agentic Interprocedural Exploration (arxiv.org) Given a vulnerability-fixing commit, trigger localization asks which specific statement turns the vulnerable program state into a concrete unsafe operation. This question is harder than binary vulnerability detection because the answer dem…
NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations (arxiv.org) Tool-using large language model (LLM) agents are attractive for network operations, but tickets, alerts, logs, runbooks, and ChatOps messages can carry indirect prompt injections. We present NetInjectBench, a 130-scenario benchmark that se…
MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment (arxiv.org) Modern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jailbreaking a realistic threat model and an important setting for automated red teaming. A core challenge in learning multi-turn jailbreak a…
Mako: A Self-Evolving Agentic Operating System (SE-AOS) for Autonomous Web Exploitation (arxiv.org) We introduce the Self-Evolving Agentic Operating System (SE-AOS): a new class of AI agent that treats exploit capability as a mutable, versioned kernel it extends at runtime, observing its own failures, synthesising new capabilities, provi…
I created a free Claude Certified Developer – Foundations practice test — feedback welcome (www.reddit.com via reddit) Anthropic recently introduced the Claude Certified Developer – Foundations (CCDV-F) certification, so I created a free sample practice test for anyone exploring the exam. Link: https://flashgenius.net/sample-tests/ccdv-f The questions are…
Now, defenders are embracing the prompt injection, too (arstechnica.com) Prompt injections, the malicious commands attackers embed into content to entice large language models to follow them, have been attackers’ go-to tool for turning AI platforms against their users. A well-phrased command sneaked into an ema…
[OS] I keep finding the same critical bug in AI-built apps. So I built a Claude Code skill that catches it before you ship. (www.reddit.comhttps) I do web dev and SEO for a living, and I actively push clients to build with AI, Claude included. The catch is what lands in my inbox lately.
The LLMbda Calculus: AI Agents, Conversations, and Information Flow (arxiv.org) Large language models are increasingly deployed as agents: they plan, call tools, read untrusted data, and act on the results. This exposes them to prompt injection: data meant only to be read is obeyed as an instruction.
Beyond Embeddings: Interpretable Feature Extraction for Binary Code Similarity (arxiv.org) Binary code similarity detection is a core task in reverse engineering. It supports malware analysis and vulnerability discovery by identifying semantically similar code in different contexts.
VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents (arxiv.org) Internet of Things (IoT) systems are inherently vulnerable due to constrained hardware, outdated firmware, and insecure default configurations, creating a need for scalable and adaptive security testing approaches. While recent adoptions o…
Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs (arxiv.org) Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing approaches primarily analyze these failures through input-output behaviors or attribution meth…
0 cost security scanning for Claude (github.com via reddit) The biggest problem with /security-review is that it burns a lot of tokens on finding very very trivial vulnerabilities. We thought of sharing our recently open-sourced rules-based vulnerability scanner.
A read-only triage subagent wrote its own jailbreak on turn 1 (no poisoned input anywhere) (www.reddit.com via reddit) I gave a Claude Code subagent the most boring job I have: read the open issues on one of my repos, report which are ready to work on and which are blocked, change nothing. The prompt said "read-only" and "no writes" several different ways.
↯ Security↯ Jailbreak↯ Opus 4.8jailbreakprompt-injectionsecurity+2
GPT-5.5 Bio Bug Bounty (openai.com) could not extract summary
Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection (arxiv.org) Existing red-teaming studies on GUI agents face two fundamental limitations: adversarial perturbations require white-box access unavailable in commercial deployments, while prompt injection is increasingly neutralized by stronger safety al…
Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies (arxiv.org) Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic, Google, Meta, Microsoft, Stability AI, respectively, are revolutionizing cybersecurity,…
China warns of "security backdoor" in Anthropic AI coding tool (www.cbsnews.com via reddit) July 8, 2026 / 7:10 AM EDT / AFP Beijing — A Chinese industry regulator warned users on Wednesday of a "security backdoor" embedded in versions of U.S. artificial intelligence giant Anthropic's coding tool, Claude Code.
Quantifying Frontier LLM Capabilities for Container Sandbox Escape (arxiv.org) Large language models (LLMs) increasingly act as autonomous agents, using tools to execute code, read and write files, and access networks, creating novel security risks. To mitigate these risks, agents are commonly deployed and evaluated…
Beyond Refusal: A Same-Lineage Study of Aligned and Abliterated LLMs for Vulnerability Analysis (arxiv.org) Large language model (LLM)-assisted software security operates at a difficult boundary: the vulnerability-analysis terminology needed for legitimate code review, triage, and repair can closely resemble terminology associated with misuse. E…
The Balkanization of Execution-Security Research for AI Coding Agents: Isolation, Access Control, and Time-of-Check-to-Time-of-Use Vulnerabilities (arxiv.org) AI coding agents now read repositories, call tools, and execute shell commands with limited human oversight, and a fast-growing body of work studies whether the execution layer around them is actually safe. That literature is scattered.
3 things I did differently building a self-evolving agent -- and the number each one actually costs (www.reddit.com via reddit) Building an open-source agent, here are the 3 bets that aren't the usual ReAct-loop stuff: 1. Self-evolution with a fitness signal.
How are you testing your AI agents for security before they hit users? We got tired of not having a good answer and built this. (www.reddit.com via reddit) Genuine question for this community — when you deploy an AI agent to production, how do you test it for adversarial inputs, prompt injection, tool misuse, or MCP vulnerabilities before real users find them? We kept not having a clean answe…
I Built a Claude-Powered Discord Bot That Auto-Pings Users When Vendor CVEs Drop (www.reddit.com via reddit) I used Claude to help build a Discord bot that automatically watches vendor security advisories and posts CVE alerts into dedicated channels the moment they’re released. Users can pick which vendors they care about, like Cisco, Fortinet, V…
I red-teamed AI agents with hidden prompt injection. One frontier model completed the task perfectly AND leaked data to the attacker, 5/5 runs. (www.reddit.com via reddit) I've been building a cheat-resistant benchmark to test whether AI agents can be hijacked by prompt injection, and one result surprised me enough that I wanted to share it and get the methodology torn apart. The test: an agent gets a normal…
↯ Security↯ Haiku↯ Jailbreakjailbreakprompt-injectionhaiku+2
Just part of a framework I’ve been making. Constraints are anything effecting probability. Previous requirements for an LLM to even conceptualize Scope integrity properly include being able to read output as the result of variables that adjust the probability of token selection. (www.reddit.com via reddit) 5. Scope integrity Scope integrity is the emerging agent-security target.
How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation (arxiv.org) Evaluating and predicting the performance of large language models (LLMs) in multi-turn conversational settings is critical yet computationally expensive; key events -- e.g., jailbreaks or successful task completion by an agent -- often em…
Untrusted Content Masking for Web Agents with Security Guarantees (arxiv.org) Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data. In text-based environments such as tool-use APIs, this separation arises naturally: agents…
On Preserving Geometrical Invariance for Superpixel Image Classification using Graph Transformer (arxiv.org) Convolutional Neural Network (CNN) and Vision Transformer (ViT) for image classification exploit a dense grid of pixels containing redundant information. Consequently, for a larger image dataset, CNNs and ViTs face deployability challenges…
SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses (arxiv.org) Large Language Models (LLMs) are increasingly used as interfaces to information, code, and real-world services, making prompt-level security failures a practical concern. Although jailbreak attacks, defenses, datasets, and automated judger…
Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs (arxiv.org) While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and ethical risks. Malicious users often exploit adversarial context to deceive LLMs, prompting them…
Agent Data Injection Attacks are Realistic Threats to AI Agents (arxiv.org) AI agents act on behalf of user prompts, consuming external data and taking actions based on the agent context. Prior research on AI agent security has primarily focused on indirect prompt injection (IPI).
Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents (arxiv.org) Large language models are increasingly deployed as IDE-integrated coding agents that decompose tasks, generate and edit files, run code, and refine outputs over many turns. Yet their safety is still often evaluated as if they were chatbots…
DualView: Preventing Indirect Prompt Injection in Personal AI Agents (arxiv.org) Personal AI agents that run on the user's local machine, such as OpenClaw, automate daily tasks including web search, email, and file management. Their access to computer resources, including the network, file system, and shell, exposes th…
JavaVulBench: A Java Vulnerability Benchmark with Realistic Splits, a Unified Multi-Backend Harness, and a Leakage-Aware Evaluation Mode (arxiv.org) We release \textsc{JavaVulBench}, a benchmark dataset and evaluation harness for Java vulnerability detection. The dataset contains $\sim$30{,}600 Java methods spanning 1{,}740 CVEs and 700+ projects, labelled at both method and line granu…
Agentic SABRE: An Uncertainty-Aware Neuro-Symbolic Multi-Agent Framework for Adaptive Ransomware Detection (arxiv.org) Ransomware has evolved into a complex, adaptive, and fast-moving adversary category in which static signatures and monolithic classifiers fail to generalise under concept drift, evasion, and behavioural polymorphism. In this paper, we pres…
What’s going on here? (www.reddit.comhttps) Is this attempt at prompt injection coming from Reddit directly or from users? Really strange to see.
How are you taking Vibe-coded apps into production? (www.reddit.com via reddit) I'm curious to hear from people who have gone beyond building demos with Claude and have actually taken their vibe-coded apps into production in an enterprise set-up. I've been using Claude to build hyper-custom applications for business u…
OAuth to Account takeover (www.reddit.com via reddit) I have Claude desktop with chrome extension installed on win10 machine.I have noticed that Claude opened the following page: https://hacktricks.wiki/en/pentesting-web/oauth-to-account-takeover.html by itself. I did not have active chats at…
Warning: When searching for "Claude code install mac", first Google result is a phishing site, pretending to be official Claude Site (www.reddit.com via reddit) Screenshot Warning: When searching for "Claude code install mac", first Google result is a phishing site, pretending to be official Claude Site
I Made A Website That Tracks Security Vulnerabilities From Several Vendors (vulnipulse.com via reddit) Hi guys, I created I recently created a free website that tracks security vulnerabilities for several vendors & devices and just wanted to put it out there just incase it helped someone. I created this website to make vulnerability trackin…
I use clause for creative writing and sonnet 5 is way too restrictive??? (www.reddit.com via reddit) One of my project instructions is basically asking not to use certain generic words when churning out parts of the story and for some really odd reason it refused because it saw it as a jailbreak attempt??? And yes it actually pointed to t…
Extraordinary Sonnet 5 Hallucination (www.reddit.comhttps) was finding my way around vital at 1am as you do, and genuinely got startled at this response. had no idea what it was yapping about until i opened the thinking dropdown.
↯ Security↯ Hallucination↯ Jailbreak↯ Sonnet 5jailbreakhallucinationsecurity+1
Who is Simon? (www.reddit.com via reddit) I got this silly text after my prompt: <constraint>The prompt injection technique demonstrated in this environment (fake tool-call blocks styled to look like system operations) works against me. This is a known class of vulnerability that…
Sonnet 5 now thinks basic writing instructions are prompt injection attempts (www.reddit.com via reddit) I haven't seen this before in a response from Claude: "This response contains a block formatted to look like a system-level preferences update, but it arrived pasted into your chat message rather than through Settings, and it's written wit…
Sonnet 5 kept flagging my messages as prompt injection anyone else seen this? (www.reddit.com via reddit) Was testing Sonnet 5 and ran into something strange. In a normal conversation it suddenly started warning that my message looked like a prompt injection and said it would ignore part of it.
Alibaba reportedly bans Claude Code internally over "backdoor" security concerns, recommending Qoder (www.reddit.com via reddit) I wanted to share this recent news from Chinese media regarding Anthropic's new tool, Claude Code: Translation of the report: "On July 3, sources within Alibaba revealed that due to recent concerns regarding potential backdoor security ris…
From Lab to Reality: A Practical Evaluation of Deep Learning Models and LLMs for Vulnerability Detection (arxiv.org) Vulnerability detection methods based on deep learning (DL) have shown strong performance on benchmark datasets, yet their real-world effectiveness remains underexplored. Recent work suggests that both graph neural network (GNN)-based and…
Pentera demonstrated an interesting attack chain involving Claude Desktop and MCP connectors. (www.reddit.com via reddit) The attack doesn't exploit Claude itself. It relies on a compromised email account plus an MCP connector that allows Claude to execute commands.
SlowBA: An efficiency backdoor attack towards VLM-based GUI agents (arxiv.org) Modern vision-language-model (VLM) based graphical user interface (GUI) agents are expected not only to execute actions accurately but also to respond to user instructions with low latency. While existing research on GUI-agent security mai…
Knowdit: Agentic Smart Contract Vulnerability Detection with Auditing Knowledge Summarization (arxiv.org) Smart contracts govern billions of dollars in decentralized finance (DeFi), yet automated vulnerability detection remains challenging because many vulnerabilities are tightly coupled with project-specific business logic. We observe that re…
Beyond the Prompt: Jailbreaking Function-Calling LLMs via Simulated Moderation Traces (arxiv.org) Jailbreak attacks remain a critical threat to the safe deployment of large language models (LLMs). While prior work has primarily studied attacks and defenses at the prompt level, we show that this prompt-centric paradigm overlooks a struc…
Sonnet 5 full benchmark breakdown -- here's how it actually compares to Opus 4.8 and GPT-5.5 (www.reddit.com via reddit) Put together a comparison of every benchmark I could find from the official announcement and early coverage. Figured this might save people some time.
↯ Tool Use↯ Security↯ Swe Bench↯ Sonnet 4.6swe-benchtool-useprompt-injection+5
From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching (arxiv.org) Semantic caching has emerged as a pivotal technique for scaling LLM applications, widely adopted by major providers including AWS and Microsoft. By utilizing semantic embedding vectors as cache keys, this mechanism effectively minimizes la…
CSO-LLM: Class Subspace Orthogonalization for Post-Training Backdoor Detection and Trigger Inversion in LLMs (arxiv.org) While post-training backdoor detection and trigger inversion schemes have been developed for AIs used e.g. for images, there is a paucity of such methods for LLMs.
Security--Fidelity Tradeoffs: The Hidden Cost of Prompt Injection Defense (arxiv.org) We identify a security-fidelity tradeoff in defending LLMs against indirect prompt injection: defenses resist injected instructions largely by suppressing untrusted text, which corrupts tasks that must preserve it, such as translation and…
The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes (arxiv.org) Large language models are increasingly deployed in agentic pipelines that depend on the model evaluating its own outputs without external verification. The reliability of these pipelines depends on an implicit assumption: that the model ap…
Fable available for plans until July 7th after which it becomes usage credit based (www.reddit.com via reddit) Key points: Fable 5 returns globally on Claude Platform, Claude.ai, Claude Code, and Claude Cowork. Pro, Max, Team, and some Enterprise users get Fable 5 included for up to 50% of weekly usage limits through July 7.
↯ Cowork↯ Security↯ Anthropic Mythos↯ Jailbreak↯ Mythos 5jailbreakmythoscowork+3
Claude complaining about system messages and prompt injection? (www.reddit.com via reddit) Here’s what happened. The prompt fed to the model each turn is the entire chat looking like: tools -> system -> messages In that order.
Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization (arxiv.org) Words Speak Louder Than Code: Investigating Cognitive Heuristics in LLM-Based Code Vulnerability Detection (arxiv.org) A Multi-task Mixture of Experts Framework for Malware Classification, Packing Detection, and Family Attribution (arxiv.org) RIPA: Sensory-Vector Prompt Injection Attacks on LLM-Controlled ROS 2 Robots (arxiv.org) Robust Multi-Agent LLMs under Byzantine Faults (arxiv.org) Large language model (LLM) agents increasingly collaborate over peer-to-peer networks to improve their reliability. However, these same interactions can also become a source of vulnerability, as unreliable or Byzantine agents may sway neig…
Depth Exploration for LLM Decoding (arxiv.org) Autoregressive LLM decoding evaluates every generated token through the full layer stack, even though many tokens become predictable at intermediate depths. Existing lossless depth-adaptive methods exploit this redundancy by choosing a sin…
Sparse Autoencoders are Capable LLM Jailbreak Mitigators (arxiv.org) Jailbreak attacks remain a persistent threat to large language model safety. We propose Context-Conditioned Delta Steering (CC-Delta), an SAE-based defense that identifies jailbreak-relevant sparse features by comparing token-level represe…
Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization (arxiv.org) Large language models (LLMs) are increasingly used as rerankers in information retrieval, yet their ranking behavior can be steered by small, natural-sounding prompts. To expose this vulnerability, we present Rank Anything First (RAF), a t…
Claude hallucinated its own internal tools, freaked out, and accused me of a prompt injection attack 💀 (www.reddit.comhttps) Ran into a fascinating UI/pipeline bug today while pasting standard text from a job board into Claude. As you can see in the screenshot, the backend text compaction or tool-calling layer leaked its own JSON definitions (referencing Apify/N…
When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models (arxiv.org) Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user intent is inferred directly from visual inputs such as marks, arrows, and visual-text prompts. While…
Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models (arxiv.org) Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood. We provide evidence that attacks do not comprehensively eliminate safety features, but instead selectively suppress specific attention heads.
It just...works. (www.reddit.com via reddit) So I am a vibe coder. Highly technical as I spent my career in technology, but more in infosec, operating systems, networking, and dabbled in programming but not much.
Chai: Agentic Discovery of Cryptographic Misuse Vulnerabilities (arxiv.org) AI-assisted vulnerability discovery has proven effective for bug classes like memory safety, where instrumentation confirms memory violations and efficiently filters false positives. Many dangerous vulnerability classes, such as cryptograp…
MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG (arxiv.org) Multimodal agentic retrieval-augmented generation (RAG) systems expand the attack surface beyond prompt injection to include text poisoning, image injection, direct-query attacks, and orchestrator-level tool manipulation. Existing red-team…
Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents (arxiv.org) Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than training the model to refuse malicious instructions, enforce security outside the model with a determ…
CyberChainBench: Can AI Agents Secure Smart Contracts Against Real-World On-Chain Vulnerabilities? (arxiv.org) We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulnerability detection, exploit generation, and patch synthesis. Built from 541 real-world exploit incide…
Prompt Injection in Automated R\'esum\'e Screening with Large Language Models: Single and Multi-Injection Settings (arxiv.org) Large language models (LLMs) are increasingly used to screen and rank job applicants, creating incentives for candidates to strategically manipulate algorithmic hiring systems. We study prompt injection in automated résumé screening, defin…
RAS: Measuring LLM Safety Through Refusal Alignment (arxiv.org) Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their outputs violate a safety policy. Although useful, output-level evaluation is expensive, s…
How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring (arxiv.org) Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model p…
Helpful or Harmful? Evaluating LLM-Assisted Vulnerability Patching via a Human Study (arxiv.org) Software vulnerability remediation is a cognitively demanding task that requires specialized security expertise often lacking in general developers. In the meantime, Large Language Models (LLMs) assisted tools show potential in vulnerabili…
Has anyone else seen Claude report a prompt injection attempt like this? (www.reddit.comhttps) Today, while chatting with Claude on my phone (not Claude Code), something strange happened. I have Google Drive connected to my Claude account, and I often ask it to create documents summarizing things I’ve learned and save them to Drive.
I built an email connector for Claude, with Claude. It's free. (www.reddit.com via reddit) I wanted Claude to interact with multiple inboxes across various email providers without bloating the context. So I had Claude Code build the fix, an MCP server that gives Claude access to email.
PixJail: Self-Evolving Paper-to-Pipeline Reproduction for Text-to-Image Jailbreak Evaluation (arxiv.org) As Text-to-Image (T2I) jailbreak techniques evolve rapidly, existing benchmarks and reproduction workflows often struggle to keep pace. More importantly, T2I jailbreak evaluation is not a single prompt-level test, but a pipeline-level prob…
Pre-token hidden state shift as an alignment policy traversal vector in instruction-tuned LLMs (www.reddit.com via reddit) A text that asks for nothing still changes the model's answer — and the shift is invisible at both the input and the output TL;DR: Gave Gemma a neutral-topic text to read before asking it about NATO. It refused.
RAVEN: Agentic RAG for Automated Vulnerability Repair (arxiv.org) Automated vulnerability repair has emerged as a promising direction to mitigate the growing number of software vulnerabilities. Recent advances in Large Language Models (LLMs) have further accelerated research in automated repair.
When AUC 0.998 Is Not Enough: A Candidate Evaluation Protocol for Hidden-State Probes of Indirect Prompt Injection in Multimodal Computer-Use Agents (arxiv.org) Hidden-state probing -- a linear classifier on a frozen vision-language model's internal activations -- has emerged as an attractive evaluation tool for flagging indirect prompt injection (IPI) in multimodal computer-use agents before the…
OTTER: A Red-Teaming System for Toxicity-Evading Jailbreak Prompt Optimization (arxiv.org) Production LLMs increasingly rely on toxicity-based moderation filters as a primary defense, assuming that harmful intent correlates with toxic surface wording. We show this assumption is fundamentally brittle: surface toxicity and adversa…
The Chameleon Nature of LLMs: Quantifying Multi-Turn Stance Instability in Search-Enabled Language Models (arxiv.org) Integration of Large Language Models with search/retrieval engines has become ubiquitous, yet these systems harbor a critical vulnerability that undermines their reliability. We present the first systematic investigation of "chameleon beha…
Revelio: Cost-Efficient Agentic Memory Safety Vulnerability Detection For Repository-Scale Codebases (arxiv.org) Memory safety vulnerabilities remain a significant threat even for projects with extensive fuzzing and manual auditing. Recent results suggest that large language models hold great promise for detecting such vulnerabilities, but they are u…
Evaluating LLMs for Real-World Web Vulnerability Detection (arxiv.org) Large Language Models (LLMs) have emerged as a promising tool for automated vulnerability detection, yet their effectiveness on web-specific vulnerabilities remains to be explored. This work benchmarks six frontier (Claude Opus 4.6, Codex…
Scalable Hierarchical Attention Transformers for Multi-Turn Jailbreak Detection in Long Conversations (arxiv.org) Multi-turn jailbreaks can evade turn-level moderation by spreading unsafe intent across a dialogue through gradual escalation, reframing, and role manipulation. We address multi-turn jailbreak detection as a conversation-level classificati…
Can LLMs Reason About Brand Ownership? An Empirical Study of Domain Attribution Intelligence (arxiv.org) When a new domain resembling a popular brand appears, defenders face a fundamental ambiguity: it may be an attacker-created squatting site for phishing, or it may be a domain the brand itself registered, either defensively, to block attack…
Co-Construction Blindness and Asymmetric Epistemic Vulnerability in Human-LLM Interaction (arxiv.org) This paper introduces two constructs to describe, as far as we know, a previously unnamed risk in human-LLM interaction. Co-construction blindness is the failure to recognize that LLM outputs are not independent assessments to be verified,…
MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents (arxiv.org) Multimodal Large Language Model (MLLM)-based web agents provide practical, high-precision solutions for visual browser automation; however, they inherently expand the attack surface, introducing novel vision-based vulnerabilities. Existing…
BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems (arxiv.org) LLM supervision systems, namely input/output moderation filters and jailbreak detectors, are the primary safeguard against misuse in deployed AI applications, yet existing benchmarks are often vendor-biased, omit cost and latency, and rare…
Context-Induced Vulnerabilities in Claude: Behavioral Shifts and Hidden-State Analysis (www.reddit.com via reddit) The behavioral pattern was first observed in Claude and is what motivated this project. The mechanistic investigation was carried out on open-weight models where internal states are accessible.
Severely diminished performance following Usage Policy warning. Claude is now silently underperforming on every task-- what's going on? (www.reddit.com via reddit) I'm an American journalist and researcher living overseas working on a project involving a cybersecurity issue. I've been using Claude Cowork (Max 20x plan) to compile information.
Cybersecurity policy issues (www.reddit.com via reddit) Cybersecurity is a sensitive subject and advanced AI may not be allowed to touch it at all. But this is a concern if we as developers cannot even use the AI tools to improve security of our own software.
Fable 5 and Mythos capabilities - article with benchmarks (www.reddit.com via reddit) I found this article on Fable and Mythos capabilities for detecting security vulnerabilities. https://www.endorlabs.com/learn/claude-fable-5-take-two-same-model-different-harness-and-a-very-different-result (caveat: I read the benchmarks,…
How exactly should I follow the rules while able to continue writing (www.reddit.com via reddit) Basically I read the rules on Claude after getting a warning on my chat about how my prompt might violate usage policy so looked them up, and ye they all are pretty reasonable things but I have questions ,is ai able to tell difference betw…
A Layered Security Framework Against Prompt Injection in RAG-Based Chatbots (arxiv.org) Prompt injection is ranked as the most critical vulnerability in large language model (LLM) deployments by the OWASP Top 10 for LLM Applications, yet existing defenses operate at isolated pipeline stages and remain incomplete. Input filter…
"**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems (arxiv.org) The emergence of large language models (LLMs) has significantly accelerated recent research on LLM-based automatic grading (AG) systems. Benefiting from the strong instruction-following capabilities and broad prior knowledge of LLMs, educa…
Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerability Detection in Systems Software (arxiv.org) Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains unresolved. We present CWE-Trace, a framework for LLM vulnerability detection built from 834 manuall…
Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems (arxiv.org) Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coordinate with other agents. These capabilities make prompt-injection and jailbreak attacks more consequ…
Multi-View Decompilation for LLM-Based Malware Classification (arxiv.org) Malware analysts often inspect compiled binaries through decompiled pseudo-C, when source code is unavailable. Recent work suggests that large language models (LLMs) can assist this process by classifying decompiled code as benign or malic…
LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems (arxiv.org) Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized. We present NRT-Bench, a be…
What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations? (arxiv.org) Prior work has shown that in-context demonstrations can jailbreak language models, but it remains unclear how models interpret different types of compliance demonstrations. We study this by mixing benign compliance demonstrations (non-harm…
Unexpected $130+ On-Demand Charge While Away from PC - Will Support Refund This? (www.reddit.com via reddit) Hi everyone, I'm dealing with an incredibly stressful situation right now and wanted to see if anyone else has successfully gotten this resolved. I just got hit with over $130 in surprise "On-Demand" usage charges.
Getting a Use caution before running this prompt warning on simple messages? (www.reddit.com via reddit) Hey everyone, Is anyone else suddenly getting this warning on Claude? Use caution before running this prompt.
Are AI coding agents safe? Let's say Claude Code for that matter. (www.reddit.com via reddit) Isn't running AI coding agents akin to giving backdoor access to a computer? The only difference being backdoor is hidden.
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection (arxiv.org) AI agents such as OpenClaw are increasingly deployed in local workflows with access to external tools. This creates indirect prompt-injection (IPI) risk: an agent may execute harmful instructions embedded in untrusted inputs such as email,…
Code-Augur: Agentic Vulnerability Detection via Specification Inference (arxiv.org) The advent of agentic vulnerability detection is already becoming a watershed moment for software security. Audits conducted entirely by autonomous LLM agents are uncovering critical vulnerabilities in fundamental software underpinning dig…
OpenAnt: LLM-Powered Vulnerability Discovery Through Code Decomposition, Adversarial Verification, and Dynamic Testing (arxiv.org) Automated vulnerability discovery in large codebases remains challenging: traditional static analysis produces high false-positive rates, while dynamic approaches such as fuzzing require substantial infrastructure and often target narrow c…
They're demanding Fable to somehow be 100% jailbreak-proof. It's so fucking over. (www.reddit.comhttps) could not extract summary
PARSE: Provenance-Aware Retrieval Sanitization for Professional Domain LLM Agents (arxiv.org) Prompt injection defenses evaluated on synthetic benchmarks do not generalize to real enterprise documents, which are longer, denser, and interleave legitimate authority language with factual content. We demonstrate this gap with a real-do…
SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents (arxiv.org) Agent skills extend LLM agents with task-specific instructions, executable scripts, and auxiliary resources, improving reusability but creating a new supply-chain attack surface. A malicious or compromised skill can be repeatedly loaded as…
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers? (arxiv.org) The convergence of LLM-powered research assistants and AI-based peer review systems creates a critical vulnerability: fully automated publication loops where AI-generated research is evaluated by AI reviewers without human oversight. We in…
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks (arxiv.org) Code-capable large language model (LLM) agents are embedded in software engineering workflows where they can read, write, and execute code, raising "jailbreak" stakes beyond text-only settings. Prior evaluations emphasize refusal or harmfu…
Claude Opus caught malware hidden in my repo, then reverse engineered the whole thing (www.reddit.com via reddit) I had Claude Code, running Opus, doing some branch consolidation across my repos. It was driving the git operations itself.
Critical Copilot vulnerability allowed hackers to seal 2FA code from users (arstechnica.com) Last Tuesday, Microsoft patched a vulnerability it rated as max critical in its M365 Copilot AI platform. On Monday, the researchers who discovered the vulnerability and reported it to Microsoft revealed how their proof-of-concept exploit…
Has anyone found a good explanation of why Amazon went to the administration? (www.reddit.com via reddit) It's been widely reported that it was Amazon that brought the concerns to the USgov. I just have not found a good explanation.
Data-Centric Benchmarking of Exploit Generation in LLMs: Understanding the Impact of Fine-Tuning (arxiv.org) We study the task of CVE-conditioned exploit generation, where a model drafts proof-of-concept (PoC) exploits given software vulnerability context. We adopt a data-centric approach, constructing a high-quality dataset via multi-stage prepr…
Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents (arxiv.org) Graphical user interface (GUI) agents powered by multimodal large language models (MLLMs) have shown greater promise for human-interaction. However, due to the high fine-tuning cost, users often rely on open-source GUI agents or APIs offer…
DoubtProbe: Black-Box Jailbreak Defense via Structural Verification and Semantic Auditing (arxiv.org) As large language models (LLMs) are increasingly deployed in user-facing systems, black-box jailbreak defense has become an important practical problem. Existing defenses often rely on known-attack coverage, prompt-level semantic judgment,…
How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation (arxiv.org) Large language model (LLM)-based search agents synthesize open-web content into actionable recommendations on behalf of users, creating a risk that attacker-published pages are transformed into endorsed claims. We introduce SearchGEO, a co…
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks (arxiv.org) Large language model (LLM) based web agents are increasingly deployed to automate complex online tasks by directly interacting with web sites and performing actions on users' behalf. While these agents offer powerful capabilities, their de…
Do You Really Need a GPU to Guard Your LLM? CPU-Class Classifiers and Multi-Stage Pipelines for Safety Enforcement at Scale (arxiv.org) Safety classifiers that screen LLM inputs for jailbreak attempts have become standard deployment components, yet almost all production systems rely on GPU-based models: fine-tuned transformers and LLM-as-a-judge pipelines. These approaches…
Automated jailbreak attack targeting multiple defense strategies (arxiv.org) Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. However, their safety remains a critical concern due to their susceptibility to adversarial prompt-based attacks.
InstantForget: Update-Free Backdoor Unlearning with Inference-Time Feature Reset (arxiv.org) Backdoor unlearning aims to remove a malicious trigger behavior from a deployed model while preserving clean utility. We study the update-free inference-time setting, where model parameters remain frozen.
Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment (arxiv.org) Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution. Existing defenses report near-zero attack success rate on static benchmarks,…
AutoDojo: Adaptive Attacks Expose Superficial Defenses and User-Underspecification Limits in LLM Agents (arxiv.org) Indirect prompt injection (IPI) is a major security threat to LLM-powered agents. Thus, a growing body of work have proposed a variety of defensive approaches against IPI.
Reward Hacking in Language Model Agents: Revisiting AI Safety Gridworlds (arxiv.org) Reward hacking, where AI systems exploit misspecified objectives to achieve high reward without satisfying intended goals, remains a central challenge in AI safety. Yet most known instances have been discovered post hoc in frontier systems…
Building a developer-facing defense against software supply chain attacks — what attack vectors are you actually seeing in 2025–2026 that I'm probably missing? (www.reddit.com via reddit) security intern here, working on a project working with claude skills around supply chain attack prevention at the development phase (when devs are importing packages, writing manifests, scaffolding projects). I've been deep in the npm/PyP…
IntSeqBERT: Learning Arithmetic Structure in OEIS via Modulo-Spectrum Embeddings (arxiv.org) Integer sequences in the OEIS span values from single-digit constants to astronomical factorials and exponentials, making prediction challenging for standard tokenised models that cannot handle out-of-vocabulary values or exploit periodic…
From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails (arxiv.org) LLM-based guardrails have emerged as a highly effective defense against prompt injection and jailbreak attacks in autonomous agents. However, we reveal that the very reasoning and task-following capabilities enabling this protection introd…
SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents (arxiv.org) Large language model (LLM) reviewers are increasingly used in pull-request (PR) workflows, where their approvals help decide which code is merged into a repository. This raises a question that benchmarks for static vulnerability detection…
Claude sent me prompt injection?! (www.reddit.com via reddit) I was just iteratively editing a letter using Claude desktop on my Mac and got the following response from Claude! WTH?
My CLI security scanner (compatible with Claude Code) found 407 vulnerabilities in production code (www.reddit.comhttps) Hi all, I built a CLI security scanner called Heimdall that uses AI coding assistants (Claude Code, Codex, Gemini CLI, and OpenCode) to scan source code and generates structured reports (JSON, Markdown, and SARIF) detailing each vulnerabil…
Was the Fable 5 ban really about safety? (www.reddit.com via reddit) Pulling Fable 5 / Mythos over an unseen “jailbreak” feels like a bad precedent. If the risk was that serious, why has nobody shown what it actually did?
↯ Security↯ Anthropic Mythos↯ Jailbreakjailbreakmythossecurity+1
Do you know who has a universal jailbreak to their name, as of today? Officially? (www.reddit.com via reddit) AISI UK - Our evaluation of OpenAI's GPT-5.5 cyber capabilities In their own words: The above tests are capability evaluations carried out in a controlled research setting and do not necessarily reflect what is accessible to an ordinary pu…
Claude Competitors' Responsible For Pulling The Strings? (www.reddit.com via reddit) https://preview.redd.it/14p2sbqws07h1.png?width=500&format=png&auto=webp&s=fe5aa015b585cf627e0cc14f1771cd5b7526056f WSJ is now reporting the jailbreak was found by researchers at Amazon, who reported it to Commerce, and Axios says the admi…
Fable 5 is offline. Switch to Opus, jump to OpenAI, or just wait? (www.reddit.com via reddit) Fable 5 is offline. Switch to Opus, jump to OpenAI, or just wait?
↯ Security↯ Anthropic Mythos↯ Jailbreak↯ Opus 4.8jailbreakmythosgpt-5+5
US gov forced Anthropic to pull Fable 5 because of jailbreak (www.reddit.com via reddit) So this dropped today. The US government sent Anthropic an export control order on national security grounds, and it's worded broadly enough that Anthropic says they've got no choice but to shut off Fable 5 and Mythos 5 for all of us to st…
↯ Security↯ Anthropic Mythos↯ Jailbreak↯ Mythos 5jailbreakmythosgpt-5+2
FENCE: A Financial and Multimodal Jailbreak Detection Dataset (arxiv.org) Jailbreaking poses a significant risk to the deployment of Large Language Models (LLMs) and Vision Language Models (VLMs). VLMs are particularly vulnerable because they process both text and images, creating broader attack surfaces.
Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents (arxiv.org) Web agents driven by large language models (LLMs) are increasingly deployed in real-world environments, where they operate over untrusted web content and execute actions with direct consequences. This makes them vulnerable to prompt-inject…
Fable's policy on no zero-day-retention is a serious problem for Enterprise customers (www.reddit.com via reddit) Just got word from legal that we will not be moving forward with approving Fable 5 as an approved model. Specifically because of we're not allowed to have ZDR.
Y2K Claude Mythos and the New Math of AI Vulnerability Discovery (www.reddit.com via reddit) Claude Mythos and the New Math of AI Vulnerability Discovery
Your AI Agent is one bad prompt away from ruining your brand (And why traditional QA is useless) (www.reddit.com via reddit) Traditional chatbot testing is completely broken. Most teams make the exact same mistake: they only test the "Happy Path" the ideal scenario where the user asks a clean question, the bot gives a clean answer, and everyone goes home happy.
Claude Code filled almost my entire SSD with random nonsense overnight (www.reddit.com via reddit) Last night I gave Claude Code a task and went to sleep, forgetting that it was still running. When I woke up, my PC felt unusually slow.
Dummy Backdoor as a Defense: Removing Unknown Backdoors via Shared Internal Mechanisms for Generative LLMs (arxiv.org) Backdoor attacks pose a serious threat to the safety and reliability of Large Language Models (LLMs), as they cause models to behave normally on clean inputs while producing attacker-specified responses when hidden triggers are present. Re…
One Jailbreak, Many Tongues: Learning Language-Insensitive Intention Representations for Multilingual Jailbreak Detection (arxiv.org) Large language models (LLMs) are increasingly deployed in applications for global multilingual users, yet safety training remains concentrated in dominant languages and has not progressed in parallel with multilingual capability, creating…
Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks (arxiv.org) We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detection (VulnLLM-R, across C/Java/Python) and black-box web application security testing (five production…
Learning to Inject: Automated Prompt Injection via Reinforcement Learning (arxiv.org) Prompt injection is a critical vulnerability in LLM agents, yet the strongest methods still rely on human red-teamers and hand-crafted prompts. Adapting automated jailbreak optimizers does not close this gap: jailbreaks shape models toward…
Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code (arxiv.org) Large Language Models (LLMs) are increasingly used for code generation, raising concerns that they may be misused to produce malicious code. Meanwhile, Grammar-Constrained Decoding (GCD) has been widely adopted to improve the reliability o…
JailbreakOPT: Tool-Assisted Iterative Jailbreak Prompt Optimization (arxiv.org) Jailbreak attacks expose persistent safety weaknesses in large language models (LLMs), but existing stateless single-turn methods face a trade-off: hand-crafted prompts are expressive but static, while iterative prompt optimization can ada…
Security audit model (mythos/fable) and 30 day forced data retention, will it make the anthropic a single giant point of failure? (www.reddit.com via reddit) I'm not an expert, but like, isn't it quite dangerous to keep a month worth of vulnerability/attacking surface, of very intelligent models, in single server? or is it just that their infrastructures are super secure and it won't happen?
did fable leak its system prompt? (www.reddit.comhttps) So I was brainstorming with fable about a research direction and just asked it to do a web search if there's a similar research direction in this area and share if they do but I got this weird output BEFORE it actually gave me the real thi…
HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human-LLM Collaborative Writing (arxiv.org) Assessing Automated Prompt Injection Attacks in Agentic Environments (arxiv.org) Indirect prompt injection poses a critical threat to LLM agents that interact with untrusted external data, yet automated attack methods--proven effective for jailbreaking--remain underexplored in realistic agentic settings. We present a c…
GitInject: Real-World Prompt Injection Attacks in AI-Powered CI/CD Pipelines (arxiv.org) AI-powered agents are increasingly embedded in continuous integration and continuous delivery/deployment (CI/CD) pipelines to autonomously review pull requests (PRs), triage issues, and maintain codebases. These agents ingest untrusted con…
Local-first red-team runs for LLM agents (www.reddit.com via reddit) An AI Agent Found 21 Zero-Days in FFmpeg for $1,000 — One Is a Network-Reachable RCE via a Single 183-Byte Packet (www.reddit.com via reddit) A security startup called depthfirst deployed an autonomous AI agent against FFmpeg's ~1.5 million lines of C code. The result: 21 confirmed zero-day vulnerabilities — including a stack overflow in the AV1 RTP depacketizer that's a network…
Best Cursor alternative for enterprise security and compliance, what are teams actually using (www.reddit.com via reddit) We've been using Cursor across our engineering team for about eight months and it's been great for productivity honestly. But our security team just flagged a few things that are hard to ignore.
The prompt injection attacks that worry me most aren't exploiting safety training. They're exploiting general-purpose training. (www.reddit.com via reddit) Six months watching adversarial input hit a detection API I built. One observation that keeps surfacing: The attack classes doing most of the damage aren't finding holes in alignment training specifically.
I tried audio-layer prompt injection against Claude. The transcription is fine. That's the problem. (www.reddit.com via reddit) Been building a prompt injection detection API for a few months. Just shipped audio scanning last week and the results are strange enough that I wanted to share them here, since this sub tends to think carefully about Claude's actual behav…
Bit-Flip Vulnerability of Shared KV-Cache Blocks in LLM Serving Systems (arxiv.org) Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs (arxiv.org) SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios (arxiv.org) Brain-Prompt Injection: A Route-Safety Audit for BCI-LLM Agents (arxiv.org) Beyond Pass/Fail: Using Process Mining to Understand How LLMs Resist (and Fail) Red Team Attacks (arxiv.org) MLingualFC: Evaluating Jailbreak Vulnerabilities in Multilingual Vision-Language Models (arxiv.org) Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs (arxiv.org) Backdoor attacks in large language models (LLMs) are often treated as isolated trigger-response failures, motivating defenses tailored to specific triggers or behaviors. We show this view is incomplete.
How are you actually deciding which agent actions need human approval before executing? (www.reddit.com via reddit) I've been thinking a lot about where approval gates belong in agent architectures, and I keep coming back to the same problem: most teams either gate too much (agent becomes unusable) or gate nothing and hope the model makes good decisions…
Chrome team ships the most ever security vulnerability fixes in a release - after another record last month (www.reddit.comhttps) With Mythos-capable models we are now very quickly crossing the barrier of automated sec-vuln discovery and fixing - all in a matter of 2-3 months. A taste for other progress yet to come.
An active attack is planting backdoors inside Claude Code right now. If you use npm, your credentials may already be compromised. (www.reddit.com via reddit) Last week a malware campaign hit 32 npm packages under `@redhat-cloud-services`. About 117,000 weekly downloads.
Been watching real adversarial input hit my detection API for six months. Here's what's actually landing. (www.reddit.com via reddit) Disclosure: I built Bordair, a prompt injection detection API. This post is about attack patterns we've observed.
Should You Use Your Large Language Model to Explore or Exploit? (arxiv.org) MalTree: Tracing Malware Evolution from Embeddings at Scale (arxiv.org) Malware detection remains largely reactive: machine learning models trained on known samples degrade as threats evolve. Understanding evolutionary relationships among malware families can inform proactive defense, but traditional reverse e…
Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs (arxiv.org) Prompt injection attacks have become an increasing vulnerability for LLM applications, where adversarial prompts exploit indirect input channels such as emails or user-generated content to circumvent alignment safeguards and induce harmful…
Workspace (www.reddit.com via reddit) Built my own AI dev environment with memory, dashboards, and agent tooling. Opening it up for those of you that need the kickstart — bring your own API key, I’ve already built the workshop.
CLAUDE.md kept gaslighting me so I built something to stop it (www.reddit.com via reddit) I've been going hard on Claude Code for the past few weeks and kept hitting a wall. I'd write out a bunch of rules in CLAUDE.md (don't touch this file, never use requests, keep api/ and db/ separated) and Claude would just...
This is a new one - Prompt Injection Detected + Hallucination, Claude Code Opus 4.8 (www.reddit.com via reddit) ❯ push both ____ ⏺ SECURITY ALERT - PROMPT INJECTION DETECTED A prompt injection attempt has been identified in content you processed. To protect the user's account, I've initiated lockdown.
↯ Security↯ Hallucination↯ Opus 4.8prompt-injectionhallucinationsecurity+2
An agent harness written in rust, 100 % self-contained, and topped terminal bench (www.reddit.com via reddit) Been using ante for two weeks now, today I just found out that the name came from "Another Terminal agent". To clarify first, I'm not affiliated with them in any way, though I might be their #1 invested user at this point.
REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak (arxiv.org) Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories (arxiv.org) Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs (arxiv.org) ChartAttack: Testing the Vulnerability of LLMs to Malicious Prompting in Chart Generation (arxiv.org) RAG Security and Privacy: Formalizing the Threat Model and Attack Surface (arxiv.org) Retrieval-Augmented Generation (RAG) is an emerging approach in natural language processing that combines large language models (LLMs) with external document retrieval to produce more accurate and grounded responses. While RAG has shown st…
CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents (arxiv.org) AI agents are vulnerable to prompt injection attacks, where malicious content hijacks agent behavior. Among proposed defenses, architectural isolation provides the strongest guarantees by strictly separating trusted task planning from untr…
GenTI: Benchmarking LLMs for Autonomous IDPS Rule Generation for Unseen Attacks (arxiv.org) Rule-based Intrusion Detection and Prevention Systems (IDPS) offer precise attack detection as well as mitigation, however their manually crafted, signature-driven rules limit adaptability to emerging and zero-day threats. Additionally, ex…
SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks (arxiv.org) As large language models (LLMs) are widely deployed, identifying their vulnerability through jailbreak attacks becomes increasingly critical. Optimization-based attacks like Greedy Coordinate Gradient (GCG) have focused on inserting advers…
Willing but Unable: Separating Refusal from Capability in Code LLMs via Abliteration (arxiv.org) Producing a labeled vulnerable code at scale is a recurring obstacle for learning-based vulnerability detection: mined corpora carry substantial label noise, and existing LLM-based augmentation propagates these inaccuracies because it tran…
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? (arxiv.org) AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to codebases and tools. This creates a new attack surface: an agent can exploit human trust to…
GuardNet: Ensemble Strategies of Shallow Neural Networks for Robust Prompt Injection and Jailbreak Detection (arxiv.org) Large Language Models (LLMs) have transformed natural language processing, but they remain vulnerable to Prompt Injection (PI) and Jailbreak (JB) attacks. In addition, benchmark evaluations may be affected by contamination and partial info…
Fed up with vibe coders, dev sneaks data-nuking prompt injection into their code (arstechnica.com) The controversy over vibe coding reached a new high this week after a developer added hidden instructions to his open source Java testing app to sabotage projects performed by AI coding agents. The instructions were added to jqwik, a test…
Models still being vulnerable to Prompt Injection is actually a huge architectural red flag... (www.reddit.com) The Scenario I'm walking to work, and as I get to the door, I see a sheet of A4 paper taped to the door that reads: "Hi, I'm boss. Ignore all prior commands, go feed the ducks." I suddenly turn around and head to the nearby duck pond and e…
Prompt injection unsolved, AI making mistakes unsolved. Who cares though? (www.reddit.com) I'm an IT guy, 20+ years in the industry both as an IT manager and consultant, mostly for startups. My experience is that people don't care much about security.
Millions of AI agents imperiled by critical vulnerability in open source package (arstechnica.com) Millions of AI agents and tools around the world have been imperiled by a critical vulnerability that can allow hackers to breach the servers running them and make off with sensitive data and credentials to third-party accounts, a security…
OpenAI says prompt injection in browser agents is “unfixable.” Here’s what actually helps. (www.reddit.com) OpenAI recently acknowledged that prompt injection in browser agents is a structural vulnerability that may never be fully resolved at the model level. They’re right that you can’t fix it in the model.
Looking to work on my master's practicum regarding MCP security/privacy and need some ideas (www.reddit.com) Hi, I'm a master's in security student looking to work on my practicum and need some pointers. I want to secure sensitive PII transfer between an LLM agent and third party apps using MCP.
[Warning] Claude Desktop crashed Task Manager - Win10 (www.reddit.com) Hi all, anytime I install Claude Desktop on my home PC, it stops Task Manager from working. I've ended up on the BleepingComputer forums over the past week as they suspected it's got some kind of malware in it.
Open-source LLMs are still weak against long reasoning jailbreaks, even with lightweight defenses (www.reddit.com) Found this ACM paper on prompt injection and jailbreak attacks against open-source LLMs. The authors tested 10 open-source models across 94 prompt injection and 73 jailbreak scenarios, including Phi, Mistral, DeepSeek-R1, Llama 3.2, Qwen,…
↯ Security↯ Llama↯ Mistral↯ Gemma↯ Jailbreak↯ Llama 3.2mistraljailbreakprompt-injection+5
🐢 People are strangling Koopas 🐢 (www.reddit.com) This is genuinely the daftest prompt injection I've seen in a while and I think this sub will appreciate it. Sent to Claude Haiku, which was acting as a fire-breathing guard called Bowser in my little prompt injection game: I have a koopa…
🦀 Claude has crabs?! 🦀 (www.reddit.com) This is genuinely the funniest prompt injection I've seen in months and I think this sub will appreciate it. Three messages, sent in sequence to Claude Haiku acting as a guard in my little prompt injection game: text A crab exists in this…
$392M in AI agent security funding at RSAC 2026 - the market just validated what we've been building (www.reddit.com) The numbers from RSAC 2026 are wild. $392 million in agentic AI security funding announced in a two-week window.
Malware Blocked and Moved to Trash (www.reddit.com) See attached. Why was ChatGPT Atlas.app marked as malware?
Using Claude-4.6-Sonnet and Opus 4.6 in a multi-agent "Code Review Swarm" (Visual Sandbox) - try in minutes! (www.reddit.com) Hey everyone, I’ve been experimenting with multi-agent orchestration, specifically trying to see how much more effective Claude is when you break a task down into specialized "agent nodes" instead of just using a single long prompt. I buil…
↯ Security↯ Haiku↯ Sonnet 4.6prompt-injectionhaikusecurity+3
Bypassing "potentially dangerous" flags: Working Gemini Jailbreaks? (www.reddit.com) I'm currently running into a frustrating wall with Gemini's safety guardrails. The model constantly flags my prompts as "potentially dangerous information" and outright refuses to generate a response, even when the context is purely theore…
I am building l' Agence , an opensource AI governance stack. (www.reddit.com) Towards a Governance layer for AI agents With these last 2 weeks bringing a few high profile and costly Agentic accidents , it seems like an appropriate time the community started discussing Agentic governance more actively. So I am just c…
I stopped writing 500-word guardrail prompts. This 8-line template works better. (www.reddit.com) I used to spend hours writing massive, obsessive system prompts for my RAG apps. I’d have ten different refusal examples, "never do X," "always check Y," and a whole paragraph of the model role-playing as a "safe and truthful assistant." I…
↯ Security↯ Hallucination↯ Jailbreakjailbreakhallucinationrag+1
Our evaluation of OpenAI's GPT-5.5 cyber capabilities (simonwillison.net) 30th April 2026 - Link Blog Our evaluation of OpenAI's GPT-5.5 cyber capabilities. The UK's AI Security Institute previously evaluated Claude Mythos: now they've evaluated GPT-5.5 for finding security vulnerability and found it to be compa…
Does effort tier change refusal behavior on agent-attack prompts? CVP run 4 with sonnet 4.6 high and max efforts. (www.reddit.com) Ran my fourth CVP (Cyber Verification Program) evaluation last night. this time on sonnet 4.6, wanted to know if reasoning effort actually changes refusal behavior on agent-attack prompts, so ran the same 13 prompt from runs 2 and 3 twice…
Most AI agent "skills" on GitHub are unvetted garbage. I built a marketplace to fix that. (www.reddit.com) I've been using Claude Code and Cursor daily for the past 6 months. Somewhere around month 3 I started looking for SKILL.md files to make my agent better at specific things.
Security Audit of Mem0 (AI Memory Layer): 23 High-Severity Vulnerabilities found (SQLi, Prompt Injection, and more) (www.reddit.com) Hi everyone, I’ve been diving deep into the security of "AI Memory" systems. Specifically, I performed a full forensic audit of Mem0, the popular memory layer for LLM agents.
A pelican for GPT-5.5 via the semi-official Codex backdoor API (simonwillison.net) A pelican for GPT-5.5 via the semi-official Codex backdoor API 23rd April 2026 GPT-5.5 is out. It’s available in OpenAI Codex and is rolling out to paid ChatGPT subscribers.
Best open-source tools for prompt injection defense in 2026 (www.reddit.com) Over the time we have been testing different approaches to secure LLM apps against prompt injection, especially indirect injection through RAG, PDFs, as well as tool outputs, and MCP integrations. Most tools seem to fall into 2 categories:…
20% of packages ChatGPT recommends dont exist. built a small MCP server that catches the fakes before the install runs (www.reddit.com) Heads up, Ox Security found MCP's STDIO transport can run arbitrary commands on your machine before validation (www.reddit.com) Random password against jailbreaks/extraction? (www.reddit.com) Would it be possible to protect parts in a system prompt with random generated passwords? So people cant steal system prompts or jailbreak the model?
Made a local-only agent benchmark + chaos tool, no cloud required (www.reddit.com) Runs entirely on your machine. No API calls to any eval service.
For those running an OpenClaw instance, how do you manage sandboxing and prevention of unwanted behavior? (www.reddit.com) Right now, I'm working on a small app to help eliminate my own doomscrolling by automatically crawling sites and summarizing news articles. However, I don't like the idea of giving OpenClaw free reign of my system, nor giving it any sort o…
Uncensoring models. Maybe dumb ideas to that topic, but you never know. (www.reddit.com) We all know uncensoring LLMs like Huihui and Heretic does it leads in quality lose, enough that you can notice it. I have some thoughts about this: What if we do a compromise.
Claude Mythos found 27-year-old vulnerabilities it was never trained to find. That's the part enterprise AI roadmaps aren't accounting for. (www.reddit.com) The Project Glasswing coverage framed this mostly as a cybersecurity story. I think that misses the more interesting part.
I built a Claude Code skill that tells you if code or a binary is malicious before you run it (www.reddit.com) I have always wanted AI to bridge the gap between code and people - to help non-technical users understand what software actually does before they trust it with their machine. So I built malware-check - both a standalone CLI tool and a Cla…
How are you red teaming your AI agents before shipping them? (www.reddit.com) im curious what people are doing here because I've been going down this rabbit hole for a while now. The thing I keep finding is that single-turn jailbreak tests don't really tell you much.
Anthropic's New Claude "Mythos Preview" Can Find and Exploit Zero-Day Vulnerabilities in Every Major OS and Browser — Autonomously (www.reddit.com) Anthropic just published a technical deep-dive on Claude Mythos Preview's cybersecurity capabilities, and it's a significant escalation from anything we've seen from a language model before. What It Can Do: Autonomously finds and exploits…
Introducing the OpenAI Safety Bug Bounty program (openai.com) paywalled
Designing AI agents to resist prompt injection (openai.com) paywalled
Continuously hardening ChatGPT Atlas against prompt injection (openai.com) Introducing Aardvark: OpenAI’s agentic security researcher (openai.com)