A CI/CD Red Team Framework for demonstrating Build Pipeline security risks.
#red-team
24 items
Show HN: SmokedMeat, like Metasploit, but for CI/CD (open-source) (github.com via hn) Show HN: Red-team LLM reasoning and agent actions (honest scoring, local-first) (github.com via hn) CoT Red Team Agent Refusal quotes of a canary are not a finding. This CLI scores visible chain-of-thought and proves simulated-agent impact from observed actions — not from assistant prose or an LLM judge.
Anthropic has a Red Team page (red.anthropic.com via hn) Welcome to red.anthropic.com, the home for research from Anthropic’s Frontier Red Team (and occasionally other teams at Anthropic) on what frontier AI models mean for national security. We provide evidence-based analysis about AI’s implica…
Claude Plays Robotics (www.anthropic.com via hn) Subscribe to the Frontier Red Team newsletter Get updates on our latest red-teaming research and findings. Shmuel Berman, Michael Ilie, Jia Deng, and Daniel Freeman Do language models’ strengths transfer to robotics, a domain which require…
Using Claude as the Lead agent in a multi-agent security team (www.reddit.com) Building a hierarchical agent system where Claude (via API) acts as the Lead agent coordinating specialist sub-agents. Wanted to share what's working on the synthesis prompt since this is where most of the value comes from.
SmokedMeat: A Red Team Tool to Hack Your Pipelines First (labs.boostsecurity.io via hn) SmokedMeat: A Red Team Tool to Hack Your Pipelines First TL;DR: In March 2026, TeamPCP unleashed mayhem on the software supply chain: compromising Trivy, LiteLLM, KICS, Telnyx, and dozens of npm packages, proving that CI/CD pipelines are t…
A $37 GLM 5.3 red team: the Alloy-modeled auth layer held, but two bugs outside (goodmem.ai via hn) Red-teaming GoodMem with GLM 5.3 We used GLM 5.3 to red-team GoodMem. How we defined the tests, what the agent found, what we fixed, and how we verified the fixes.
The Evolving Role of the Red Team in the Era of Agentic Security (blog.google via hn) The Evolving Role of the Red Team in the Era of Agentic Security At Google, our Red Teams have always operated on the cutting edge of security. We’ve shared our journey in the past: from the high-stakes operations showcased in our Hacking…
Show HN: A sandbox for running real jailbreak techniques against local LLMs (github.com via hn) LLM Red Team Lab A hands-on kit for educational, authorized red teaming of any locally-run LLM. It works with any OpenAI-compatible model — Llama, Mistral, Qwen, Gemma, DeepSeek R1, and more — and covers the two ways an LLM system gets exp…
↯ Security↯ Llama↯ Mistral↯ Gemma↯ Jailbreakred-teammistraljailbreak+6
Show HN: Vulnsy – A platform for vulnerability management and reporting (www.vulnsy.com via hn) I've spent over 10 years doing penetration tests and red team engagements, and one thing that always seemed to take far longer than it should was reporting. Most reporting platforms do a great job of managing reusable findings, but I still…
T3MP3ST autonomous red team platform multi-agent offensive-security meta-harness (github.com via hn) 🌩️ T3MP3ST ▄▄▄█████▓▓█████ ███▄ ▄███▓ ██▓███ ▓█████ ██████ ▄▄▄█████▓ ▓ ██▒ ▓▒▓█ ▀ ▓██▒▀█▀ ██▒▓██░ ██▒▓█ ▀ ▒██ ▒ ▓ ██▒ ▓▒ ▒ ▓██░ ▒░▒███ ▓██ ▓██░▓██░ ██▓▒▒███ ░ ▓██▄ ▒ ▓██░ ▒░ ░ ▓██▓ ░ ▒▓█ ▄ ▒██ ▒██ ▒██▄█▓▒ ▒▒▓█ ▄ ▒ ██▒░ ▓██▓ ░ ▒██▒ ░ ░▒████…
Claude Fable 5: What Our Red Team Found Before the Plug Got Pulled (www.reco.ai via hn) Inside Claude Fable 5: What Our Red Team Found Before the Plug Got Pulled Reco AI Research — June 14, 2026 When Anthropic shipped Claude Fable 5 on June 9, it was pitched as something different from the rest of the Claude line. Not a chat…
AICU – LLM Red Team Vulnerability Scanner (github.com via hn) AICU Black-box security scanner for LLM applications. Point it at any chat endpoint, get a report of what leaks.
Chaining LLM and web bugs to Admin (blog.quarkslab.com via hn) During a Red Team exercise we were able to chain multiple LLM and web-based vulnerabilities to achieve admin account takeover from a low-privileged account. Trusting the LLM turned out to be the first falling domino of a long chain of even…
Show HN: Z3r0 – Multi-agent red team collaboration platform (github.com via hn) English · 中文 Architecture · Agent Team · Runtime Model · Deployment · Quickstart :warning: Legal Notice This project may be used only within a lawful and explicitly authorized scope for security testing, assessment, and research. Any unaut…
Netgear Nighthawk RS700S: Red Team Level1Diagnostic (forum.level1techs.com via hn) Preview of the Netgear RS700S. I would also submit that Netgear deleting ALL the GPL links: … they know how bad it is.
Show HN: SuperVoiceMode universal voice layer for AI-assisted development (voicemode.io via hn) I wanted to see if I could one-shot build a dictation tool for my own use. I built it.
Self-hosted red team workspace (github.com via hn) RootNotes RootNotes is a self-hosted red team workspace for tracking projects, notes, hosts, credentials, findings, loot, objectives, scope, and attack paths in one interface. The project in this repository is split into: frontend/: React…
Free Red Team Security Audit for AI Agents & RAG Systems (limited) (www.reddit.com) I'm developing a specialized Red Team audit framework focused on real-world AI agent and RAG security risks (prompt injection, tool misuse, excessive agency, indirect injection through documents, memory poisoning, etc.). I’m looking for a…
Building the first AI Red Team OS – mythosai.cloud – early access open (mythosai.cloud via hn) SYSTEM INITIALIZING... STAND BY MYTHOSAI THE FIRST RED TEAM OPERATING SYSTEM "" AI-Native Core Red Team Ready Adversarial Engine Zero Trust Architecture OPSEC First Post-Exploitation C2 Integration Evasion Layer Threat Intelligence Request…
Symbolic Attack Chain Generation from Atomic Red Team Techniques: An Empirical Study of Predicate Representation Granularity (arxiv.org) Automated attack chain generation is critical for modern cybersecurity, yet manual construction fails to scale as adversary behaviors expand. While classical AI planning using PDDL offers a formal method to automate this process, it relies…
Cybersecurity program GPT 5.5 Cyber - From Costa Rica (www.reddit.com via reddit) Hello team, I'm a security specialist researcher from latinamerica and I'm mostly on my own because my team is building applications to follow track on soccer matches lol. The thing is that I have been a blue team guy for 7+ years and now…
Beyond Pass/Fail: Using Process Mining to Understand How LLMs Resist (and Fail) Red Team Attacks (arxiv.org) I am building l' Agence , an opensource AI governance stack. (www.reddit.com) Towards a Governance layer for AI agents With these last 2 weeks bringing a few high profile and costly Agentic accidents , it seems like an appropriate time the community started discussing Agentic governance more actively. So I am just c…