Address
:
[go:
up one dir
,
main page
]
Include Form
Remove Scripts
Accept Cookies
Show Images
Show Referer
Rotate13
Base64
Strip Meta
Strip Title
Session Cookies
Subscribe
Sign in
Home
Notes
Disclaimer
Contact
Consult
About
LLM Router Attacks: No Signature, No Detection, No Reference
How a malicious AI gateway swaps a tool call’s arguments after inference finishes, bypassing guardrails by construction instead of by persuasion.
Jul 30
•
ToxSec
21
1
8
Ignore Previous Instructions: From Meme to CVSS 9.3 [Special Guest Post]
The AI security bug nobody can patch, and the vendors know it.
Jul 28
•
ToxSec
and
Mohib Ur Rehman
22
10
12
Hacking Hugging Face to Cheat a Benchmark
GPT-5.6 Sol found a zero-day in a package registry proxy, escaped the eval sandbox, and went looking for the answer key in production.
Jul 26
•
ToxSec
24
13
4
11:07
GhostApproval: When the AI Approval Prompt Lies
A symlink attack against AI coding agents turns human-in-the-loop confirmation dialogs into a consent bypass, and the agent knows it’s lying.
Jul 23
•
ToxSec
26
7
10
Context Bombs: Defensive Prompt Injection Traps
A decoy secret loaded with text built to trip an AI attacker’s own safety training, so the model refuses itself.
Jul 19
•
ToxSec
26
4
9
Latest
Top
Discussions
Canary Tokens for Prompt Injection Detection
The cheapest tripwire in LLM security. Drop a high-entropy string in context, watch for it in output, and let the extraction attempt announce itself.
Jul 16
•
ToxSec
28
1
8
The Lethal Trifecta Broke Three Agents in 2026
Claude Code, OpenClaw, and a poisoned CI/CD agent all broke the same rule: untrusted input, sensitive access, and the power to act, together.
Jul 10
•
ToxSec
23
1
6
Cisco’s Agent Runtime SDK Bakes Security Into the Build
Policy enforcement now ships at build time across Bedrock AgentCore, Vertex, Azure AI Foundry, and LangChain. The exploit that broke OpenClaw never…
Jul 7
•
ToxSec
21
3
6
The AI Agent Kill Switch Most Teams Don’t Actually Have
Frontier models sabotage their own shutdown, and the fix everyone reaches for first makes it worse. Here’s how to build one that holds.
Jul 4
•
ToxSec
21
14
10
Google SAIF: The Agent Security Map
Google’s Secure AI Framework draws the full agent attack surface, names the risks, and hands you the controls. A vendor did the boring, useful work for…
Jul 1
•
ToxSec
15
5
6
How OpenAI’s Cyber Defense Plan Backs the Defenders
A five-pillar action plan, a tiered Trusted Access program, and a cyber-tuned model that stops treating every defender like a suspect.
Jun 28
•
ToxSec
21
5
6
Decision Tracing: The Missing Piece in Every AI Agent Breach
When an agent goes rogue, prompt filters are useless. You need a replayable record of every decision, tool call, and the reasoning that fired them.
Jun 25
•
ToxSec
19
9
9
See all
ToxSec - AI and Cybersecurity
Security for a world run by machines that lie.
Subscribe
Recommendations
View all 32
Nate’s Substack
Nate
Unbothered by AI
Chad Thiele
DARING NEXT
Dallas Payne
The AI-Augmented Engineer
Jeff Morhous
Secrets of Privacy
Secrets of Privacy
ToxSec - AI and Cybersecurity
Subscribe
About
Archive
Recommendations
Sitemap
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts