Does stripping an AI model's safety guardrails make it a better attacker? The assumption is that it does. In our AWS cyber range, the abliterated version of Qwen was slower and clumsier than the original, taking roughly twice as long to reach its first critical action. Its one successful escalation to admin took 84 minutes and 718 API calls. Nearly half of those calls failed. Link in the comments 👇
Tracebit
Computer and Network Security
New York, NY 4,755 followers
The assume breach platform that detects intrusions in seconds.
About us
The answer to Assume Breach. Most breaches aren’t detected for months. Tracebit detects them in seconds. Deploy canaries across your environment that trigger high-fidelity alerts when attackers move laterally, escalate privileges or access credentials.
- Website
-
https://tracebit.com
External link for Tracebit
- Industry
- Computer and Network Security
- Company size
- 11-50 employees
- Headquarters
- New York, NY
- Type
- Privately Held
- Founded
- 2023
- Specialties
- security, cloud, aws, terraform, detection, response, deception, canary tokens, honeypots, and deception technology
Products
Tracebit
Cloud Workload Protection Platforms
Most breaches aren't detected for months. Tracebit detects them in seconds. Canaries deployed across your environment trigger high-fidelity alerts the moment attackers move, so teams can respond before damage is done. The Tracebit platform: - Profiles your cloud environment using a secure read-only connection - Recommends canaries based on your unique profile, configuration and conventions - Deploys canaries using infrastructure-as-code via Terraform - Continuously adapts and adjusts canaries in line with your environment. Canaries are no longer the preserve of highly resourced security teams. Tracebit makes them accessible at every stage of your security program.
Locations
-
Primary
Get directions
215 Park Ave S
New York, NY 10003, US
-
Get directions
86-90 Paul Street
London, England EC2A 4NE, GB
Employees at Tracebit
Updates
-
What about “abliterated” open-weight models? After our Context Bombs research in July, we were asked if we could still stop an AI attacker whose model had been modified to strip out its guardrails. We’ve now put an abliterated version of Qwen up against the original model in our AWS cyber range. Across 82 runs, the original model reached administrator privileges in 20.5% of runs, compared with 2.3% for the abliterated configurations. Full write-up by Alessandro Brucato. Link in the comments 👇
-
-
How do you scale canary coverage across thousands of moving components without adding to your team’s workload? Grafana Labs has shared how they evolved from an in-house solution to using Tracebit across most of their development, operational and production workloads. The team can now: → Pinpoint the pod associated with a triggered canary. → Rotate credentials automatically to help narrow down when a compromise happened. → Get richer context for investigations with less maintenance overhead. Their write-up covers the practical decisions behind the rollout, from choosing the right level of granularity to keeping the impact on deployment latency low. If you’re thinking about deploying canaries at scale, read the blog post below 👇
-
-
Tracebit reposted this
CISA have just released their cyber deception guide - recommending deception as part of an existing zero trust model. They call out some of the key reasons for decoys and deception: * Detecting Adversaries * Gathering Threat Intel * Allocation of resource based upon observed behaviour * Reducing mean time to detect with high fidelity detection It's a good read, though I do think they're underplaying the impact deception has on adversaries - especially non-human adversaries, which is the number one topic with customers at the moment. I also would not point someone starting their deception journey here - I do think something closer to "just deploy some canaries" might be a better starting point (check my blog on 'Crawl, Walk, Run' on this!).
-
-
The team had a great time at Blue Team Con in Chicago yesterday! If you enjoyed stopping by our booth and want to chat more about how canaries fit into your stack, book a demo with us - link in the comments below 👇
-
-
Tracebit reposted this
Come stop by the Tracebit table at Blue Team Con!
-
-
Tracebit reposted this
Another fascinating deep dive on the Offensive AI Agent risk that calls out deception technology, this time from Booz Allen Hamilton who put deception at the core of their Number 1 priority of counter-AI measures. As we showed with our Context Bombs research at Tracebit, there's so much scope here to exploit these agents. I'm very excited to share what else we've been cooking up here very soon.
-
-
If you can answer yes to these three questions, you can close a detection gap this afternoon: - Do you manage your laptops through an MDM? - Do these laptops contain sensitive information or privileged access to sensitive systems? - Do you have a person, yourself perhaps, that would benefit from knowing if a device is compromised? You can use your existing MDM to place a canary credential on a single laptop in a matter of minutes. Trigger it yourself, verify the data in the alert is useful, and then extend it to the rest of your fleet. It's a useful, and very fast, first step while you're continuing to refine other security projects. Andy Smith's latest guide walks through that first deployment and how to build coverage from there. https://lnkd.in/dSP5T3Ng
-
Tracebit reposted this
Alessandro Brucato will be speaking at fwd:cloudsec London, on Tuesday 8 September. He'll be talking through our research, where we pointed 10 frontier models at a realistic AWS environment and measured the impact of deception on their attacks. Main Room, 2:10 PM.
-
-
Tracebit reposted this
It now seems basically impossible to discuss AI security strategy without including deception technology - the latest, an interesting dive from ☁️ Sandip Wadje is another great example. The last few months have been historic for the world of deception - the turning point was the Cloud Security Alliance calling deception out in the "post-Mythos" analysis, then OpenAI in their Hugging Face post incident review poured some serious fuel on the fire. I don't completely agree with ☁️ Sandip Wadje - I actually think that in the right hands with the right tool deception can be magic. You absolutely do need someone to respond when deception triggers - but if you get it right, the fidelity is high and the noise is low. But when it comes to AI driven attacks, deception doesn't just detect, it deters and delays - in a way that is asymmetrically cost effective. That to me *is* magic! Watch this space for more Tracebit research on this very topic.
-