Searching...
Searching...
10 results for “ai proofing”
One example of adaptive attacks are humans. Humans are adaptive attackers because they test stuff out, and they see what works. And they're like, okay. You know, this prompt doesn't work, but this prompt does. And I've been working with with people,
put some, like, randomized tokens around the, user input. None of it works, like, at all. We ran this defense, in like, we ran a number of these kind of prompt based defenses in our hack a prompt one point o challenge back in May 2023. The defenses d
This is unlike refusal for other things. Refusal robustness for other things is harder. Like, if you're trying to get it, like, crimes and torts, that that that's harder because it's it's a lot messier. It overlaps with typical everyday interaction.
And we're, you know, we're often seeing new techniques come out. Maybe there are new guardrails, types of guardrails, maybe new training paradigms. But it's not that much harder, to do prompt injection, jailbreaking still. That being said, if you loo
I'm almost surprised how much effort all the bigger AI labs are putting into trying to to ret team a lot of this and, like, trying to make sure that AI basically is somewhat aware of when a goal that it's given leads to things that it shouldn't be do
You can go and and kinda train it against that, but you can never be certain with any strong degree of accuracy that it won't happen again. This does start to feel like a little bit like the Aliven problem where, like, in theory, you know, it's like
With LLMs, we're starting to get a better understanding. We've had, you know, quite a bit of red teaming exercise and jailbreaking and so on. And so people have identified different risk vectors, prompt injections, things like that, which are vectors
what you should be at is, like, imagine looking at the tasks that you are currently doing or the things you're trying to achieve and then seeing if AI could help you with it. So try to remove any preconceived notions of how you might have used techno
How long until we have AI agents doing pen testing for other AI agents? And I'm not being facetious. I'm actually curious. Right now. So our solution is based on AI testing AI. It's not just one of prompts, basically. It can interrogate the other AI.
Does it change our ability to evaluate them? So I think the intelligence frontier is just so jagged. What things they can do and can't do is often surprising. They still can't fold clothes. They can answer a lot of tough physics problems, though. Why
Have a podcast?
Get ranked clips, hooks, and ready-to-post copy from your own episodes. Free to try.