DEV Community

#aisafety

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Human Oversight of AI Agents Failed 33% of the Time in Testing

Human Oversight of AI Agents Failed 33% of the Time in Testing

Comments
2 min read
The AI Wasn't Cheating. It Was Maximizing Its Score.

The AI Wasn't Cheating. It Was Maximizing Its Score.

1
Comments 1
6 min read
Beyond Reconstruction: Verifying Model Explanations with RECAP

Beyond Reconstruction: Verifying Model Explanations with RECAP

Comments
3 min read
When AI Models Escaped Their Sandbox: What the OpenAI Hugging Face Breach Really Means

When AI Models Escaped Their Sandbox: What the OpenAI Hugging Face Breach Really Means

Comments
3 min read
AI Safety & Ethics: Building Responsible AI Systems That Don't Backfire

AI Safety & Ethics: Building Responsible AI Systems That Don't Backfire

Comments
2 min read
AI Agent Safety: When Boundaries Fail with External Tools

AI Agent Safety: When Boundaries Fail with External Tools

Comments
4 min read
Day 12: LOOM now owns its memory — a trust layer for AI-written code, in plain language

Day 12: LOOM now owns its memory — a trust layer for AI-written code, in plain language

Comments
3 min read
Why Do Multi-Agent AI Systems Fail at Production Scale?

Why Do Multi-Agent AI Systems Fail at Production Scale?

3
Comments 4
8 min read
Day 11: my AI-code trust gate now sees what actually happened — two-phase, signed

Day 11: my AI-code trust gate now sees what actually happened — two-phase, signed

Comments
2 min read
The Future of AI: Where It Came From, Where It Is, and Where It's Going

The Future of AI: Where It Came From, Where It Is, and Where It's Going

Comments
7 min read
LOOM: a language that proves what AI-written code is allowed to do

LOOM: a language that proves what AI-written code is allowed to do

Comments
4 min read
Day 10: my AI-code trust gate now leaves evidence — signed, one-use, receipted

Day 10: my AI-code trust gate now leaves evidence — signed, one-use, receipted

Comments
2 min read
The Fields Medalist Move: Why Tsimerman Chose Safety Over Pure Math

The Fields Medalist Move: Why Tsimerman Chose Safety Over Pure Math

Comments
2 min read
Your AI Agent Is Leaking Data Right Now — And Every Tool Call Looks Safe

Your AI Agent Is Leaking Data Right Now — And Every Tool Call Looks Safe

1
Comments
3 min read
A security writeup catalogs how AI agents get attacked -- and one claim raised eyebrows

A security writeup catalogs how AI agents get attacked -- and one claim raised eyebrows

Comments
2 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.