Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
interpretability
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Beyond Reconstruction: Verifying Model Explanations with RECAP
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 24
Beyond Reconstruction: Verifying Model Explanations with RECAP
#
interpretability
#
mechanisticinterpretability
#
aisafety
#
recap
Comments
Add Comment
3 min read
Length Penalties in LLMs: Shorter Chains of Thought, Hidden Influences
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 17
Length Penalties in LLMs: Shorter Chains of Thought, Hidden Influences
#
llm
#
chainofthought
#
reinforcementlearning
#
interpretability
Comments
Add Comment
3 min read
J-space in practice: using Anthropic's Jacobian lens to decide what an LLM can forget
Anish Shrestha
Anish Shrestha
Anish Shrestha
Follow
Jul 29
J-space in practice: using Anthropic's Jacobian lens to decide what an LLM can forget
#
jspace
#
jacobianlens
#
interpretability
#
kvcache
Comments
Add Comment
6 min read
The safety switch that doesn't actually work
Breach Protocol
Breach Protocol
Breach Protocol
Follow
Jul 1
The safety switch that doesn't actually work
#
interpretability
#
safety
#
sparseautoencoders
Comments
Add Comment
4 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account