home.social

#interpretability — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #interpretability, aggregated by home.social.

  1. Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?

    benjaminhan.net/posts/20260927

    #Paper #AI #Interpretability

  2. Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?

    benjaminhan.net/posts/20260927

    #Paper #AI #Interpretability

  3. Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?

    benjaminhan.net/posts/20260927

    #Paper #AI #Interpretability

  4. Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?

    benjaminhan.net/posts/20260927

    #Paper #AI #Interpretability

  5. Can we read what #LLMs are thinking but not saying? A paper from #Anthropic finds a small set of word-linked patterns inside #Claude that it can report, control on request, and use to reason silently, like the global workspace theory in #neuroscience. How do these findings connect to #LatentReasoning, #AISafety, and #KnowledgeEditing?

    benjaminhan.net/posts/20260927

    #Paper #AI #Interpretability