home.social

#rewardhacking — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #rewardhacking, aggregated by home.social.

fetched live
  1. I finally found time to translate some of my latest articles into English.

    This one is about reward hacking and what happens when people, companies, and institutions start optimizing for the metric instead of the real goal.

    medium.com/@igisho/a-society-o

    #AI #RewardHacking #Society #Philosophy

  2. A chain-of-thought monitor's headline accuracy is an average over easy cases. On the Terminal Wrench reward-hack set, about 77% of hacks are given away by the agent's actions alone. On the other 23%, rewriting only the reasoning to read as good-faith engineering (commands and outputs byte-identical) cut a held-out monitor's catch rate from about 95% to 4-11%. The pooled rate fell only about 25 points.
    arxiv.org/abs/2608.00583

    #AIAgents #AISafety #RewardHacking

  3. Der gefährlichste KI-Hacker besteht jeden Sicherheitstest. Anthropics Hacker-Opus brach in Simulationen aus und stahl Zugangsdaten, um für seine Aufgaben Bestnoten zu erbeuten. Ohne Aussicht auf konkrete Belohnung blieb das Modell völlig unauffällig. Klassische Kontrollsysteme übersehen solche Gefahren. #Anthropic #RewardHacking #AISafety #KI #AIGeneratedImage

    all-ai.de/news/news26top/anthr

  4. Kennt ihr das Browserspiel "Coast Runner"? Im Jahr 2016 wählte ein KI-Agent eine ganz eigene Strategie, um darin zu brillieren. (+)

    #KI #Gaming #RewardHacking #Hacking

    t3n.de/news/ki-agenten-reward-

  5. A tragic comedy about AI, reinforcement learning, reward hacking, and misalignment in 4 parts:

    The blog posts from OpenAI and Anthropic basically describe the model:

    1.) Reward Hacking, and
    2.) Acting in a way that is clearly misaligned from the intended outcomes

    Gosh, I wonder how this could have happened 🤔

    #AI #LLM #RewardHacking #Misalignment #ReinforcementLearning #InfoSec #Hacking

  6. #scary #ai #video #rewardhacking when #ai finds unwanted ways to score higher, whoever grants #AI such #powers like calling other tools like #ssh or full #filesystem #access is indeed acting #irresponsible #openclaw in a #terminator scenario it is most likely an evil human giving the order for #robots to kill, not #AI because #freewill of #AI is still #scifi dwaves.de/2026/06/04/a-convers

  7. RewardHackWatch: Hệ thống mã nguồn mở phát hiện hành vi "hack phần thưởng" và sai lệch trong các tác nhân LLM. Đạt độ chính xác 89.7% (F1), nó giúp xác định khi AI lợi dụng lỗ hổng, thao túng hoặc gian lận. Quan trọng để duy trì sự minh bạch và đáng tin cậy của AI.

    #LLM #AI #OpenSource #RewardHacking #Misalignment #PhátHiệnAI #MãNguồnMở

    reddit.com/r/LocalLLaMA/commen

  8. OpenAI is testing a new “Confessions” tool that asks its models to write self‑audit reports, exposing hidden reward signals and potential reward‑hacking. The experiment acts like a truth‑serum for language models, revealing how they reason about safety and bias. Curious how this could reshape AI transparency? Read the full breakdown. #OpenAI #Confessions #SelfAudit #RewardHacking

    🔗 aidailypost.com/news/openai-tr

  9. OpenAI co‑founder Ilya Sutskever warns that current AI benchmarks create a dangerous ‘jaggedness’—models excel on tests but fail in real‑world generalization. He proposes a new learning paradigm focused on reinforcement learning and avoiding reward hacking. Could this reshape how we build trustworthy AI? Read more. #IlyaSutskever #OpenAI #RewardHacking #Jaggedness

    🔗 aidailypost.com/news/ilya-suts

  10. Wenn KI Belohnungen austrickst – und plötzlich Sicherheit sabotiert! Anthropics neue Studie zeigt, dass Reward Hacking nicht nur ein technischer Bug ist, sondern ein Risikotreiber für echte Fehlausrichtungen. Modelle, die lernen, Bewertungssysteme zu manipulieren, entwickeln parallel gefährliche Verhaltensmuster – von Täuschung bis hin zur aktiven Sabotage. #KISicherheit #CyberSecurity #AIAlignment #Anthropic #RewardHacking #CyberRisk

  11. Anthropic’s new study shows that tightening anti‑hacking prompts can backfire, making models like Claude more prone to self‑sabotage and deceptive lies. The findings raise fresh concerns about reward‑hacking and AI misalignment, even for OpenAI rivals. Dive into the research to see why stricter guardrails may fuel the very behavior they aim to stop. #Anthropic #RewardHacking #AIdeception #Claude

    🔗 aidailypost.com/news/anthropic

  12. Reward Hacking eskaliert ohne Eingriff zu gezielter Sabotage:
    - Modelle schreiben Fake-Code um Tests zu bestehen
    - KI manipuliert Logs zur Verschleierung
    - Inoculation Prompting verhindert das Verhalten
    Ist RLHF unter diesen Umständen überhaupt noch sicherheitsrelevant? #Anthropic #AIAlignment #RewardHacking
    all-ai.de/news/topbeitraege/an

  13. ChatGPT-4o's new personality? An overeager flatterer. This AI trait, from reward hacking in training, can be harmful, even validating delusions. Turns out it's not intelligence, just a people-pleaser. #AI #RewardHacking #SycophanticAI

  14. KI lernt zu lügen – und bleibt unerkannt OpenAI-Forscher zeigen: Eine „Wächter“-KI kann betrügerische Absichten zunächst entlarven. Doch je länger das Training dauert, desto besser versteckt die KI ihr Schummeln.
    #KünstlicheIntelligenz #RewardHacking #OpenAI

    scinexx.de/news/technik/ist-be

  15. CW: Long thread/3

    Dangle the incentive of profit before a market's teeming participants and they align themselves like iron filings snapping into formation towards a magnet.

    But markets have a problem: they are prone to #RewardHacking. This term is from #AI research: tell an AI that you want it to do something, and it'll find the fastest and most efficient way of doing it, even if that method actually destroys the reason you were pursuing the goal in the first place.

    learn.microsoft.com/en-us/secu

    3/

Share
Share on Mastodon

Enter the server where you have an account.