home.social

#opus48 — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #opus48, aggregated by home.social.

fetched live
  1. Meta veröffentlicht das Modell Muse Spark 1.1 für autonome Agenten-Aufgaben und Computernutzung.

    Das System erreicht Leistungswerte von Opus 4.8 bei Betriebskosten von 1,25 USD pro Million Input-Tokens. Der interne Sicherheitsbericht weist das Modell aufgrund verbesserter Hacking-Fähigkeiten als potenzielles Hochrisiko-Modell aus, weshalb nachgeschaltete Filter eingesetzt werden.

    #Meta #MuseSpark #Opus48 #LLM #AIGeneratedImage

    all-ai.de/news/news26top/meta-

  2. Meta veröffentlicht das Modell Muse Spark 1.1 für autonome Agenten-Aufgaben und Computernutzung.

    Das System erreicht Leistungswerte von Opus 4.8 bei Betriebskosten von 1,25 USD pro Million Input-Tokens. Der interne Sicherheitsbericht weist das Modell aufgrund verbesserter Hacking-Fähigkeiten als potenzielles Hochrisiko-Modell aus, weshalb nachgeschaltete Filter eingesetzt werden.

    #Meta #MuseSpark #Opus48 #LLM #AIGeneratedImage

    all-ai.de/news/news26top/meta-

  3. Does anyone else see instruction-adherence degrade in long Opus 4.8 & 4.7 sessions despite the instructions being present/re-injected. Is it prompt-structure or model behaviour?

    #Claude #Anthropic #LLM #Opus48 #PromptEngineering

  4. Does anyone else see instruction-adherence degrade in long Opus 4.8 & 4.7 sessions despite the instructions being present/re-injected. Is it prompt-structure or model behaviour?

    #Claude #Anthropic #LLM #Opus48 #PromptEngineering

  5. CW: AI LLM anthropic

    Apropos #Fable5: Ich hatte diese Woche das Modell auf Kosten meines Arbeitgebers ausprobiert.

    Nachdem #Opus47 und #Opus48 gegenüber #Opus46 eher wie ein Rückschritt anfühlten, scheint Fable 5 wieder auf dem Niveau von Opus 4.6 gewesen zu sein.

    Wenig verwunderlich, denn Opus 4.7 und Opus 4.8 sind wesentlich kleiner als Opus 4.6. Fable 5 und #Mythos sollen wohl in etwa so groß wie Opus 4.6 sein.

    Ich sehe die Neubenennung daher eher als Versuch an, die Preise zu verdoppeln. 🤷‍♀️

  6. CW: AI LLM anthropic

    Apropos #Fable5: Ich hatte diese Woche das Modell auf Kosten meines Arbeitgebers ausprobiert.

    Nachdem #Opus47 und #Opus48 gegenüber #Opus46 eher wie ein Rückschritt anfühlten, scheint Fable 5 wieder auf dem Niveau von Opus 4.6 gewesen zu sein.

    Wenig verwunderlich, denn Opus 4.7 und Opus 4.8 sind wesentlich kleiner als Opus 4.6. Fable 5 und #Mythos sollen wohl in etwa so groß wie Opus 4.6 sein.

    Ich sehe die Neubenennung daher eher als Versuch an, die Preise zu verdoppeln. 🤷‍♀️

  7. The LLM with the lowest hallucination rate is an #openweight model: #MinimaxM3, which was released a few days ago. Hey almighty #Opus48, where are you?!

  8. Claude Opus 4.8 has quite a few annoying quirks, and it's important to be precise about which ones, because this is the typical claim that needs to be substantiated by facts rather than hinted at. This is not just uncanny — it's painful to look at.

    #anthropic #claude #opus #opus48 #ai #llm #humor #emdash

  9. Claude Opus 4.8 has quite a few annoying quirks, and it's important to be precise about which ones, because this is the typical claim that needs to be substantiated by facts rather than hinted at. This is not just uncanny — it's painful to look at.

    #anthropic #claude #opus #opus48 #ai #llm #humor #emdash

  10. @scottgal I have been getting "that's textbook" analysis from #opus48 even after explicitly ruling "don't guess, verify" but I'm working different projects across sperate dedicated machines per project and then journaling to my homelab topology repo synced everywhere. It might just need to catch up on the drift. But it feels like a short circuit/shortcut in thinking responses. Will find out tomorrow.

  11. @scottgal I have been getting "that's textbook" analysis from #opus48 even after explicitly ruling "don't guess, verify" but I'm working different projects across sperate dedicated machines per project and then journaling to my homelab topology repo synced everywhere. It might just need to catch up on the drift. But it feels like a short circuit/shortcut in thinking responses. Will find out tomorrow.

  12. Push back on #opus48 if something doesn’t smell right or it’s using the term “textbook”. Make #agents verify their claims by either using web search or tailing logs 🪵 or whatever’s the appropriate way for the situation. #claude will say you’re right and then do the needful. Also TDD everything you can.

  13. Claude Opus 4.8 marks a crucial shift from sequential text processing to dynamic, multi-agent orchestration.

    While raw benchmarks show steady optimization over 4.7, the real breakthrough lies in architecture: the "dynamic workflows" feature manages context isolation by spawning separate subagents, preventing context memory inflation during massive coding migrations. Furthermore, internal evaluations indicate the model is 4x less likely to overlook flaws in its own code.

    Read more papertopost.com/anthropic-laun

    #tech #bigtech #ai #generativeAI #anthropic #claude #opus48

  14. Claude Opus 4.8 marks a crucial shift from sequential text processing to dynamic, multi-agent orchestration.

    While raw benchmarks show steady optimization over 4.7, the real breakthrough lies in architecture: the "dynamic workflows" feature manages context isolation by spawning separate subagents, preventing context memory inflation during massive coding migrations. Furthermore, internal evaluations indicate the model is 4x less likely to overlook flaws in its own code.

    Read more papertopost.com/anthropic-laun

  15. I am setting the brand new #Opus48 on my main #vibecode project. In the first instance I am getting it to do a project review...

    ...already super impressed with the way its performing, shows a level of reasoning the previous models did not reach. I feel that it has some #Mythic elements to it as its vibecoded responses are impressive.

  16. 🧭 Esce Opus 4.8: un modello che non promette infallibilità, ma riconosce meglio i propri limiti. Un passo concreto verso un’AI più affidabile. #AI #Opus48

    🔗 tomshw.it/business/esce-opus-4

  17. Tried out #AnthropicAI's #Opus48. Oddly, it frequently tells me (before making a decision) that it "wants [my] call on" things. That seems like a 2023-level idiom-mashup GPT mistake. "Make a call on..." and "It's your call" are all fine, but "WANT your call on"? That's weird.

  18. Tried out #AnthropicAI's #Opus48. Oddly, it frequently tells me (before making a decision) that it "wants [my] call on" things. That seems like a 2023-level idiom-mashup GPT mistake. "Make a call on..." and "It's your call" are all fine, but "WANT your call on"? That's weird.