home.social

#evaluationbenchmarks — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #evaluationbenchmarks, aggregated by home.social.

  1. #OpenAI adjusted several #evaluationbenchmarks for its #GPT6 #Astra model after the initial announcement, with some metrics showing Astra performing better and others worse. The changes, which included altering hallucination rates and maths scores, sparked concerns about #benchmaxxing and the accuracy of #benchmark tests. fortune.com/2026/09/04/openai- #tech #news #ainews