#evaluationbenchmarks — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #evaluationbenchmarks, aggregated by home.social.
-
#OpenAI adjusted several #evaluationbenchmarks for its #GPT6 #Astra model after the initial announcement, with some metrics showing Astra performing better and others worse. The changes, which included altering hallucination rates and maths scores, sparked concerns about #benchmaxxing and the accuracy of #benchmark tests. https://fortune.com/2026/09/04/openai-quietly-boosts-some-of-astras-evaluation-metrics-amid-rare-delay-in-publication-of-the-modeblog-post-announcement/?eicker.news #tech #news #ainews