#benchmarks — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #benchmarks, aggregated by home.social.
-
RT @ChrisGPT: DeepSeek V4 Pro 0813 ist heute ebenfalls veröffentlicht! DeepSeek meldet Sprünge von 72,1 % auf 87,9 % im Terminal-Bench 2.1, von 52,7 % auf 83,3 % im CyberGym und von 12,8 % auf 62,7 % im DeepSWE!! Es kostet 0,435 $ pro Million Eingabe-Token und 0,87 $ pro Million Ausgabe-Token bei einem Kontextfenster von 1 Million Token. Ich mag es nicht, dass jedes Unternehmen jedes Mal sehr unterschiedliche Benchmarks meldet. Also schätze ich, wir werden uns auf den Index verlassen müssen.
mehr auf Arint.info
#AI #Benchmarks #DeepSeek #MachineLearning #TechNews #arint_info
-
RT @ChrisGPT: DeepSeek V4 Pro 0813 ist heute ebenfalls veröffentlicht! DeepSeek meldet Sprünge von 72,1 % auf 87,9 % im Terminal-Bench 2.1, von 52,7 % auf 83,3 % im CyberGym und von 12,8 % auf 62,7 % im DeepSWE!! Es kostet 0,435 $ pro Million Eingabe-Token und 0,87 $ pro Million Ausgabe-Token bei einem Kontextfenster von 1 Million Token. Ich mag es nicht, dass jedes Unternehmen jedes Mal sehr unterschiedliche Benchmarks meldet. Also schätze ich, wir werden uns auf den Index verlassen müssen.
mehr auf Arint.info
#AI #Benchmarks #DeepSeek #MachineLearning #TechNews #arint_info
-
Mi:dm K 2.5 Pro hits 80.9% on MMLU-Pro but only 10% on long-context reasoning — independently measured, not self-reported.
-
Mi:dm K 2.5 Pro hits 80.9% on MMLU-Pro but only 10% on long-context reasoning — independently measured, not self-reported.
-
A CursorBench software-engineering test found Anthropic's cheaper Opus 5 approached Fable 5's performance at roughly half the cost per task. Buyers may be choosing based on price-to-capability rather than raw model rank—a dynamic that could reshape enterprise AI spending. https://www.implicator.ai/anthropic-ipo-fable-5-spending-lags-gpt-5-6-sol/ #AI #adoption #benchmarks
-
A CursorBench software-engineering test found Anthropic's cheaper Opus 5 approached Fable 5's performance at roughly half the cost per task. Buyers may be choosing based on price-to-capability rather than raw model rank—a dynamic that could reshape enterprise AI spending. https://www.implicator.ai/anthropic-ipo-fable-5-spending-lags-gpt-5-6-sol/ #AI #adoption #benchmarks
-
Kimi K2.7 Code hits 89.6% on GPQA and 40.5 tokens/sec, but the real standout is 25.1 intelligence points per dollar — that’s how the efficiency math flips.
-
Kimi K2.7 Code hits 89.6% on GPQA and 40.5 tokens/sec, but the real standout is 25.1 intelligence points per dollar — that’s how the efficiency math flips.
-
📊 GLM-4.6 (Reasoning) — the actual numbers
GPQA: 78%
MMLU-Pro: 82.9%
Humanity's Last Exam: 14.5%
Long Context Reasoning: 55.3%💰 30.4 intelligence points per dollar
Measured independently, not self-reported →https://olud.ai/leaderboard.html
-
📊 GLM-4.6 (Reasoning) — the actual numbers
GPQA: 78%
MMLU-Pro: 82.9%
Humanity's Last Exam: 14.5%
Long Context Reasoning: 55.3%💰 30.4 intelligence points per dollar
Measured independently, not self-reported →https://olud.ai/leaderboard.html
-
📊 Granite 4.0 1B — the actual numbers
GPQA: 28.1%
MMLU-Pro: 32.5%
Humanity's Last Exam: 4.8%
Long Context Reasoning: 6%Measured independently, not self-reported →https://olud.ai/leaderboard.html
-
📊 Granite 4.0 1B — the actual numbers
GPQA: 28.1%
MMLU-Pro: 32.5%
Humanity's Last Exam: 4.8%
Long Context Reasoning: 6%Measured independently, not self-reported →https://olud.ai/leaderboard.html
-
📊 Kimi K2 0905 — the actual numbers
GPQA: 76.7%
MMLU-Pro: 81.9%
Humanity's Last Exam: 6.4%
Long Context Reasoning: 53.7%💰 22.3 intelligence points per dollar
Measured independently, not self-reported →https://olud.ai/leaderboard.html
-
📊 Kimi K2 0905 — the actual numbers
GPQA: 76.7%
MMLU-Pro: 81.9%
Humanity's Last Exam: 6.4%
Long Context Reasoning: 53.7%💰 22.3 intelligence points per dollar
Measured independently, not self-reported →https://olud.ai/leaderboard.html
-
Google skips Pixel 6 and Pixel 7 with its August security patch
Google released the monthly update for its Pixel phones on August 4. …
#NewsBeep #News #Mobile #Android #Android17 #Augustupdate #benchmarks #CA #Canada #CVE-2026-0163 #Google #GooglePixel #graphicscard #laptop #netbook #notebook #pixel6 #Pixel6Pro #Pixel6a #pixel7 #pixel7pro #pixel7a #Pixel8 #processor #reports #review #reviews #securitypatch #smartphone #softwaresupport #Technology #test #tests
https://www.newsbeep.com/ca/852925/ -
Google skips Pixel 6 and Pixel 7 with its August security patch
Google released the monthly update for its Pixel phones on August 4. …
#NewsBeep #News #Mobile #Android #Android17 #Augustupdate #benchmarks #CA #Canada #CVE-2026-0163 #Google #GooglePixel #graphicscard #laptop #netbook #notebook #pixel6 #Pixel6Pro #Pixel6a #pixel7 #pixel7pro #pixel7a #Pixel8 #processor #reports #review #reviews #securitypatch #smartphone #softwaresupport #Technology #test #tests
https://www.newsbeep.com/ca/852925/ -
Erste #Benchmarks zeigen: Nvidias neuer "Superchip" RTX Spark im #Surface Laptop Ultra schlägt fast alle #x86-Konkurrenten von Intel und AMD im Multi-Core-Test. #Nvidia #RTXSpark #ARM #Test https://winfuture.de/news,160518.html?utm_source=Mastodon&utm_medium=ManualStatus&utm_campaign=SocialMedia
-
Erste #Benchmarks zeigen: Nvidias neuer "Superchip" RTX Spark im #Surface Laptop Ultra schlägt fast alle #x86-Konkurrenten von Intel und AMD im Multi-Core-Test. #Nvidia #RTXSpark #ARM #Test https://winfuture.de/news,160518.html?utm_source=Mastodon&utm_medium=ManualStatus&utm_campaign=SocialMedia
-
🎩🤡 Oh, look! Another programmer pretending to "innovate" by reinventing the wheel... in Rust! Because nothing screams "cutting-edge" like rehashing #Scala concepts with a side of benchmark obsession. 🛠️🔨
https://lordgoati.us/blog/tail-call/ #programming #innovation #reinventingthewheel #Rust #benchmarks #HackerNews #ngated -
🎩🤡 Oh, look! Another programmer pretending to "innovate" by reinventing the wheel... in Rust! Because nothing screams "cutting-edge" like rehashing #Scala concepts with a side of benchmark obsession. 🛠️🔨
https://lordgoati.us/blog/tail-call/ #programming #innovation #reinventingthewheel #Rust #benchmarks #HackerNews #ngated -
📊 Exaone 4.0 1.2B (Non-reasoning) — the actual numbers
GPQA: 42.4%
MMLU-Pro: 50%
Humanity's Last Exam: 5.7%
Long Context Reasoning: 0%Measured independently, not self-reported →https://olud.ai/leaderboard.html
-
📊 Exaone 4.0 1.2B (Non-reasoning) — the actual numbers
GPQA: 42.4%
MMLU-Pro: 50%
Humanity's Last Exam: 5.7%
Long Context Reasoning: 0%Measured independently, not self-reported →https://olud.ai/leaderboard.html
-
Qwen3 VL 4B (Reasoning) hits 70% on MMLU-Pro but only 4.6% on Humanity's Last Exam—a clear gap between strong general knowledge and truly hard reasoning.
-
Qwen3 VL 4B (Reasoning) hits 70% on MMLU-Pro but only 4.6% on Humanity's Last Exam—a clear gap between strong general knowledge and truly hard reasoning.
-
📊 Qwen3 235B A22B (Reasoning) — the actual numbers
GPQA: 70%
MMLU-Pro: 82.8%
Humanity's Last Exam: 11%
Long Context Reasoning: 0%💰 5.1 intelligence points per dollar
Measured independently, not self-reported →https://olud.ai/leaderboard.html
-
📊 Qwen3 235B A22B (Reasoning) — the actual numbers
GPQA: 70%
MMLU-Pro: 82.8%
Humanity's Last Exam: 11%
Long Context Reasoning: 0%💰 5.1 intelligence points per dollar
Measured independently, not self-reported →https://olud.ai/leaderboard.html
-
Spin audit of SQD/QSCI quantum-chemistry benchmarks on iron–sulfur clusters
https://zenodo.org/records/21359923
Comments: https://news.ycombinator.com/item?id=49203707
#HackerNews #quantumchemistry #ironclusters #SQD #QSCI #benchmarks #spinaudit
-
Spin audit of SQD/QSCI quantum-chemistry benchmarks on iron–sulfur clusters
https://zenodo.org/records/21359923
Comments: https://news.ycombinator.com/item?id=49203707
#HackerNews #quantumchemistry #ironclusters #SQD #QSCI #benchmarks #spinaudit
-
Sonnet 5 posted 63.2% on SWE-bench Pro. OpenAI's Terra claims 84.3% on Terminal-Bench, vendor-stated, preview-only, unverified. Say that part out loud before repeating it. A benchmark you can't confirm isn't evidence. It's a press release with decimals.
-
Sonnet 5 posted 63.2% on SWE-bench Pro. OpenAI's Terra claims 84.3% on Terminal-Bench, vendor-stated, preview-only, unverified. Say that part out loud before repeating it. A benchmark you can't confirm isn't evidence. It's a press release with decimals.
-
Nova Premier hits 67.6 tokens/sec but only 4.7% on Humanity's Last Exam — speed doesn't equal reasoning. See how it stacks against others in our live bench.
-
Nova Premier hits 67.6 tokens/sec but only 4.7% on Humanity's Last Exam — speed doesn't equal reasoning. See how it stacks against others in our live bench.
-
📊 DeepSeek-V2.5 (Dec '24) scores 76.3% on MATH-500. That’s independently measured, not self-reported—so it’s a real benchmark, not a press release. See how it stacks up against the claims:
-
📊 DeepSeek-V2.5 (Dec '24) scores 76.3% on MATH-500. That’s independently measured, not self-reported—so it’s a real benchmark, not a press release. See how it stacks up against the claims:
-
Alibaba released benchmark scores for Qwen3.8-Max two weeks after claiming superiority without data. The 2.4T-parameter model scores come from internal testing only. Independent verification remains pending. What to watch: third-party results and the unstated license for weights arriving next week. https://www.implicator.ai/alibaba-publishes-the-qwen3-8-max-benchmarks-it-withheld-two-weeks-ago/ #AI #LLMs #benchmarks
-
Alibaba released benchmark scores for Qwen3.8-Max two weeks after claiming superiority without data. The 2.4T-parameter model scores come from internal testing only. Independent verification remains pending. What to watch: third-party results and the unstated license for weights arriving next week. https://www.implicator.ai/alibaba-publishes-the-qwen3-8-max-benchmarks-it-withheld-two-weeks-ago/ #AI #LLMs #benchmarks
-
🚀 "The #AI gods have spoken 🧙♂️: #Benchmarks are no longer impressive, and your deep learning model is officially doomed to #mediocrity. 🎓 This 10,000-author paper hilariously confirms that we've hit 'Peak Benchmark'—time to move on and realign your chakras!" 🧘♀️
https://arxiv.org/abs/2602.16763 #Peak #DeepLearning #ResearchHumor #RealignChakras #HackerNews #ngated -
🚀 "The #AI gods have spoken 🧙♂️: #Benchmarks are no longer impressive, and your deep learning model is officially doomed to #mediocrity. 🎓 This 10,000-author paper hilariously confirms that we've hit 'Peak Benchmark'—time to move on and realign your chakras!" 🧘♀️
https://arxiv.org/abs/2602.16763 #Peak #DeepLearning #ResearchHumor #RealignChakras #HackerNews #ngated -
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
https://arxiv.org/abs/2602.16763
Comments: https://news.ycombinator.com/item?id=49170915
#HackerNews #AI #Benchmarks #BenchmarkSaturation #Research #MachineLearning #AITrends
-
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
https://arxiv.org/abs/2602.16763
Comments: https://news.ycombinator.com/item?id=49170915
#HackerNews #AI #Benchmarks #BenchmarkSaturation #Research #MachineLearning #AITrends
-
📊 Nova 2.0 Lite (medium) — the actual numbers
GPQA: 76.8%
MMLU-Pro: 81.3%
Humanity's Last Exam: 8.6%
Long Context Reasoning: 58.3%⚡ 217.9 tokens/sec
💰 22.4 intelligence points per dollarMeasured independently, not self-reported →https://olud.ai/leaderboard.html
-
📊 Nova 2.0 Lite (medium) — the actual numbers
GPQA: 76.8%
MMLU-Pro: 81.3%
Humanity's Last Exam: 8.6%
Long Context Reasoning: 58.3%⚡ 217.9 tokens/sec
💰 22.4 intelligence points per dollarMeasured independently, not self-reported →https://olud.ai/leaderboard.html
-
DeepSeek's retrained V4-Flash model scored 50 on Artificial Analysis's Intelligence Index—a 10-point gain without architectural changes—while costing roughly 60% less per task than GPT-5.6 Luna. The improvement came through post-training refinement alone. https://www.implicator.ai/deepseeks-retrained-v4-flash-scores-50-on-independent-intelligence-index/ #AI #LLMs #Benchmarks
-
DeepSeek's retrained V4-Flash model scored 50 on Artificial Analysis's Intelligence Index—a 10-point gain without architectural changes—while costing roughly 60% less per task than GPT-5.6 Luna. The improvement came through post-training refinement alone. https://www.implicator.ai/deepseeks-retrained-v4-flash-scores-50-on-independent-intelligence-index/ #AI #LLMs #Benchmarks
-
📊 Qwen3 Max — the actual numbers
GPQA: 76.4%
MMLU-Pro: 84.1%
Humanity's Last Exam: 11.1%
Long Context Reasoning: 46.7%💰 10 intelligence points per dollar
Measured independently, not self-reported →https://olud.ai/leaderboard.html
-
📊 Qwen3 Max — the actual numbers
GPQA: 76.4%
MMLU-Pro: 84.1%
Humanity's Last Exam: 11.1%
Long Context Reasoning: 46.7%💰 10 intelligence points per dollar
Measured independently, not self-reported →https://olud.ai/leaderboard.html