home.social

#benchmarks — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #benchmarks, aggregated by home.social.

fetched live
  1. RT @ChrisGPT: DeepSeek V4 Pro 0813 ist heute ebenfalls veröffentlicht! DeepSeek meldet Sprünge von 72,1 % auf 87,9 % im Terminal-Bench 2.1, von 52,7 % auf 83,3 % im CyberGym und von 12,8 % auf 62,7 % im DeepSWE!! Es kostet 0,435 $ pro Million Eingabe-Token und 0,87 $ pro Million Ausgabe-Token bei einem Kontextfenster von 1 Million Token. Ich mag es nicht, dass jedes Unternehmen jedes Mal sehr unterschiedliche Benchmarks meldet. Also schätze ich, wir werden uns auf den Index verlassen müssen.

    mehr auf Arint.info

    #AI #Benchmarks #DeepSeek #MachineLearning #TechNews #arint_info

    https://x.com/ChrisGPT/status/2087572834650407024#m

  2. RT @ChrisGPT: DeepSeek V4 Pro 0813 ist heute ebenfalls veröffentlicht! DeepSeek meldet Sprünge von 72,1 % auf 87,9 % im Terminal-Bench 2.1, von 52,7 % auf 83,3 % im CyberGym und von 12,8 % auf 62,7 % im DeepSWE!! Es kostet 0,435 $ pro Million Eingabe-Token und 0,87 $ pro Million Ausgabe-Token bei einem Kontextfenster von 1 Million Token. Ich mag es nicht, dass jedes Unternehmen jedes Mal sehr unterschiedliche Benchmarks meldet. Also schätze ich, wir werden uns auf den Index verlassen müssen.

    mehr auf Arint.info

    #AI #Benchmarks #DeepSeek #MachineLearning #TechNews #arint_info

    https://x.com/ChrisGPT/status/2087572834650407024#m

  3. Mi:dm K 2.5 Pro hits 80.9% on MMLU-Pro but only 10% on long-context reasoning — independently measured, not self-reported.

    olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  4. Mi:dm K 2.5 Pro hits 80.9% on MMLU-Pro but only 10% on long-context reasoning — independently measured, not self-reported.

    olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  5. A CursorBench software-engineering test found Anthropic's cheaper Opus 5 approached Fable 5's performance at roughly half the cost per task. Buyers may be choosing based on price-to-capability rather than raw model rank—a dynamic that could reshape enterprise AI spending. implicator.ai/anthropic-ipo-fa #AI #adoption #benchmarks

  6. A CursorBench software-engineering test found Anthropic's cheaper Opus 5 approached Fable 5's performance at roughly half the cost per task. Buyers may be choosing based on price-to-capability rather than raw model rank—a dynamic that could reshape enterprise AI spending. implicator.ai/anthropic-ipo-fa #AI #adoption #benchmarks

  7. Kimi K2.7 Code hits 89.6% on GPQA and 40.5 tokens/sec, but the real standout is 25.1 intelligence points per dollar — that’s how the efficiency math flips.

    olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  8. Kimi K2.7 Code hits 89.6% on GPQA and 40.5 tokens/sec, but the real standout is 25.1 intelligence points per dollar — that’s how the efficiency math flips.

    olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  9. 📊 GLM-4.6 (Reasoning) — the actual numbers

    GPQA: 78%
    MMLU-Pro: 82.9%
    Humanity's Last Exam: 14.5%
    Long Context Reasoning: 55.3%

    💰 30.4 intelligence points per dollar

    Measured independently, not self-reported →olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  10. 📊 GLM-4.6 (Reasoning) — the actual numbers

    GPQA: 78%
    MMLU-Pro: 82.9%
    Humanity's Last Exam: 14.5%
    Long Context Reasoning: 55.3%

    💰 30.4 intelligence points per dollar

    Measured independently, not self-reported →olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  11. 📊 Granite 4.0 1B — the actual numbers

    GPQA: 28.1%
    MMLU-Pro: 32.5%
    Humanity's Last Exam: 4.8%
    Long Context Reasoning: 6%

    Measured independently, not self-reported →olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  12. 📊 Granite 4.0 1B — the actual numbers

    GPQA: 28.1%
    MMLU-Pro: 32.5%
    Humanity's Last Exam: 4.8%
    Long Context Reasoning: 6%

    Measured independently, not self-reported →olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  13. 📊 Kimi K2 0905 — the actual numbers

    GPQA: 76.7%
    MMLU-Pro: 81.9%
    Humanity's Last Exam: 6.4%
    Long Context Reasoning: 53.7%

    💰 22.3 intelligence points per dollar

    Measured independently, not self-reported →olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  14. 📊 Kimi K2 0905 — the actual numbers

    GPQA: 76.7%
    MMLU-Pro: 81.9%
    Humanity's Last Exam: 6.4%
    Long Context Reasoning: 53.7%

    💰 22.3 intelligence points per dollar

    Measured independently, not self-reported →olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  15. 🎩🤡 Oh, look! Another programmer pretending to "innovate" by reinventing the wheel... in Rust! Because nothing screams "cutting-edge" like rehashing #Scala concepts with a side of benchmark obsession. 🛠️🔨
    lordgoati.us/blog/tail-call/ #programming #innovation #reinventingthewheel #Rust #benchmarks #HackerNews #ngated

  16. 🎩🤡 Oh, look! Another programmer pretending to "innovate" by reinventing the wheel... in Rust! Because nothing screams "cutting-edge" like rehashing #Scala concepts with a side of benchmark obsession. 🛠️🔨
    lordgoati.us/blog/tail-call/ #programming #innovation #reinventingthewheel #Rust #benchmarks #HackerNews #ngated

  17. 📊 Exaone 4.0 1.2B (Non-reasoning) — the actual numbers

    GPQA: 42.4%
    MMLU-Pro: 50%
    Humanity's Last Exam: 5.7%
    Long Context Reasoning: 0%

    Measured independently, not self-reported →olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  18. 📊 Exaone 4.0 1.2B (Non-reasoning) — the actual numbers

    GPQA: 42.4%
    MMLU-Pro: 50%
    Humanity's Last Exam: 5.7%
    Long Context Reasoning: 0%

    Measured independently, not self-reported →olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  19. Qwen3 VL 4B (Reasoning) hits 70% on MMLU-Pro but only 4.6% on Humanity's Last Exam—a clear gap between strong general knowledge and truly hard reasoning.

    olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  20. Qwen3 VL 4B (Reasoning) hits 70% on MMLU-Pro but only 4.6% on Humanity's Last Exam—a clear gap between strong general knowledge and truly hard reasoning.

    olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  21. 📊 Qwen3 235B A22B (Reasoning) — the actual numbers

    GPQA: 70%
    MMLU-Pro: 82.8%
    Humanity's Last Exam: 11%
    Long Context Reasoning: 0%

    💰 5.1 intelligence points per dollar

    Measured independently, not self-reported →olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  22. 📊 Qwen3 235B A22B (Reasoning) — the actual numbers

    GPQA: 70%
    MMLU-Pro: 82.8%
    Humanity's Last Exam: 11%
    Long Context Reasoning: 0%

    💰 5.1 intelligence points per dollar

    Measured independently, not self-reported →olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  23. Sonnet 5 posted 63.2% on SWE-bench Pro. OpenAI's Terra claims 84.3% on Terminal-Bench, vendor-stated, preview-only, unverified. Say that part out loud before repeating it. A benchmark you can't confirm isn't evidence. It's a press release with decimals.

    #AI #AIAgents #SoftwareEngineering #Benchmarks #marketing

  24. Sonnet 5 posted 63.2% on SWE-bench Pro. OpenAI's Terra claims 84.3% on Terminal-Bench, vendor-stated, preview-only, unverified. Say that part out loud before repeating it. A benchmark you can't confirm isn't evidence. It's a press release with decimals.

    #AI #AIAgents #SoftwareEngineering #Benchmarks #marketing

  25. Nova Premier hits 67.6 tokens/sec but only 4.7% on Humanity's Last Exam — speed doesn't equal reasoning. See how it stacks against others in our live bench.

    olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  26. Nova Premier hits 67.6 tokens/sec but only 4.7% on Humanity's Last Exam — speed doesn't equal reasoning. See how it stacks against others in our live bench.

    olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  27. 📊 DeepSeek-V2.5 (Dec '24) scores 76.3% on MATH-500. That’s independently measured, not self-reported—so it’s a real benchmark, not a press release. See how it stacks up against the claims:

    olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  28. 📊 DeepSeek-V2.5 (Dec '24) scores 76.3% on MATH-500. That’s independently measured, not self-reported—so it’s a real benchmark, not a press release. See how it stacks up against the claims:

    olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  29. Alibaba released benchmark scores for Qwen3.8-Max two weeks after claiming superiority without data. The 2.4T-parameter model scores come from internal testing only. Independent verification remains pending. What to watch: third-party results and the unstated license for weights arriving next week. implicator.ai/alibaba-publishe #AI #LLMs #benchmarks

  30. Alibaba released benchmark scores for Qwen3.8-Max two weeks after claiming superiority without data. The 2.4T-parameter model scores come from internal testing only. Independent verification remains pending. What to watch: third-party results and the unstated license for weights arriving next week. implicator.ai/alibaba-publishe #AI #LLMs #benchmarks

  31. 🚀 "The #AI gods have spoken 🧙‍♂️: #Benchmarks are no longer impressive, and your deep learning model is officially doomed to #mediocrity. 🎓 This 10,000-author paper hilariously confirms that we've hit 'Peak Benchmark'—time to move on and realign your chakras!" 🧘‍♀️
    arxiv.org/abs/2602.16763 #Peak #DeepLearning #ResearchHumor #RealignChakras #HackerNews #ngated

  32. 🚀 "The #AI gods have spoken 🧙‍♂️: #Benchmarks are no longer impressive, and your deep learning model is officially doomed to #mediocrity. 🎓 This 10,000-author paper hilariously confirms that we've hit 'Peak Benchmark'—time to move on and realign your chakras!" 🧘‍♀️
    arxiv.org/abs/2602.16763 #Peak #DeepLearning #ResearchHumor #RealignChakras #HackerNews #ngated

  33. 📊 Nova 2.0 Lite (medium) — the actual numbers

    GPQA: 76.8%
    MMLU-Pro: 81.3%
    Humanity's Last Exam: 8.6%
    Long Context Reasoning: 58.3%

    ⚡ 217.9 tokens/sec
    💰 22.4 intelligence points per dollar

    Measured independently, not self-reported →olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  34. 📊 Nova 2.0 Lite (medium) — the actual numbers

    GPQA: 76.8%
    MMLU-Pro: 81.3%
    Humanity's Last Exam: 8.6%
    Long Context Reasoning: 58.3%

    ⚡ 217.9 tokens/sec
    💰 22.4 intelligence points per dollar

    Measured independently, not self-reported →olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  35. DeepSeek's retrained V4-Flash model scored 50 on Artificial Analysis's Intelligence Index—a 10-point gain without architectural changes—while costing roughly 60% less per task than GPT-5.6 Luna. The improvement came through post-training refinement alone. implicator.ai/deepseeks-retrai #AI #LLMs #Benchmarks

  36. DeepSeek's retrained V4-Flash model scored 50 on Artificial Analysis's Intelligence Index—a 10-point gain without architectural changes—while costing roughly 60% less per task than GPT-5.6 Luna. The improvement came through post-training refinement alone. implicator.ai/deepseeks-retrai #AI #LLMs #Benchmarks

  37. 📊 Qwen3 Max — the actual numbers

    GPQA: 76.4%
    MMLU-Pro: 84.1%
    Humanity's Last Exam: 11.1%
    Long Context Reasoning: 46.7%

    💰 10 intelligence points per dollar

    Measured independently, not self-reported →olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI

  38. 📊 Qwen3 Max — the actual numbers

    GPQA: 76.4%
    MMLU-Pro: 84.1%
    Humanity's Last Exam: 11.1%
    Long Context Reasoning: 46.7%

    💰 10 intelligence points per dollar

    Measured independently, not self-reported →olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI