home.social

#benchmarks — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #benchmarks, aggregated by home.social.

  1. 📊 DeepSeek V3.2 Exp (Non-reasoning) — the actual numbers

    GPQA: 73.8%
    MMLU-Pro: 83.6%
    Humanity's Last Exam: 9%
    Long Context Reasoning: 44.7%

    💰 44.1 intelligence points per dollar

    Measured independently, not self-reported →olud.ai/leaderboard.html

    #LLM #Benchmarks #OpenSource #AI