#aibenchmark — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #aibenchmark, aggregated by home.social.
-
Kimi K3: second only to Fable 5 on AA-Briefcase
https://artificialanalysis.ai/articles/kimi-k3-agentic-knowledge-benchmark
Comments: https://news.ycombinator.com/item?id=49001930
#HackerNews #KimiK3 #Fable5 #AABriefcase #AIbenchmark #technology
-
Tested Cogito V1 14B Qwen on my Linux server. 45 t/s, 9.7GB VRAM, and the same IDA self-awareness trick its 8B sibling pulled -- Run 2 deliberately stepped back to brute force because a beginner probably needed simpler first. Run 3 came back stronger with a nice candy analogy. That's DeepCogito's IDA training making a transformation of Qwen into something way better.
Read the full breakdown below.
-
Tested Cogito V1 8B on my Linux server. 83 t/s, 5.4GB VRAM, 131k context. The real story is where it deliberately wrote worse code because it decided a beginner needed simplicity over efficiency -- and admitted it! That's IDA self-reflection making a live call.
I guess a 5GB model with a conscience is worth more than a 70B model with none?Read the full breakdown below.
-
Windows 11 est le dernier des Windows
https://fed.brid.gy/r/https://korben.info/windows-11-performances-degradation-benchmark.html
-
AA-Omniscience: New AI Reliability Benchmark Reveals Top Models Are More Likely to Hallucinate
#AI #LLM #GenAI #AIBenchmark #Hallucination #AISafety #OpenAI #Anthropic #xAI #Grok #GPT51 #ClaudeAI
-
OpenAI launches GDPval to measure AI performance on real-world economic tasks
https://web.brid.gy/r/https://nerds.xyz/2025/09/openai-gdpval/
-
AI Job Takeover? Not Yet, Agents Disappoint with Low 25% Success Rate in Business Automation Study
#AIAgents #Automation #FutureOfWork #LLM #CarnegieMellon #TheAgentCompany #AIbenchmark #TechNews #AIethics #JobAutomation
-
Die Grenzen von KI austesten
Reuters & die New York Times berichten über einen neuen Test: Humanity's Last Exam. Mit 3.000 Fragen aus über 100 Themengebieten werden hier die Grenzen moderner KI-Systeme ausgetestet. Thorben Jansen vom IPN war an der Entwicklung beteiligt.
🔗 Mehr: https://lastexam.ai
New York Times: https://www.reuters.com/technology/artificial-intelligence/ai-experts-ready-humanitys-last-exam-stump-powerful-tech-2024-09-16/