#ai-benchmarks — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #ai-benchmarks, aggregated by home.social.
-
Sakana AI Plans to Add Nvidia Nemotron Models to Fugu
#Nemotron #SakanaAI #NVIDIA #ArtificialIntelligenceAI #AIModels #LargeLanguageModelsLLMs #AIIntegration #AIAgents #OpenSourceAI #AgenticAI #AICoding #AIBenchmarks #AIInference #AIPartnerships #GenerativeAI
-
Sakana AI Plans to Add Nvidia Nemotron Models to Fugu
#Nemotron #SakanaAI #NVIDIA #ArtificialIntelligenceAI #AIModels #LargeLanguageModelsLLMs #AIIntegration #AIAgents #OpenSourceAI #AgenticAI #AICoding #AIBenchmarks #AIInference #AIPartnerships #GenerativeAI
-
Sakana AI Plans to Add Nvidia Nemotron Models to Fugu
#Nemotron #SakanaAI #NVIDIA #ArtificialIntelligenceAI #AIModels #LargeLanguageModelsLLMs #AIIntegration #AIAgents #OpenSourceAI #AgenticAI #AICoding #AIBenchmarks #AIInference #AIPartnerships #GenerativeAI
-
Sakana AI Plans to Add Nvidia Nemotron Models to Fugu
#Nemotron #SakanaAI #NVIDIA #ArtificialIntelligenceAI #AIModels #LargeLanguageModelsLLMs #AIIntegration #AIAgents #OpenSourceAI #AgenticAI #AICoding #AIBenchmarks #AIInference #AIPartnerships #GenerativeAI
-
Sakana AI Plans to Add Nvidia Nemotron Models to Fugu
#Nemotron #SakanaAI #NVIDIA #ArtificialIntelligenceAI #AIModels #LargeLanguageModelsLLMs #AIIntegration #AIAgents #OpenSourceAI #AgenticAI #AICoding #AIBenchmarks #AIInference #AIPartnerships #GenerativeAI
-
https://winbuzzer.com/2026/07/17/moonshot-ai-unveils-28t-parameter-kimi-k3-ai-model-xcxwbn/
Moonshot AI has launched its 2.8-trillion-parameter Kimi K3 model, but a high hallucination rate might tempers its frontier-model pitch.
#AI #MoonshotAI #KimiK3 #AIModels #MultimodalAI #MixtureOfExperts #AIReasoningModels #AIBenchmarks #ChinaAI
-
https://winbuzzer.com/2026/07/17/moonshot-ai-unveils-28t-parameter-kimi-k3-ai-model-xcxwbn/
Moonshot AI has launched its 2.8-trillion-parameter Kimi K3 model, but a high hallucination rate might tempers its frontier-model pitch.
#AI #MoonshotAI #KimiK3 #AIModels #MultimodalAI #MixtureOfExperts #AIReasoningModels #AIBenchmarks #ChinaAI
-
https://winbuzzer.com/2026/07/17/moonshot-ai-unveils-28t-parameter-kimi-k3-ai-model-xcxwbn/
Moonshot AI has launched its 2.8-trillion-parameter Kimi K3 model, but a high hallucination rate might tempers its frontier-model pitch.
#AI #MoonshotAI #KimiK3 #AIModels #MultimodalAI #MixtureOfExperts #AIReasoningModels #AIBenchmarks #ChinaAI
-
https://winbuzzer.com/2026/07/17/moonshot-ai-unveils-28t-parameter-kimi-k3-ai-model-xcxwbn/
Moonshot AI has launched its 2.8-trillion-parameter Kimi K3 model, but a high hallucination rate might tempers its frontier-model pitch.
#AI #MoonshotAI #KimiK3 #AIModels #MultimodalAI #MixtureOfExperts #AIReasoningModels #AIBenchmarks #ChinaAI
-
https://winbuzzer.com/2026/07/17/moonshot-ai-unveils-28t-parameter-kimi-k3-ai-model-xcxwbn/
Moonshot AI has launched its 2.8-trillion-parameter Kimi K3 model, but a high hallucination rate might tempers its frontier-model pitch.
#AI #MoonshotAI #KimiK3 #AIModels #MultimodalAI #MixtureOfExperts #AIReasoningModels #AIBenchmarks #ChinaAI
-
https://winbuzzer.com/2026/07/14/german-consortium-launches-soofi-s-for-sparse-industrial-ai-xcxwbn/
Germany’s Soofi S AI model pairs sparse architecture with strong project-run benchmarks, but licensing gaps and long-context limits temper its promise.
#AI #SoofiS30BA3B #AIModels #MixtureOfExperts #AITraining #AICompute #AIBenchmarks #AIInference #AIModelDevelopment
-
https://winbuzzer.com/2026/07/14/german-consortium-launches-soofi-s-for-sparse-industrial-ai-xcxwbn/
Germany’s Soofi S AI model pairs sparse architecture with strong project-run benchmarks, but licensing gaps and long-context limits temper its promise.
#AI #SoofiS30BA3B #AIModels #MixtureOfExperts #AITraining #AICompute #AIBenchmarks #AIInference #AIModelDevelopment
-
https://winbuzzer.com/2026/07/14/german-consortium-launches-soofi-s-for-sparse-industrial-ai-xcxwbn/
Germany’s Soofi S AI model pairs sparse architecture with strong project-run benchmarks, but licensing gaps and long-context limits temper its promise.
#AI #SoofiS30BA3B #AIModels #MixtureOfExperts #AITraining #AICompute #AIBenchmarks #AIInference #AIModelDevelopment
-
https://winbuzzer.com/2026/07/14/german-consortium-launches-soofi-s-for-sparse-industrial-ai-xcxwbn/
Germany’s Soofi S AI model pairs sparse architecture with strong project-run benchmarks, but licensing gaps and long-context limits temper its promise.
#AI #SoofiS30BA3B #AIModels #MixtureOfExperts #AITraining #AICompute #AIBenchmarks #AIInference #AIModelDevelopment
-
https://winbuzzer.com/2026/07/14/german-consortium-launches-soofi-s-for-sparse-industrial-ai-xcxwbn/
Germany’s Soofi S AI model pairs sparse architecture with strong project-run benchmarks, but licensing gaps and long-context limits temper its promise.
#AI #SoofiS30BA3B #AIModels #MixtureOfExperts #AITraining #AICompute #AIBenchmarks #AIInference #AIModelDevelopment
-
Current AI has launched Gap Map v0.1 to catalog open source AI projects, detailing 24,626 projects and 421 products.
#AI #OpenSourceAiGapMap #CurrentAi #OpenSourceAI #AIBenchmarks #AIModels #AIResearch
-
Current AI has launched Gap Map v0.1 to catalog open source AI projects, detailing 24,626 projects and 421 products.
#AI #OpenSourceAiGapMap #CurrentAi #OpenSourceAI #AIBenchmarks #AIModels #AIResearch
-
Current AI has launched Gap Map v0.1 to catalog open source AI projects, detailing 24,626 projects and 421 products.
#AI #OpenSourceAiGapMap #CurrentAi #OpenSourceAI #AIBenchmarks #AIModels #AIResearch
-
Current AI has launched Gap Map v0.1 to catalog open source AI projects, detailing 24,626 projects and 421 products.
#AI #OpenSourceAiGapMap #CurrentAi #OpenSourceAI #AIBenchmarks #AIModels #AIResearch
-
Current AI has launched Gap Map v0.1 to catalog open source AI projects, detailing 24,626 projects and 421 products.
#AI #OpenSourceAiGapMap #CurrentAi #OpenSourceAI #AIBenchmarks #AIModels #AIResearch
-
https://winbuzzer.com/2026/07/04/alibaba-skillweaver-claims-99-agent-token-cut-in-benchmark-xcxwbn/
Alibaba Cloud's SkillWeaver framewordk routes AI-agent tasks to relevant tools and claims 99% lower benchmark token use, but code and production proof remain open.
#AI #SkillWeaver #AlibabaCloud #Alibaba #AIAgents #AgenticAI #AIBenchmarks #Qwen #AIResearch
-
https://winbuzzer.com/2026/07/04/alibaba-skillweaver-claims-99-agent-token-cut-in-benchmark-xcxwbn/
Alibaba Cloud's SkillWeaver framewordk routes AI-agent tasks to relevant tools and claims 99% lower benchmark token use, but code and production proof remain open.
#AI #SkillWeaver #AlibabaCloud #Alibaba #AIAgents #AgenticAI #AIBenchmarks #Qwen #AIResearch
-
https://winbuzzer.com/2026/07/04/alibaba-skillweaver-claims-99-agent-token-cut-in-benchmark-xcxwbn/
Alibaba Cloud's SkillWeaver framewordk routes AI-agent tasks to relevant tools and claims 99% lower benchmark token use, but code and production proof remain open.
#AI #SkillWeaver #AlibabaCloud #Alibaba #AIAgents #AgenticAI #AIBenchmarks #Qwen #AIResearch
-
https://winbuzzer.com/2026/07/04/alibaba-skillweaver-claims-99-agent-token-cut-in-benchmark-xcxwbn/
Alibaba Cloud's SkillWeaver framewordk routes AI-agent tasks to relevant tools and claims 99% lower benchmark token use, but code and production proof remain open.
#AI #SkillWeaver #AlibabaCloud #Alibaba #AIAgents #AgenticAI #AIBenchmarks #Qwen #AIResearch
-
https://winbuzzer.com/2026/07/04/alibaba-skillweaver-claims-99-agent-token-cut-in-benchmark-xcxwbn/
Alibaba Cloud's SkillWeaver framewordk routes AI-agent tasks to relevant tools and claims 99% lower benchmark token use, but code and production proof remain open.
#AI #SkillWeaver #AlibabaCloud #Alibaba #AIAgents #AgenticAI #AIBenchmarks #Qwen #AIResearch
-
https://winbuzzer.com/2026/06/29/glm-52-tops-claude-code-in-semgrep-idor-benchmark-xcxwbn/
Z.ai's open-weight model GLM-5.2 has reached 39% F1 on Semgrep's IDOR benchmark, beating Anthropic's Claude Opus 4.8 in Claude Code
#AI #GLM52 #ClaudeCode #AIModels #AICoding #AIBenchmarks #OpenSourceAI #Anthropic #Claude
-
https://winbuzzer.com/2026/06/29/glm-52-tops-claude-code-in-semgrep-idor-benchmark-xcxwbn/
Z.ai's open-weight model GLM-5.2 has reached 39% F1 on Semgrep's IDOR benchmark, beating Anthropic's Claude Opus 4.8 in Claude Code
#AI #GLM52 #ClaudeCode #AIModels #AICoding #AIBenchmarks #OpenSourceAI #Anthropic #Claude
-
https://winbuzzer.com/2026/06/29/glm-52-tops-claude-code-in-semgrep-idor-benchmark-xcxwbn/
Z.ai's open-weight model GLM-5.2 has reached 39% F1 on Semgrep's IDOR benchmark, beating Anthropic's Claude Opus 4.8 in Claude Code
#AI #GLM52 #ClaudeCode #AIModels #AICoding #AIBenchmarks #OpenSourceAI #Anthropic #Claude
-
https://winbuzzer.com/2026/06/29/glm-52-tops-claude-code-in-semgrep-idor-benchmark-xcxwbn/
Z.ai's open-weight model GLM-5.2 has reached 39% F1 on Semgrep's IDOR benchmark, beating Anthropic's Claude Opus 4.8 in Claude Code
#AI #GLM52 #ClaudeCode #AIModels #AICoding #AIBenchmarks #OpenSourceAI #Anthropic #Claude
-
https://winbuzzer.com/2026/06/29/glm-52-tops-claude-code-in-semgrep-idor-benchmark-xcxwbn/
Z.ai's open-weight model GLM-5.2 has reached 39% F1 on Semgrep's IDOR benchmark, beating Anthropic's Claude Opus 4.8 in Claude Code
#AI #GLM52 #ClaudeCode #AIModels #AICoding #AIBenchmarks #OpenSourceAI #Anthropic #Claude
-
Patronus AI lands $50M to build ‘digital worlds’ that stress-test AI agents
AI agents are becoming more sophisticated. They are evolving from answering questions to autonomously executing multi-step complex tasks.…
#NewsBeep #News #US #USA #UnitedStates #UnitedStatesOfAmerica #Artificialintelligence #AI #aibenchmarks #ArtificialIntelligence #evaluation #GreenfieldPartners #lightspeed #NotableCapital #PatronusAI #Technology
https://www.newsbeep.com/us/726846/ -
Patronus AI lands $50M to build ‘digital worlds’ that stress-test AI agents
AI agents are becoming more sophisticated. They are evolving from answering questions to autonomously executing multi-step complex tasks.…
#NewsBeep #News #US #USA #UnitedStates #UnitedStatesOfAmerica #Artificialintelligence #AI #aibenchmarks #ArtificialIntelligence #evaluation #GreenfieldPartners #lightspeed #NotableCapital #PatronusAI #Technology
https://www.newsbeep.com/us/726846/ -
https://www.europesays.com/ie/554156/ Patronus AI lands $50M to build ‘digital worlds’ that stress-test AI agents #AI #AiBenchmarks #ArtificialIntelligence #ArtificialIntelligence #Éire #evaluation #GreenfieldPartners #IE #Ireland #lightspeed #NotableCapital #PatronusAI #Technology
-
https://winbuzzer.com/2026/06/18/glm-52-tops-open-weights-ai-ranking-as-coding-race-tightens-xcxwbn/
Z.ai's GLM-5.2 models takes the lead among open-weight models on Artificial Analysis' index, with public weights, a 1M-token window, and deployment caveats for coding teams.
#AI #GLM5 #AICoding #ChinaAI #AIModels #OpenSourceAI #AIBenchmarks
-
https://winbuzzer.com/2026/06/18/glm-52-tops-open-weights-ai-ranking-as-coding-race-tightens-xcxwbn/
Z.ai's GLM-5.2 models takes the lead among open-weight models on Artificial Analysis' index, with public weights, a 1M-token window, and deployment caveats for coding teams.
#AI #GLM5 #AICoding #ChinaAI #AIModels #OpenSourceAI #AIBenchmarks
-
https://winbuzzer.com/2026/06/18/glm-52-tops-open-weights-ai-ranking-as-coding-race-tightens-xcxwbn/
Z.ai's GLM-5.2 models takes the lead among open-weight models on Artificial Analysis' index, with public weights, a 1M-token window, and deployment caveats for coding teams.
#AI #GLM5 #AICoding #ChinaAI #AIModels #OpenSourceAI #AIBenchmarks
-
https://winbuzzer.com/2026/06/18/glm-52-tops-open-weights-ai-ranking-as-coding-race-tightens-xcxwbn/
Z.ai's GLM-5.2 models takes the lead among open-weight models on Artificial Analysis' index, with public weights, a 1M-token window, and deployment caveats for coding teams.
#AI #GLM5 #AICoding #ChinaAI #AIModels #OpenSourceAI #AIBenchmarks
-
https://winbuzzer.com/2026/06/18/glm-52-tops-open-weights-ai-ranking-as-coding-race-tightens-xcxwbn/
Z.ai's GLM-5.2 models takes the lead among open-weight models on Artificial Analysis' index, with public weights, a 1M-token window, and deployment caveats for coding teams.
#AI #GLM5 #AICoding #ChinaAI #AIModels #OpenSourceAI #AIBenchmarks
-
https://winbuzzer.com/2026/06/10/anthropic-opens-claude-fable-5-with-safety-routing-xcxwbn/
Anthropic has launched Claude Fable 5, bringing Mythos-class AI to regular Claude users with safety routing, a discounted June 22 access window, and usage-credit pricing.
#AI #ClaudeFable5 #Anthropic #Claude #ClaudeMythos #ProjectGlasswing #AIModels #AISafety #AISecurity #AIBenchmarks #EnterpriseAI #Cybersecurity
-
https://winbuzzer.com/2026/06/10/anthropic-opens-claude-fable-5-with-safety-routing-xcxwbn/
Anthropic has launched Claude Fable 5, bringing Mythos-class AI to regular Claude users with safety routing, a discounted June 22 access window, and usage-credit pricing.
#AI #ClaudeFable5 #Anthropic #Claude #ClaudeMythos #ProjectGlasswing #AIModels #AISafety #AISecurity #AIBenchmarks #EnterpriseAI #Cybersecurity
-
https://winbuzzer.com/2026/06/10/anthropic-opens-claude-fable-5-with-safety-routing-xcxwbn/
Anthropic has launched Claude Fable 5, bringing Mythos-class AI to regular Claude users with safety routing, a discounted June 22 access window, and usage-credit pricing.
#AI #ClaudeFable5 #Anthropic #Claude #ClaudeMythos #ProjectGlasswing #AIModels #AISafety #AISecurity #AIBenchmarks #EnterpriseAI #Cybersecurity
-
https://winbuzzer.com/2026/06/10/anthropic-opens-claude-fable-5-with-safety-routing-xcxwbn/
Anthropic has launched Claude Fable 5, bringing Mythos-class AI to regular Claude users with safety routing, a discounted June 22 access window, and usage-credit pricing.
#AI #ClaudeFable5 #Anthropic #Claude #ClaudeMythos #ProjectGlasswing #AIModels #AISafety #AISecurity #AIBenchmarks #EnterpriseAI #Cybersecurity
-
https://winbuzzer.com/2026/06/10/anthropic-opens-claude-fable-5-with-safety-routing-xcxwbn/
Anthropic has launched Claude Fable 5, bringing Mythos-class AI to regular Claude users with safety routing, a discounted June 22 access window, and usage-credit pricing.
#AI #ClaudeFable5 #Anthropic #Claude #ClaudeMythos #ProjectGlasswing #AIModels #AISafety #AISecurity #AIBenchmarks #EnterpriseAI #Cybersecurity
-
https://winbuzzer.com/2026/06/07/perplexity-lets-ai-agents-write-their-own-search-code-xcxwbn/
Perplexity's Search as Code lets AI agents generate Python search workflows, but claimed token savings and benchmark gains still need outside validation.
#AI #SearchAsCode #PerplexityAI #AISearch #AIAgents #AgenticAI #AIBenchmarks #SearchEngines
-
https://winbuzzer.com/2026/06/07/perplexity-lets-ai-agents-write-their-own-search-code-xcxwbn/
Perplexity's Search as Code lets AI agents generate Python search workflows, but claimed token savings and benchmark gains still need outside validation.
#AI #SearchAsCode #PerplexityAI #AISearch #AIAgents #AgenticAI #AIBenchmarks #SearchEngines
-
https://winbuzzer.com/2026/06/07/perplexity-lets-ai-agents-write-their-own-search-code-xcxwbn/
Perplexity's Search as Code lets AI agents generate Python search workflows, but claimed token savings and benchmark gains still need outside validation.
#AI #SearchAsCode #PerplexityAI #AISearch #AIAgents #AgenticAI #AIBenchmarks #SearchEngines
-
https://winbuzzer.com/2026/06/07/perplexity-lets-ai-agents-write-their-own-search-code-xcxwbn/
Perplexity's Search as Code lets AI agents generate Python search workflows, but claimed token savings and benchmark gains still need outside validation.
#AI #SearchAsCode #PerplexityAI #AISearch #AIAgents #AgenticAI #AIBenchmarks #SearchEngines
-
https://winbuzzer.com/2026/06/07/perplexity-lets-ai-agents-write-their-own-search-code-xcxwbn/
Perplexity's Search as Code lets AI agents generate Python search workflows, but claimed token savings and benchmark gains still need outside validation.
#AI #SearchAsCode #PerplexityAI #AISearch #AIAgents #AgenticAI #AIBenchmarks #SearchEngines
-
https://winbuzzer.com/2026/06/01/study-says-ai-search-agents-guess-before-they-search-xcxwbn/
Benchmark Reveals AI Search Agents Guess Before They Search
#AI #LiveBrowseComp #BrowseComp #AISearch #AIAgents #AgenticAI #AIBenchmarks #AIResearch
-
https://winbuzzer.com/2026/06/01/study-says-ai-search-agents-guess-before-they-search-xcxwbn/
Benchmark Reveals AI Search Agents Guess Before They Search
#AI #LiveBrowseComp #BrowseComp #AISearch #AIAgents #AgenticAI #AIBenchmarks #AIResearch
-
https://winbuzzer.com/2026/06/01/study-says-ai-search-agents-guess-before-they-search-xcxwbn/
Benchmark Reveals AI Search Agents Guess Before They Search
#AI #LiveBrowseComp #BrowseComp #AISearch #AIAgents #AgenticAI #AIBenchmarks #AIResearch
-
https://winbuzzer.com/2026/06/01/study-says-ai-search-agents-guess-before-they-search-xcxwbn/
Benchmark Reveals AI Search Agents Guess Before They Search
#AI #LiveBrowseComp #BrowseComp #AISearch #AIAgents #AgenticAI #AIBenchmarks #AIResearch
-
https://winbuzzer.com/2026/06/01/study-says-ai-search-agents-guess-before-they-search-xcxwbn/
Benchmark Reveals AI Search Agents Guess Before They Search
#AI #LiveBrowseComp #BrowseComp #AISearch #AIAgents #AgenticAI #AIBenchmarks #AIResearch
-
This article explores the instability metric, a benchmark designed to measure how consistently AI models reason through mathematical and logical problems. https://hackernoon.com/new-ai-benchmarks-are-testing-consistency-instead-of-memorization #aibenchmarks
-
This article explores the instability metric, a benchmark designed to measure how consistently AI models reason through mathematical and logical problems. https://hackernoon.com/new-ai-benchmarks-are-testing-consistency-instead-of-memorization #aibenchmarks
-
This article explores the instability metric, a benchmark designed to measure how consistently AI models reason through mathematical and logical problems. https://hackernoon.com/new-ai-benchmarks-are-testing-consistency-instead-of-memorization #aibenchmarks
-
This article explores the instability metric, a benchmark designed to measure how consistently AI models reason through mathematical and logical problems. https://hackernoon.com/new-ai-benchmarks-are-testing-consistency-instead-of-memorization #aibenchmarks
-
This article explores the instability metric, a benchmark designed to measure how consistently AI models reason through mathematical and logical problems. https://hackernoon.com/new-ai-benchmarks-are-testing-consistency-instead-of-memorization #aibenchmarks
-
https://winbuzzer.com/2026/05/28/deepswe-puts-gpt-55-ahead-in-ai-coding-tests-xcxwbn/
Datacurve's new DeepSWE benchmark puts GPT-5.5 ahead of Claude and challenges older AI coding rankings by arguing verifier design can distort results.
#AI #CodingBenchmarks #AIBenchmarks #AICoding #AIModels #OpenAI #Anthropic #GPT55 #ClaudeOpus47
-
https://winbuzzer.com/2026/05/28/deepswe-puts-gpt-55-ahead-in-ai-coding-tests-xcxwbn/
Datacurve's new DeepSWE benchmark puts GPT-5.5 ahead of Claude and challenges older AI coding rankings by arguing verifier design can distort results.
#AI #CodingBenchmarks #AIBenchmarks #AICoding #AIModels #OpenAI #Anthropic #GPT55 #ClaudeOpus47