#gsm8k — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #gsm8k, aggregated by home.social.
-
🧪 LLM Benchmark Showdown: 5 lokale Ollama-Modelle im Vergleich
Getestet auf derselben Hardware (#gmktecevo2 #AMDRyzenAIMaxPlus395 #strixhalo):
• #GSM8K (100 Samples) — Math
• #BFCL (100/Kategorie) — Function Calling
• #MBPP+ (50) — Python Coding
• #HumanEval+ (20) — Python Coding📊 Ergebnisse (Accuracy / Output TK/s / VRAM):
**qwen3.8:27b**
GSM8K 82% | BFCL 91.5% | MBPP+ 100% | HE+ 100%
⚡ 25.5 TK/s | 💾 18 GB VRAM**qwen3.6:27b**
GSM8K 83% | BFCL 93% | MBPP+ 98% | HE+ 75%
⚡ 12.7 TK/s | 💾 33 GB VRAM**qwen3.6:35b**
GSM8K 84% | BFCL 90% | MBPP+ 98% | HE+ 55%
⚡ 61.8 TK/s | 💾 27 GB VRAM**ornith-1.5:35b**
GSM8K 75% | BFCL 92.5% | MBPP+ 78% | HE+ 0%
⚡ 63.6 TK/s | 💾 26 GB VRAM**nemotron-3.5-lightning:30b**
GSM8K 59% | BFCL 74% | MBPP+ 94% | HE+ 0%
⚡ 91.9 TK/s | 💾 26 GB VRAM🏆 Fazit:
qwen3.8:27b ist der klare Sieger — als einziges Modell 100% bei beiden Coding-Benchmarks, bei GSM8K/BFCL gleichauf mit den anderen Qwen-Modellen. Bei 25.5 TK/s und nur 18 GB VRAM das beste Qualität/Speed/Effizienz-Verhältnis.
qwen3.6:27b ist qualitativ nah dran (BFCL sogar 93%), aber mit 12.7 TK/s unerträglich langsam und frisst 33 GB VRAM — fast 2× so viel wie qwen3.8 bei halber Speed.
qwen3.6:35b ist mit 61.8 TK/s 2.4× schneller als qwen3.8, aber HE+ nur 55% (vs 100%). Trading Code-Qualität für Speed.
ornith-1.5:35b und nemotron-3.5-lightning:30b fallen bei Coding komplett durch (HE+ 0%), sind aber die schnellsten Modelle im Feld (64 / 92 TK/s).
💡 TK/s = generierte Tokens/Sekunde (Warm-Run, ollama --verbose).
💾 VRAM = GPU-Speicher bei max context (262K bzw. 1M bei nemotron). -
RT @HowToAI_: Es stellte sich heraus, dass KI-Modelle keine Mathematik beherrschen – nicht einmal solche auf Grundschulniveau, wie sie ein 10-Jähriger lösen würde.
mehr auf Arint.info
#Forschung #GSM8K #KI #KünstlicheIntelligenz #MaschinellesLernen #TechNews #arint_info
-
Large language models (LLMs) have stormed onto the scene, dazzling us with their linguistic prowess and seeming intelligence. From crafting creative text formats to tackling complex coding challenges, they've left many wondering: are these machines truly thinking? The spotlight, in particular, has fallen on their mathematical reasoning abilities, with many claiming these models are on par with human problem-solvers. But a new study throws some serious shade on these claims, suggesting LLMs might be more about sophisticated mimicry than genuine understanding.
The Illusion of Mathematical Mastery
A popular benchmark for gauging the mathematical chops of LLMs is the GSM8K dataset. This collection of grade-school math problems has seen LLMs acing the test with impressive scores, fuelling the narrative of their growing mathematical intelligence. However, researchers are now questioning the validity of these results, arguing they offer a superficial view of LLMs' true capabilities. The study's authors introduce GSM-Symbolic, a souped-up benchmark crafted from symbolic templates. This framework allows for the generation of diverse variations of the same problem, providing a more nuanced and comprehensive evaluation. And what did they find? The performance of LLMs is anything but consistent. Across various model architectures, accuracy fluctuates wildly when faced with different instantiations of the same problem, even when only the numerical values are tweaked. This inconsistency is particularly alarming considering that genuine mathematical reasoning should be impervious to such superficial changes. A human student wouldn't suddenly forget how to solve a problem just because the numbers involved are different. This suggests that LLMs are not engaging in true logical deduction but rather relying on a form of probabilistic pattern matching.Fragile Foundations: The Sensitivity of LLMs
Further investigation into the fragility of LLM reasoning revealed a critical weakness: sensitivity to changes in the problem's presentation. While models showed some resilience to variations in proper names, their performance took a nosedive when numerical values were altered. As the complexity ramped up, with additional clauses introduced, accuracy plummeted, and performance variability shot up. This trend, consistent across various LLMs, reinforces the notion that their reasoning is highly dependent on the specific problem format they've encountered during training.The "No-Op" Test: Exposing the Limits of Understanding
To truly put LLMs' mathematical comprehension to the test, researchers concocted a cunning challenge: GSM-NoOp. This dataset features problems peppered with seemingly relevant but ultimately inconsequential statements – think adding details about fruit size in a problem about counting total fruit. The results were startling. Across the board, LLMs tripped up, blindly incorporating these extraneous details into their calculations. This tendency to translate statements into operations without grasping their true significance highlights a fundamental flaw in their understanding of mathematical concepts. Even when provided with examples demonstrating the irrelevance of these "No-Op" statements, the models remained stubbornly fixated on incorporating them, revealing a deep-seated limitation in their reasoning processes. These findings cast serious doubt on the ability of current LLMs to perform genuine mathematical reasoning, suggesting they might be masters of imitation rather than true mathematical minds.The Quest for Genuine Reasoning
While LLMs have undoubtedly made remarkable strides, the study's findings urge a reassessment of their true capabilities. Their fragility, sensitivity to superficial changes, and inability to discern relevant information underscore the limitations of their current reasoning abilities. The quest for AI systems that can truly reason, going beyond mimicking patterns to achieve genuine problem-solving prowess, remains a formidable challenge. This pursuit demands new approaches to model development and a more critical evaluation of their performance. Only then can we move closer to creating AI that can truly comprehend and reason about the world around us.Unlock the Future of Business with AI
Dive into our immersive workshops and equip your team with the tools and knowledge to lead in the AI era.
Get in touch with ushttps://www.ikangai.com/unmasking-the-mathematical-minds-of-llms-are-they-really-reasoning/
-
Can LLMs truly reason? Are they just sophisticated pattern matchers? Quite a cool pre-print @ https://arxiv.org/abs/2410.05229
#LLM #Reasoning #Mathematics #AGI #GSM8k -
[Перевод] Самые популярные LLM бенчмарки
Зачем использовать бенчмарки для оценки LLM? Бенчмарки LLM помогают оценивать точность больших языковых моделей, обеспечивая стандартизированную процедуру измерения метрик выполнения различных задач. Бенчмарки содержат все структуры и данные , необходимые для оценки LLM, в том числе: «Эталонные» датасеты (релевантные задачи/вопросы/промты с ожидаемыми ответами) Способы передачи входных промтов в LLM Способы интерпретации/сбора ответов Вычисляемые метрики и оценки (а также способы их вычисления) Всё вместе это позволяет согласованным образом сравнивать точность разных моделей. Но какой же бенчмарк LLM стоит использовать? В основном это зависит от сценария использования, то есть от того, для чего вы намереваетесь применять LLM. Давайте разбираться!
-
🧮 MathDial is based on #GSM8k and annotated with ground-truth solutions, student guesses, and plenty of annotations from teachers about the student solution, confusion, quality of dialog, and many more. (3/🧵) #EMNLP2023