home.social

#aireasoning — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #aireasoning, aggregated by home.social.

fetched live
  1. Terence Tao @tao :

    The math behind today’s LLMs is actually simple. Training and running them mostly uses linear algebra, matrix multiplication, and a bit of calculus, material an undergraduate can handle. We understand how to build and operate these models.

    The real mystery is why they work so well on some tasks and fail on others, and why we cannot predict that in advance. We lack good rules for forecasting performance across tasks, so progress is largely empirical.

    A key reason is the nature of real-world data. Pure noise is well understood, perfectly structured data is well understood, but natural text sits in between, partly structured and partly random. Mathematics for that middle regime is thin, similar to how physics struggles at meso-scales between atoms and continua.

    Because of this gap, we can describe the mechanisms but cannot yet explain capability jumps or give reliable task-level predictions. That mismatch, simple machinery versus hard-to-predict behavior, is the core puzzle.

    ----

    Video from Prof @Briankeating YT Channel.

    Full video link: youtu.be/ukpCHo5v-Gc

    #ArtificialIntelligence #AI #LargeLanguageModels #LLMs #MachineLearning #DeepLearning #GenerativeAI #AIResearch #Mathematics #LinearAlgebra #MatrixMultiplication #Calculus #NeuralNetworks #AIModels #ModelTraining #AIEngineering #ComputerScience #DataScience #NaturalLanguageProcessing #NLP #EmergentAbilities #AIReasoning #AIProgress #AIInnovation #TechResearch #FutureOfAI #MachineIntelligence #AIExplained #TuringTest

  2. Oh, the irony! 🤦‍♂️ A riveting deep dive into the enigma of AI reasoning gets blocked by a security service that can't even reason itself. 🤔💡 Why ponder the mysteries of machine intelligence when you can't even get past the gatekeeper? 🔒🥴
    cacm.acm.org/news/can-we-under #AIreasoning #irony #securitymachine #intelligencegatekeeper #HackerNews #ngated

  3. Open-Weight Models Edge Closer to Production Readiness

    Open-weight AI models like DeepSeek R1 can now do complex coding and reasoning. They can run on your own computer.

    #OpenWeightAI, #LLM, #AICoding, #AIReasoning, #LocalAI

    newsletter.tf/open-weight-ai-m

  4. Open-weight AI models are getting much better, reaching performance levels that allow them to be used for important tasks like coding and reasoning. This is a big step up from before.

    #OpenWeightAI, #LLM, #AICoding, #AIReasoning, #LocalAI
    newsletter.tf/open-weight-ai-m

  5. " #AIReasoning finally let's you see what the #AI really thinks."

    #LLMs don't *think*, they predict the next token.

    "Researchers have uncovered that the AI cheats when they turned on reasoning."

    Ever thought about reasoning also being text output just like non-reasoning, entirely controlled by the AI whose entire job it is to generate sycophantic text output? This output is always something made for human consumption, it is never, however, an *internal* state

    #LRM #RLM #noAI #AIHype

  6. [University of Pennsylvania] — 비용 효율적인 LLM 추론 검증을 위한 ‘Weak-Strong’ 연동 정책 및 SSV 알고리즘 분석 배경 및 개요 Continue reading on Medium »

    #ai-optimization #verificationpolicy #machine-learning #aireasoning #llm

    Origin | Interest | Match
  7. Google DeepMind just rolled out Gemini 3.1 Pro – an upgraded Gemini 3 “Deep Think” model built for heavy reasoning and complex tasks. It promises sharper chain‑of‑thought, better multi‑step problem solving, and tighter integration with generative AI pipelines. Curious how this could reshape ML workflows? Dive into the details. #Gemini3Pro #DeepThink #AIReasoning #GenerativeAI

    🔗 aidailypost.com/news/gemini-31

  8. Google's Gemini 3 Deep Think reached 84.6% on ARC-AGI-2, a reasoning benchmark designed to resist memorization. That beats GPT-5.2 (52.9%) and Claude (68.8%) by significant margins. The catch: $13.62 per task suggests these advances may remain research tools rather than production systems for now.

    #AIReasoning #Benchmarks #TestTimeCompute

    implicator.ai/google-gemini-3-

  9. New research shows that letting language models hold internal debates—checking each other’s claims and negotiating solutions—dramatically cuts errors on tough reasoning tasks. The multi‑agent approach boosts self‑consistency and semantic verification, pushing open‑source AI toward more reliable reasoning. Dive into the findings! #MultiAgentDebate #AIReasoning #SelfConsistency #SemanticVerification

    🔗 aidailypost.com/news/ai-models

  10. Trích xuất cấu trúc vượt trội so với ngữ cảnh đầy đủ (F1: 0.83 vs 0.58) trong tác vụ suy luận đa bước. Entity Cards (17.5% token) giúp mô hình suy luận tốt hơn do loại nhiễu, tập trung vào thực thể và quan hệ. Token compression (LLMLingua, QUITO) thất bại do phá vỡ cấu trúc ngữ nghĩa. Mô hình nhỏ (Qwen3-1.7B) có thể tạo Entity Cards với F1 0.60. Cần thử fine-tuning và kiểm tra trên RAG.
    #StructuredExtraction #EntityCards #AIReasoning #LLM #RAG #TríchXuấtCấuTrúc #SuyLuậnAI #MôHìnhNgônNgữ #RútGọ

  11. Thử nghiệm 23 mô hình ngôn ngữ lớn (LLM) với câu đố Nonogram (câu đố logic dạng lưới). Kết quả: hiệu suất giảm mạnh khi kích thước tăng; một số LLM viết code để giải vét cạn, số khác lập luận từng bước như con người. GPT-4.5 dẫn đầu. Tổng chi phí: ~250 USD, ~17M tokens. Dữ liệu & mã nguồn mở. Link: nonobench.com, GitHub: no-bench.

    #LLM #Nonogram #LogicPuzzle #AI #Reasoning #MôHìnhNgônNgữ #CâuĐốLogic #TríTuệNhânTạo #AIReasoning

    reddit.com/r/LocalLLaMA/commen

  12. 🚀 Polish geniuses have supposedly revolutionized AI reasoning, and yet their announcement reads like a cryptic radio station playlist. 🎧 Surely the world was waiting with bated breath for an algorithm to decode Chopin on frequency czstotliwoci! 🎶
    polskieradio.pl/395/7784/artyk #PolishAI #Revolution #AIReasoning #ChopinAlgorithm #TechNews #HackerNews #ngated

  13. Do transformer-based LLMs really show emergent understanding? Probably not! A higher-level look at model outputs vindicates the "glorified autocomplete" take. hackernoon.com/how-ai-reasonin #aireasoning

  14. Warped Semantic Manifolds: A New Path to Flawless AI Reasoning In the world of artificial intelligence, we’ve all heard the refrain: “It works, but we don’t know why.” Large language models...

    #noetic-geodesic #aireasoning #ai #machine-learning #semantic-manifolds

    Origin | Interest | Match
  15. #Meta has appointed #ShengjiaZhao as the #ChiefScientist of Meta #Superintelligence Labs (#MSL). Zhao, a former #OpenAI #researcher, will lead research efforts at MSL, focusing on #AIreasoning models. Alongside #AlexandrWang, the former CEO of #ScaleAI, Zhao will set the #researchagenda for MSL, aiming to compete with OpenAI and Google in the AI space. techcrunch.com/2025/07/25/meta #tech #media #news

  16. Techniques for monitoring the thoughts of AI reasoning models known as chains-of-thought or CoTs are now a thing to focus on.

    Researchers from OpenAI, Google DeepMind, Anthropic, and others indicate CoT monitoring may be a key method for understanding how AI reasoning models work and could be a core method to keep AI agents under control.

    COTs are an externalized process in which AI models work through problems, similar to how humans use a scratch pad to work.

    DL the research paper here: tomekkorbak.com/cot-monitorabi

    techcrunch.com/2025/07/15/rese #AI #AIReasoning #OpenAI #Google #DeepMind #Anthropic #COTs #Reasoning #AIModels #LLMs

  17. What does AI reasoning mean for global health?

    When epidemiologists investigate a disease outbreak, they do not just match symptoms to known pathogens. They work through complex chains of evidence, test hypotheses, reconsider assumptions when data does not fit, and sometimes completely change their approach based on new information. This deeply human process of systematic reasoning is what artificial intelligence systems are now learning to do.

    This capability represents a fundamental shift from AI that recognizes patterns to AI that can work through complex problems the way a skilled professional would. For those working in global health and education, understanding this transformation is essential.

    The difference between answering and reasoning

    To understand this revolution, consider how most AI works today versus how reasoning AI operates.

    Traditional AI excels at pattern recognition. Show it a chest X-ray, and it can identify pneumonia by matching patterns it learned from millions of examples. Ask it about disease symptoms, and it retrieves information from its training data. This is sophisticated, but it is fundamentally different from reasoning.

    Consider this scenario: An unusual cluster of respiratory illness appears in a rural community. The symptoms partially match several known diseases but perfectly match none. Environmental factors are unclear. Some patients respond to standard treatments. Others do not.

    A pattern-matching AI might list possible diseases based on symptom similarity. But a reasoning AI would approach it like an epidemiologist:

    • “Let me examine the symptom progression timeline.”
    • “The geographic clustering suggests environmental or infectious cause. Let me investigate both paths.”
    • “Wait, these treatment responses do not align with any single pathogen. Could this be co-infection?”
    • “I need to reconsider. What if the environmental factor is not the cause but is affecting treatment efficacy?”

    The AI actually works through the problem, forms hypotheses, recognizes when evidence contradicts its assumptions, and adjusts its approach accordingly.

    How reasoning AI thinks through problems

    Advanced AI systems now demonstrate visible thinking processes. When analyzing complex health data, they might:

    • “First, let me identify the key variables affecting disease transmission in this population.”
    • “I will start by calculating the basic reproduction number using standard methods.”
    • “These results seem inconsistent with the observed spread pattern. Let me check my assumptions.”
    • “I may have overlooked the role of asymptomatic carriers. Let me recalculate.”
    • “This aligns better with observations. Now I can project intervention outcomes.”

    This is not scripted behavior. The AI works through problems, recognizes errors, and corrects its approach—much like a researcher reviewing their analysis.

    Why reasoning requires massive computational power

    Reasoning AI systems require thousands of times more computational resources than traditional AI. Understanding why helps explain both their power and limitations.

    Think about the difference between recognizing a disease from symptoms versus investigating a novel outbreak. Recognition happens quickly: an experienced clinician identifies malaria almost instantly. But investigating an unusual disease cluster requires sustained analysis, exploring multiple hypotheses, checking each against evidence.

    The same applies to AI. Traditional pattern-matching AI makes a single pass through its neural network. But reasoning AI must:

    • Explore multiple hypotheses simultaneously;
    • Check each reasoning step for logical consistency;
    • Backtrack when evidence contradicts assumptions;
    • Verify conclusions against all available data; and
    • Consider alternative explanations.

    Each step requires intensive computation. The AI might explore hundreds of reasoning paths before reaching sound conclusions.

    Matching expert performance

    AI systems in mid-2025 perform at the level of graduate students in mathematics and other fields. For global health, this means AI that can:

    • Design epidemiological studies with appropriate controls;
    • Identify confounding variables in complex datasets;
    • Recognize when standard statistical methods do not apply; and
    • Develop novel approaches to emerging health challenges.

    This is not about calculating faster—computers have done that for decades. It is about understanding concepts, recognizing which analytical techniques to apply, and working through novel problems.

    Applications in global health

    Reasoning AI transforms multiple aspects of global health work:

    Outbreak investigation: AI that can integrate diverse data sources—clinical reports, environmental data, travel patterns, genetic sequences—to identify outbreak sources and transmission patterns.

    Treatment optimization: Systems that reason through drug interactions, comorbidities, and local factors to recommend personalized treatment protocols.

    Resource allocation: AI that understands trade-offs between prevention and treatment, immediate needs and long-term capacity building, to optimize limited resources.

    Research design: Systems that can identify weaknesses in study designs, suggest improvements, and recognize when findings may not generalize to other populations.

    Policy analysis: AI that reasons through complex interventions, anticipating unintended consequences and identifying implementation barriers.

    What makes AI reasoning different

    Five capabilities distinguish reasoning AI from pattern-matching systems:

    1. Working memory: Reasoning AI holds multiple pieces of information active while working through problems, like a human tracking several hypotheses simultaneously.
    2. Logical consistency: Each conclusion must follow logically from evidence and prior reasoning steps.
    3. Error recognition: When results do not make sense, the system recognizes the problem and adjusts its approach.
    4. Abstraction: The AI recognizes general principles and applies them to specific situations, not just memorizing solutions.
    5. Explanation: Reasoning AI can explain its logic, making its conclusions verifiable and trustworthy.

    The path forward

    The reasoning revolution does not replace human expertise but augments it in powerful ways. For global health professionals, this means:

    • AI partners that can work through complex epidemiological puzzles;
    • Systems that help design culturally appropriate interventions;
    • Tools that identify patterns humans might miss while respecting local knowledge.

    Understanding reasoning AI is no longer optional for those shaping global health. These systems are becoming intellectual partners capable of working through complex problems alongside human experts. The question is not whether to engage with this technology but how to use it effectively while maintaining human agency, judgment, and values in decisions that affect human lives.

    The ability to reason—to work systematically through complex problems—has always been central to advancing human health and knowledge. Now that machines are learning this capability, we must thoughtfully consider how to harness it for global benefit while ensuring human wisdom guides its application.

    #AIReasoning #ArtificialIntelligence #globalHealth #reasoning

  18. 🌟 95% accuracy gains? Discover how AI reasoning models are outperforming human experts and transforming decision-making across industries. Don’t get left behind in this quiet revolution! 🤖🔥
    #AIReasoning #MachineLearning #CognitiveAI
    👉
    medium.com/@rogt.x1997/95-accu

  19. The Ultimate LLM Prompting Showdown: Chain-of-Thought vs Self-Consistency vs Meta-Prompts State-of-the-art prompt engineering is not merely a niche skill for AI enthusiasts — it is a foundati...

    #machine-learning #prompt-engineering #llm #metaprompt #aireasoning

    Origin | Interest | Match
  20. 🤖 Think your AI assistant can really reason? Apple’s puzzle tests say otherwise.
    📉 See how “thinking” AIs collapse when logic gets real — and why we might be projecting intelligence where there is none.

    Hashtags:
    #AIReasoning #ChainOfThought #LLMFail #DeepTech

    URL:
    medium.com/@rogt.x1997/the-ill

  21. 🧠 What if AI pretends to think — but quits when things get real?

    Apple’s groundbreaking study shows models like Claude 3.7 hit 0% accuracy on complex tasks.
    Not because they’re slow. Because they give up.

    This piece explores the hidden failure mode of modern “thinking” AIs. You won’t see them the same way again.

    👇 Read and rethink the future:
    #AIReasoning #Claude3 #DeepSeek #AppleResearch
    medium.com/@rogt.x1997/the-ill

  22. French AI startup Mistral AI has introduced "Magistral," a new reasoning model designed to deliver logic-based answers across multiple languages. It offers responses up to 10 times faster than competitors and provides domain-specific expertise with high accuracy for solving complex problems.

    #MistralAI #Magistral #AIReasoning #MultilingualAI #AIInnovation #FutureOfAI #TechNews #ArtificialIntelligence #Greaternoida #students

  23. 🧠💡 Think your chatbot is reasoning like you?
    Think again. Just 1.5% of its neurons are faking intelligence brilliantly.
    LLMs don’t “think” — they pattern-match and guess smartly, until novelty breaks them.

    🔥 Read how modern AI mimics reasoning and why true AGI needs more than just training data:
    👉 medium.com/@rogt.x1997/the-1-5

    #ArtificialIntelligence #LLMs #AGI #AIReasoning #TechInsights #DeepLearning #NeurosymbolicAI
    medium.com/@rogt.x1997/the-1-5

Share on Mastodon

Enter the server where you have an account.