home.social

#rlvr — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #rlvr, aggregated by home.social.

fetched live
  1. [Перевод] Как управлять reasoning effort в LLM

    Как современные LLM получают разные настройки reasoning effort: системные промпты, SFT, RLVR, штрафы за длину, reasoning budgets и on-policy-дистилляция. Разбираем DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5 и K3, GLM-5, Qwen3 и Inkling.

    habr.com/ru/articles/1071250/

    #LLM #reasoning_models #reasoning_effort #RLVR #reinforcement_learning #inference_scaling

  2. [Перевод] Как управлять reasoning effort в LLM

    Как современные LLM получают разные настройки reasoning effort: системные промпты, SFT, RLVR, штрафы за длину, reasoning budgets и on-policy-дистилляция. Разбираем DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5 и K3, GLM-5, Qwen3 и Inkling.

    habr.com/ru/articles/1071250/

    #LLM #reasoning_models #reasoning_effort #RLVR #reinforcement_learning #inference_scaling

  3. Данные, которые нельзя выдумать: как мы собираем трейсы для тренировки агентских навыков Cotype

    Привет, Хабр! На днях мы выкатили третье поколение языковых моделей Cotype — флагманскую Cotype Pro 3 и ее облегченную версию Cotype Light 3 (к слову, лайт мы выкатили чутка раньше, но это не имеет значения). Почитать о характеристиках моделей можно вот в этом посте. Тут же мы расскажем вам о нюансах обучения — точнее о том, как мы учим наши модели выполнять многошаговые сценарии без запинок, что совершенно необходимо, так как они выступают ключевым компонентом наших корпоративных ИИ-агентов и мультиагентных систем. Велком под кат

    habr.com/ru/companies/ru_mts/a

    #agents #llm #rlvr #grpo #harness #synthetic_data #искусственный_интеллект #агентский_режим #ииагенты #ииагенты_для_бизнеса

  4. Данные, которые нельзя выдумать: как мы собираем трейсы для тренировки агентских навыков Cotype

    Привет, Хабр! На днях мы выкатили третье поколение языковых моделей Cotype — флагманскую Cotype Pro 3 и ее облегченную версию Cotype Light 3 (к слову, лайт мы выкатили чутка раньше, но это не имеет значения). Почитать о характеристиках моделей можно вот в этом посте. Тут же мы расскажем вам о нюансах обучения — точнее о том, как мы учим наши модели выполнять многошаговые сценарии без запинок, что совершенно необходимо, так как они выступают ключевым компонентом наших корпоративных ИИ-агентов и мультиагентных систем. Велком под кат

    habr.com/ru/companies/ru_mts/a

    #agents #llm #rlvr #grpo #harness #synthetic_data #искусственный_интеллект #агентский_режим #ииагенты #ииагенты_для_бизнеса

  5. [Перевод] Внутри Claude нашли сознание у моделей. J-пространство в LLM

    Привет, Хабр. 6 июля 2026 года Anthropic опубликовала исследование Verbalizable Representations Form a Global Workspace in Language Models . Разбираю технические детали и практические следствия. Пока вы читаете это предложение, ваш мозг делает десятки вещей, которых вы не замечаете: держит осанку, регулирует дыхание, превращает линии на экране в слова. Но часть работы мозга вам всё-таки доступна - всплывший образ, обдуманное решение, куда направить внимание. Нейробиологи называют такую активность сознательно доступной и отличают её от всей остальной, идущей без нашего ведома. У неё особые свойства: её можно описать словами, ею можно управлять и рассуждать -> в отличие от автоматических процессов, которые текут сами собой. В новой работе Anthropic показывает, что похожее разделение появилось и в языковых моделях вроде Claude. У модели обнаружился небольшой набор внутренних нейронных паттернов, которые - на фоне всей остальной обработки играют особую роль.

    habr.com/ru/articles/1056320/

    #claude #anthropic #chatgpt #evals #archive #codex #claudecode #research #antigravity #rlvr

  6. [Перевод] Внутри Claude нашли сознание у моделей. J-пространство в LLM

    Привет, Хабр. 6 июля 2026 года Anthropic опубликовала исследование Verbalizable Representations Form a Global Workspace in Language Models . Разбираю технические детали и практические следствия. Пока вы читаете это предложение, ваш мозг делает десятки вещей, которых вы не замечаете: держит осанку, регулирует дыхание, превращает линии на экране в слова. Но часть работы мозга вам всё-таки доступна - всплывший образ, обдуманное решение, куда направить внимание. Нейробиологи называют такую активность сознательно доступной и отличают её от всей остальной, идущей без нашего ведома. У неё особые свойства: её можно описать словами, ею можно управлять и рассуждать -> в отличие от автоматических процессов, которые текут сами собой. В новой работе Anthropic показывает, что похожее разделение появилось и в языковых моделях вроде Claude. У модели обнаружился небольшой набор внутренних нейронных паттернов, которые - на фоне всей остальной обработки играют особую роль.

    habr.com/ru/articles/1056320/

    #claude #anthropic #chatgpt #evals #archive #codex #claudecode #research #antigravity #rlvr

  7. ICLR 2026 tổng hợp: Cộng đồng nghiên cứu tập trung vào GRPO (157 bài) thay vì DPO, ưu tiên RLVR (125 bài) thay vì RLHF, và 202 bài về Mamba/SSMs. Nait (tuning thông minh chỉ 10% dữ liệu) giúp tối ưu hiệu quả. 257 bài về tính toán lúc test, 123 bài về hallucination. Cảnh báo: mô hình tuân thủ tốt dễ bị tấn công injection. #AI #HọcMáy #ICLR2026 #NCKH #DeepLearning #Mamba #RLVR #GRPO #MạngNeural #BảoMậtAI #ViễnTưởngAI

    reddit.com/r/LocalLLaMA/commen

  8. 🧠 Mới! Notebook code RLVR kết hợp GRPO từ đầu, được chia sẻ trong dự án “Reasoning‑from‑Scratch”. Hữu ích cho những ai muốn khám phá mô hình RL và tối ưu hoá trong AI/ML. #AI #MachineLearning #RLVR #GRPO #LậpTrình #MãNguồn

    reddit.com/r/LocalLLaMA/commen

  9. RLVR promises faster sampling but leaves reasoning untouched—base LLMs still carry the heavy‑lifting of trajectories. The paper (NeurIPS 2025) shows that gains come from smarter teacher‑distillation and minor architectural tweaks, not a new reasoning engine. Curious how sampling efficiency separates from true understanding? Dive into the details. #RLVR #SamplingEfficiency #LLMReasoning #NeurIPS2025

    🔗 aidailypost.com/news/rlvr-lift

  10. 2025 saw significant advancements in #LLMs, with #ReinforcementLearning from #VerifiableRewards (#RLVR) emerging as a key stage in training, leading to improved #reasoning capabilities. The industry also began to understand the unique “jagged” intelligence of LLMs, excelling in specific domains but lacking generalisation. karpathy.bearblog.dev/year-in- #tech #media #news

  11. 2025 saw significant advancements in #LLMs, with #ReinforcementLearning from #VerifiableRewards (#RLVR) emerging as a key stage in training, leading to improved #reasoning capabilities. The industry also began to understand the unique “jagged” intelligence of LLMs, excelling in specific domains but lacking generalisation. karpathy.bearblog.dev/year-in- #tech #media #news

  12. 2025 saw significant advancements in #LLMs, with #ReinforcementLearning from #VerifiableRewards (#RLVR) emerging as a key stage in training, leading to improved #reasoning capabilities. The industry also began to understand the unique “jagged” intelligence of LLMs, excelling in specific domains but lacking generalisation. karpathy.bearblog.dev/year-in- #tech #media #news

  13. 2025 saw significant advancements in #LLMs, with #ReinforcementLearning from #VerifiableRewards (#RLVR) emerging as a key stage in training, leading to improved #reasoning capabilities. The industry also began to understand the unique “jagged” intelligence of LLMs, excelling in specific domains but lacking generalisation. karpathy.bearblog.dev/year-in- #tech #media #news

  14. New research from Tsinghua shows that reasoning‑augmented LLMs solve tasks with fewer calls but don’t surpass raw capability. The study compares chain‑of‑thought prompting, RL‑based RLVR, and pass@1 metrics, highlighting efficiency gains for open‑source models. Worth a read for anyone tracking LLM benchmarks. #LLM #ChainOfThought #RLVR #PassAt1

    🔗 aidailypost.com/news/study-fin

  15. New research from Tsinghua shows that reasoning‑augmented LLMs solve tasks with fewer calls but don’t surpass raw capability. The study compares chain‑of‑thought prompting, RL‑based RLVR, and pass@1 metrics, highlighting efficiency gains for open‑source models. Worth a read for anyone tracking LLM benchmarks. #LLM #ChainOfThought #RLVR #PassAt1

    🔗 aidailypost.com/news/study-fin

  16. New research from Tsinghua shows that reasoning‑augmented LLMs solve tasks with fewer calls but don’t surpass raw capability. The study compares chain‑of‑thought prompting, RL‑based RLVR, and pass@1 metrics, highlighting efficiency gains for open‑source models. Worth a read for anyone tracking LLM benchmarks. #LLM #ChainOfThought #RLVR #PassAt1

    🔗 aidailypost.com/news/study-fin

  17. → Les 4 étapes pour entrainer un LLM
    scienceetonnante.com/blog/2025

    « Voilà le principe de l'apprentissage par renforcement avec une récompense vérifiable [RLVR], qui permet de se passer d'humains qui doivent juger si la réponse est conforme ou pas. »

    #entrainer #LLM #apprentissage #RLVR #humains

  18. → Les 4 étapes pour entrainer un LLM
    scienceetonnante.com/blog/2025

    « Voilà le principe de l'apprentissage par renforcement avec une récompense vérifiable [RLVR], qui permet de se passer d'humains qui doivent juger si la réponse est conforme ou pas. »

    #entrainer #LLM #apprentissage #RLVR #humains

  19. → Les 4 étapes pour entrainer un LLM
    scienceetonnante.com/blog/2025

    « Voilà le principe de l'apprentissage par renforcement avec une récompense vérifiable [RLVR], qui permet de se passer d'humains qui doivent juger si la réponse est conforme ou pas. »

    #entrainer #LLM #apprentissage #RLVR #humains

  20. → Les 4 étapes pour entrainer un LLM
    scienceetonnante.com/blog/2025

    « Voilà le principe de l'apprentissage par renforcement avec une récompense vérifiable [RLVR], qui permet de se passer d'humains qui doivent juger si la réponse est conforme ou pas. »

    #entrainer #LLM #apprentissage #RLVR #humains

  21. New study challenges a key belief about Reinforcement Learning with Verifiable Rewards (RLVR) for #LLMs:
    #RLVR boosts efficiency but doesn't create new reasoning skills — #AI base models already had them!
    arxiv.org/abs/2504.13837

  22. New study challenges a key belief about Reinforcement Learning with Verifiable Rewards (RLVR) for #LLMs:
    #RLVR boosts efficiency but doesn't create new reasoning skills — #AI base models already had them!
    arxiv.org/abs/2504.13837

  23. New study challenges a key belief about Reinforcement Learning with Verifiable Rewards (RLVR) for #LLMs:
    #RLVR boosts efficiency but doesn't create new reasoning skills — #AI base models already had them!
    arxiv.org/abs/2504.13837

  24. New study challenges a key belief about Reinforcement Learning with Verifiable Rewards (RLVR) for #LLMs:
    #RLVR boosts efficiency but doesn't create new reasoning skills — #AI base models already had them!
    arxiv.org/abs/2504.13837

  25. Forschende der Tsinghua University und der Shanghai Jiao Tong University zeigen in einer Studie, dass #RLVR zwar die Wahrscheinlichkeit erhöht, beim ersten Versuch eine richtige Antwort zu generieren – das sogenannte pass@1 –, aber keine neuen Problemlösestrategien erschließt.
    the-decoder.de/forscher-zweife
    #KI

  26. Forschende der Tsinghua University und der Shanghai Jiao Tong University zeigen in einer Studie, dass #RLVR zwar die Wahrscheinlichkeit erhöht, beim ersten Versuch eine richtige Antwort zu generieren – das sogenannte pass@1 –, aber keine neuen Problemlösestrategien erschließt.
    the-decoder.de/forscher-zweife
    #KI

  27. Forschende der Tsinghua University und der Shanghai Jiao Tong University zeigen in einer Studie, dass #RLVR zwar die Wahrscheinlichkeit erhöht, beim ersten Versuch eine richtige Antwort zu generieren – das sogenannte pass@1 –, aber keine neuen Problemlösestrategien erschließt.
    the-decoder.de/forscher-zweife
    #KI

  28. Forschende der Tsinghua University und der Shanghai Jiao Tong University zeigen in einer Studie, dass #RLVR zwar die Wahrscheinlichkeit erhöht, beim ersten Versuch eine richtige Antwort zu generieren – das sogenannte pass@1 –, aber keine neuen Problemlösestrategien erschließt.
    the-decoder.de/forscher-zweife
    #KI

  29. • 🧠 Advanced post-training with reinforcement learning with verifiable rewards (#RLVR) using Group Relative Policy Optimization

    • 🔮 All models available in 7B, 13B, and 32B sizes, can be fine-tuned on a single H100 GPU

  30. • 🧠 Advanced post-training with reinforcement learning with verifiable rewards (#RLVR) using Group Relative Policy Optimization

    • 🔮 All models available in 7B, 13B, and 32B sizes, can be fine-tuned on a single H100 GPU

  31. • 🧠 Advanced post-training with reinforcement learning with verifiable rewards (#RLVR) using Group Relative Policy Optimization

    • 🔮 All models available in 7B, 13B, and 32B sizes, can be fine-tuned on a single H100 GPU

  32. • 🧠 Advanced post-training with reinforcement learning with verifiable rewards (#RLVR) using Group Relative Policy Optimization

    • 🔮 All models available in 7B, 13B, and 32B sizes, can be fine-tuned on a single H100 GPU