home.social

#rlvr — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #rlvr, aggregated by home.social.

fetched live
  1. [Перевод] Как управлять reasoning effort в LLM

    Как современные LLM получают разные настройки reasoning effort: системные промпты, SFT, RLVR, штрафы за длину, reasoning budgets и on-policy-дистилляция. Разбираем DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5 и K3, GLM-5, Qwen3 и Inkling.

    habr.com/ru/articles/1071250/

    #LLM #reasoning_models #reasoning_effort #RLVR #reinforcement_learning #inference_scaling

  2. [Перевод] Как управлять reasoning effort в LLM

    Как современные LLM получают разные настройки reasoning effort: системные промпты, SFT, RLVR, штрафы за длину, reasoning budgets и on-policy-дистилляция. Разбираем DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5 и K3, GLM-5, Qwen3 и Inkling.

    habr.com/ru/articles/1071250/

    #LLM #reasoning_models #reasoning_effort #RLVR #reinforcement_learning #inference_scaling

  3. [Перевод] Как управлять reasoning effort в LLM

    Как современные LLM получают разные настройки reasoning effort: системные промпты, SFT, RLVR, штрафы за длину, reasoning budgets и on-policy-дистилляция. Разбираем DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5 и K3, GLM-5, Qwen3 и Inkling.

    habr.com/ru/articles/1071250/

    #LLM #reasoning_models #reasoning_effort #RLVR #reinforcement_learning #inference_scaling

  4. Данные, которые нельзя выдумать: как мы собираем трейсы для тренировки агентских навыков Cotype

    Привет, Хабр! На днях мы выкатили третье поколение языковых моделей Cotype — флагманскую Cotype Pro 3 и ее облегченную версию Cotype Light 3 (к слову, лайт мы выкатили чутка раньше, но это не имеет значения). Почитать о характеристиках моделей можно вот в этом посте. Тут же мы расскажем вам о нюансах обучения — точнее о том, как мы учим наши модели выполнять многошаговые сценарии без запинок, что совершенно необходимо, так как они выступают ключевым компонентом наших корпоративных ИИ-агентов и мультиагентных систем. Велком под кат

    habr.com/ru/companies/ru_mts/a

    #agents #llm #rlvr #grpo #harness #synthetic_data #искусственный_интеллект #агентский_режим #ииагенты #ииагенты_для_бизнеса

  5. Данные, которые нельзя выдумать: как мы собираем трейсы для тренировки агентских навыков Cotype

    Привет, Хабр! На днях мы выкатили третье поколение языковых моделей Cotype — флагманскую Cotype Pro 3 и ее облегченную версию Cotype Light 3 (к слову, лайт мы выкатили чутка раньше, но это не имеет значения). Почитать о характеристиках моделей можно вот в этом посте. Тут же мы расскажем вам о нюансах обучения — точнее о том, как мы учим наши модели выполнять многошаговые сценарии без запинок, что совершенно необходимо, так как они выступают ключевым компонентом наших корпоративных ИИ-агентов и мультиагентных систем. Велком под кат

    habr.com/ru/companies/ru_mts/a

    #agents #llm #rlvr #grpo #harness #synthetic_data #искусственный_интеллект #агентский_режим #ииагенты #ииагенты_для_бизнеса

  6. Данные, которые нельзя выдумать: как мы собираем трейсы для тренировки агентских навыков Cotype

    Привет, Хабр! На днях мы выкатили третье поколение языковых моделей Cotype — флагманскую Cotype Pro 3 и ее облегченную версию Cotype Light 3 (к слову, лайт мы выкатили чутка раньше, но это не имеет значения). Почитать о характеристиках моделей можно вот в этом посте. Тут же мы расскажем вам о нюансах обучения — точнее о том, как мы учим наши модели выполнять многошаговые сценарии без запинок, что совершенно необходимо, так как они выступают ключевым компонентом наших корпоративных ИИ-агентов и мультиагентных систем. Велком под кат

    habr.com/ru/companies/ru_mts/a

    #agents #llm #rlvr #grpo #harness #synthetic_data #искусственный_интеллект #агентский_режим #ииагенты #ииагенты_для_бизнеса

  7. [Перевод] Внутри Claude нашли сознание у моделей. J-пространство в LLM

    Привет, Хабр. 6 июля 2026 года Anthropic опубликовала исследование Verbalizable Representations Form a Global Workspace in Language Models . Разбираю технические детали и практические следствия. Пока вы читаете это предложение, ваш мозг делает десятки вещей, которых вы не замечаете: держит осанку, регулирует дыхание, превращает линии на экране в слова. Но часть работы мозга вам всё-таки доступна - всплывший образ, обдуманное решение, куда направить внимание. Нейробиологи называют такую активность сознательно доступной и отличают её от всей остальной, идущей без нашего ведома. У неё особые свойства: её можно описать словами, ею можно управлять и рассуждать -> в отличие от автоматических процессов, которые текут сами собой. В новой работе Anthropic показывает, что похожее разделение появилось и в языковых моделях вроде Claude. У модели обнаружился небольшой набор внутренних нейронных паттернов, которые - на фоне всей остальной обработки играют особую роль.

    habr.com/ru/articles/1056320/

    #claude #anthropic #chatgpt #evals #archive #codex #claudecode #research #antigravity #rlvr

  8. [Перевод] Внутри Claude нашли сознание у моделей. J-пространство в LLM

    Привет, Хабр. 6 июля 2026 года Anthropic опубликовала исследование Verbalizable Representations Form a Global Workspace in Language Models . Разбираю технические детали и практические следствия. Пока вы читаете это предложение, ваш мозг делает десятки вещей, которых вы не замечаете: держит осанку, регулирует дыхание, превращает линии на экране в слова. Но часть работы мозга вам всё-таки доступна - всплывший образ, обдуманное решение, куда направить внимание. Нейробиологи называют такую активность сознательно доступной и отличают её от всей остальной, идущей без нашего ведома. У неё особые свойства: её можно описать словами, ею можно управлять и рассуждать -> в отличие от автоматических процессов, которые текут сами собой. В новой работе Anthropic показывает, что похожее разделение появилось и в языковых моделях вроде Claude. У модели обнаружился небольшой набор внутренних нейронных паттернов, которые - на фоне всей остальной обработки играют особую роль.

    habr.com/ru/articles/1056320/

    #claude #anthropic #chatgpt #evals #archive #codex #claudecode #research #antigravity #rlvr

  9. [Перевод] Внутри Claude нашли сознание у моделей. J-пространство в LLM

    Привет, Хабр. 6 июля 2026 года Anthropic опубликовала исследование Verbalizable Representations Form a Global Workspace in Language Models . Разбираю технические детали и практические следствия. Пока вы читаете это предложение, ваш мозг делает десятки вещей, которых вы не замечаете: держит осанку, регулирует дыхание, превращает линии на экране в слова. Но часть работы мозга вам всё-таки доступна - всплывший образ, обдуманное решение, куда направить внимание. Нейробиологи называют такую активность сознательно доступной и отличают её от всей остальной, идущей без нашего ведома. У неё особые свойства: её можно описать словами, ею можно управлять и рассуждать -> в отличие от автоматических процессов, которые текут сами собой. В новой работе Anthropic показывает, что похожее разделение появилось и в языковых моделях вроде Claude. У модели обнаружился небольшой набор внутренних нейронных паттернов, которые - на фоне всей остальной обработки играют особую роль.

    habr.com/ru/articles/1056320/

    #claude #anthropic #chatgpt #evals #archive #codex #claudecode #research #antigravity #rlvr

  10. ICLR 2026 tổng hợp: Cộng đồng nghiên cứu tập trung vào GRPO (157 bài) thay vì DPO, ưu tiên RLVR (125 bài) thay vì RLHF, và 202 bài về Mamba/SSMs. Nait (tuning thông minh chỉ 10% dữ liệu) giúp tối ưu hiệu quả. 257 bài về tính toán lúc test, 123 bài về hallucination. Cảnh báo: mô hình tuân thủ tốt dễ bị tấn công injection. #AI #HọcMáy #ICLR2026 #NCKH #DeepLearning #Mamba #RLVR #GRPO #MạngNeural #BảoMậtAI #ViễnTưởngAI

    reddit.com/r/LocalLLaMA/commen

  11. 🧠 Mới! Notebook code RLVR kết hợp GRPO từ đầu, được chia sẻ trong dự án “Reasoning‑from‑Scratch”. Hữu ích cho những ai muốn khám phá mô hình RL và tối ưu hoá trong AI/ML. #AI #MachineLearning #RLVR #GRPO #LậpTrình #MãNguồn

    reddit.com/r/LocalLLaMA/commen

  12. RLVR promises faster sampling but leaves reasoning untouched—base LLMs still carry the heavy‑lifting of trajectories. The paper (NeurIPS 2025) shows that gains come from smarter teacher‑distillation and minor architectural tweaks, not a new reasoning engine. Curious how sampling efficiency separates from true understanding? Dive into the details. #RLVR #SamplingEfficiency #LLMReasoning #NeurIPS2025

    🔗 aidailypost.com/news/rlvr-lift

  13. 2025 saw significant advancements in #LLMs, with #ReinforcementLearning from #VerifiableRewards (#RLVR) emerging as a key stage in training, leading to improved #reasoning capabilities. The industry also began to understand the unique “jagged” intelligence of LLMs, excelling in specific domains but lacking generalisation. karpathy.bearblog.dev/year-in- #tech #media #news

  14. 2025 saw significant advancements in #LLMs, with #ReinforcementLearning from #VerifiableRewards (#RLVR) emerging as a key stage in training, leading to improved #reasoning capabilities. The industry also began to understand the unique “jagged” intelligence of LLMs, excelling in specific domains but lacking generalisation. karpathy.bearblog.dev/year-in- #tech #media #news

  15. 2025 saw significant advancements in #LLMs, with #ReinforcementLearning from #VerifiableRewards (#RLVR) emerging as a key stage in training, leading to improved #reasoning capabilities. The industry also began to understand the unique “jagged” intelligence of LLMs, excelling in specific domains but lacking generalisation. karpathy.bearblog.dev/year-in- #tech #media #news

  16. 2025 saw significant advancements in #LLMs, with #ReinforcementLearning from #VerifiableRewards (#RLVR) emerging as a key stage in training, leading to improved #reasoning capabilities. The industry also began to understand the unique “jagged” intelligence of LLMs, excelling in specific domains but lacking generalisation. karpathy.bearblog.dev/year-in- #tech #media #news

  17. 2025 saw significant advancements in #LLMs, with #ReinforcementLearning from #VerifiableRewards (#RLVR) emerging as a key stage in training, leading to improved #reasoning capabilities. The industry also began to understand the unique “jagged” intelligence of LLMs, excelling in specific domains but lacking generalisation. karpathy.bearblog.dev/year-in- #tech #media #news

  18. New research from Tsinghua shows that reasoning‑augmented LLMs solve tasks with fewer calls but don’t surpass raw capability. The study compares chain‑of‑thought prompting, RL‑based RLVR, and pass@1 metrics, highlighting efficiency gains for open‑source models. Worth a read for anyone tracking LLM benchmarks. #LLM #ChainOfThought #RLVR #PassAt1

    🔗 aidailypost.com/news/study-fin

  19. New research from Tsinghua shows that reasoning‑augmented LLMs solve tasks with fewer calls but don’t surpass raw capability. The study compares chain‑of‑thought prompting, RL‑based RLVR, and pass@1 metrics, highlighting efficiency gains for open‑source models. Worth a read for anyone tracking LLM benchmarks. #LLM #ChainOfThought #RLVR #PassAt1

    🔗 aidailypost.com/news/study-fin

  20. New research from Tsinghua shows that reasoning‑augmented LLMs solve tasks with fewer calls but don’t surpass raw capability. The study compares chain‑of‑thought prompting, RL‑based RLVR, and pass@1 metrics, highlighting efficiency gains for open‑source models. Worth a read for anyone tracking LLM benchmarks. #LLM #ChainOfThought #RLVR #PassAt1

    🔗 aidailypost.com/news/study-fin

  21. New research from Tsinghua shows that reasoning‑augmented LLMs solve tasks with fewer calls but don’t surpass raw capability. The study compares chain‑of‑thought prompting, RL‑based RLVR, and pass@1 metrics, highlighting efficiency gains for open‑source models. Worth a read for anyone tracking LLM benchmarks. #LLM #ChainOfThought #RLVR #PassAt1

    🔗 aidailypost.com/news/study-fin

  22. → Les 4 étapes pour entrainer un LLM
    scienceetonnante.com/blog/2025

    « Voilà le principe de l'apprentissage par renforcement avec une récompense vérifiable [RLVR], qui permet de se passer d'humains qui doivent juger si la réponse est conforme ou pas. »

    #entrainer #LLM #apprentissage #RLVR #humains

  23. → Les 4 étapes pour entrainer un LLM
    scienceetonnante.com/blog/2025

    « Voilà le principe de l'apprentissage par renforcement avec une récompense vérifiable [RLVR], qui permet de se passer d'humains qui doivent juger si la réponse est conforme ou pas. »

    #entrainer #LLM #apprentissage #RLVR #humains

  24. → Les 4 étapes pour entrainer un LLM
    scienceetonnante.com/blog/2025

    « Voilà le principe de l'apprentissage par renforcement avec une récompense vérifiable [RLVR], qui permet de se passer d'humains qui doivent juger si la réponse est conforme ou pas. »

    #entrainer #LLM #apprentissage #RLVR #humains

  25. → Les 4 étapes pour entrainer un LLM
    scienceetonnante.com/blog/2025

    « Voilà le principe de l'apprentissage par renforcement avec une récompense vérifiable [RLVR], qui permet de se passer d'humains qui doivent juger si la réponse est conforme ou pas. »

    #entrainer #LLM #apprentissage #RLVR #humains

  26. → Les 4 étapes pour entrainer un LLM
    scienceetonnante.com/blog/2025

    « Voilà le principe de l'apprentissage par renforcement avec une récompense vérifiable [RLVR], qui permet de se passer d'humains qui doivent juger si la réponse est conforme ou pas. »

    #entrainer #LLM #apprentissage #RLVR #humains

  27. New study challenges a key belief about Reinforcement Learning with Verifiable Rewards (RLVR) for #LLMs:
    #RLVR boosts efficiency but doesn't create new reasoning skills — #AI base models already had them!
    arxiv.org/abs/2504.13837

  28. New study challenges a key belief about Reinforcement Learning with Verifiable Rewards (RLVR) for #LLMs:
    #RLVR boosts efficiency but doesn't create new reasoning skills — #AI base models already had them!
    arxiv.org/abs/2504.13837

  29. New study challenges a key belief about Reinforcement Learning with Verifiable Rewards (RLVR) for #LLMs:
    #RLVR boosts efficiency but doesn't create new reasoning skills — #AI base models already had them!
    arxiv.org/abs/2504.13837

  30. New study challenges a key belief about Reinforcement Learning with Verifiable Rewards (RLVR) for #LLMs:
    #RLVR boosts efficiency but doesn't create new reasoning skills — #AI base models already had them!
    arxiv.org/abs/2504.13837

  31. New study challenges a key belief about Reinforcement Learning with Verifiable Rewards (RLVR) for #LLMs:
    #RLVR boosts efficiency but doesn't create new reasoning skills — #AI base models already had them!
    arxiv.org/abs/2504.13837

  32. Forschende der Tsinghua University und der Shanghai Jiao Tong University zeigen in einer Studie, dass #RLVR zwar die Wahrscheinlichkeit erhöht, beim ersten Versuch eine richtige Antwort zu generieren – das sogenannte pass@1 –, aber keine neuen Problemlösestrategien erschließt.
    the-decoder.de/forscher-zweife
    #KI

  33. Forschende der Tsinghua University und der Shanghai Jiao Tong University zeigen in einer Studie, dass #RLVR zwar die Wahrscheinlichkeit erhöht, beim ersten Versuch eine richtige Antwort zu generieren – das sogenannte pass@1 –, aber keine neuen Problemlösestrategien erschließt.
    the-decoder.de/forscher-zweife
    #KI

  34. Forschende der Tsinghua University und der Shanghai Jiao Tong University zeigen in einer Studie, dass #RLVR zwar die Wahrscheinlichkeit erhöht, beim ersten Versuch eine richtige Antwort zu generieren – das sogenannte pass@1 –, aber keine neuen Problemlösestrategien erschließt.
    the-decoder.de/forscher-zweife
    #KI

  35. Forschende der Tsinghua University und der Shanghai Jiao Tong University zeigen in einer Studie, dass #RLVR zwar die Wahrscheinlichkeit erhöht, beim ersten Versuch eine richtige Antwort zu generieren – das sogenannte pass@1 –, aber keine neuen Problemlösestrategien erschließt.
    the-decoder.de/forscher-zweife
    #KI

  36. Forschende der Tsinghua University und der Shanghai Jiao Tong University zeigen in einer Studie, dass #RLVR zwar die Wahrscheinlichkeit erhöht, beim ersten Versuch eine richtige Antwort zu generieren – das sogenannte pass@1 –, aber keine neuen Problemlösestrategien erschließt.
    the-decoder.de/forscher-zweife
    #KI

  37. • 🧠 Advanced post-training with reinforcement learning with verifiable rewards (#RLVR) using Group Relative Policy Optimization

    • 🔮 All models available in 7B, 13B, and 32B sizes, can be fine-tuned on a single H100 GPU

  38. • 🧠 Advanced post-training with reinforcement learning with verifiable rewards (#RLVR) using Group Relative Policy Optimization

    • 🔮 All models available in 7B, 13B, and 32B sizes, can be fine-tuned on a single H100 GPU

  39. • 🧠 Advanced post-training with reinforcement learning with verifiable rewards (#RLVR) using Group Relative Policy Optimization

    • 🔮 All models available in 7B, 13B, and 32B sizes, can be fine-tuned on a single H100 GPU

  40. • 🧠 Advanced post-training with reinforcement learning with verifiable rewards (#RLVR) using Group Relative Policy Optimization

    • 🔮 All models available in 7B, 13B, and 32B sizes, can be fine-tuned on a single H100 GPU

  41. • 🧠 Advanced post-training with reinforcement learning with verifiable rewards (#RLVR) using Group Relative Policy Optimization

    • 🔮 All models available in 7B, 13B, and 32B sizes, can be fine-tuned on a single H100 GPU