#rlvr — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #rlvr, aggregated by home.social.
-
[Перевод] Как управлять reasoning effort в LLM
Как современные LLM получают разные настройки reasoning effort: системные промпты, SFT, RLVR, штрафы за длину, reasoning budgets и on-policy-дистилляция. Разбираем DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5 и K3, GLM-5, Qwen3 и Inkling.
https://habr.com/ru/articles/1071250/
#LLM #reasoning_models #reasoning_effort #RLVR #reinforcement_learning #inference_scaling
-
[Перевод] Как управлять reasoning effort в LLM
Как современные LLM получают разные настройки reasoning effort: системные промпты, SFT, RLVR, штрафы за длину, reasoning budgets и on-policy-дистилляция. Разбираем DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5 и K3, GLM-5, Qwen3 и Inkling.
https://habr.com/ru/articles/1071250/
#LLM #reasoning_models #reasoning_effort #RLVR #reinforcement_learning #inference_scaling
-
[Перевод] Как управлять reasoning effort в LLM
Как современные LLM получают разные настройки reasoning effort: системные промпты, SFT, RLVR, штрафы за длину, reasoning budgets и on-policy-дистилляция. Разбираем DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5 и K3, GLM-5, Qwen3 и Inkling.
https://habr.com/ru/articles/1071250/
#LLM #reasoning_models #reasoning_effort #RLVR #reinforcement_learning #inference_scaling
-
Данные, которые нельзя выдумать: как мы собираем трейсы для тренировки агентских навыков Cotype
Привет, Хабр! На днях мы выкатили третье поколение языковых моделей Cotype — флагманскую Cotype Pro 3 и ее облегченную версию Cotype Light 3 (к слову, лайт мы выкатили чутка раньше, но это не имеет значения). Почитать о характеристиках моделей можно вот в этом посте. Тут же мы расскажем вам о нюансах обучения — точнее о том, как мы учим наши модели выполнять многошаговые сценарии без запинок, что совершенно необходимо, так как они выступают ключевым компонентом наших корпоративных ИИ-агентов и мультиагентных систем. Велком под кат
https://habr.com/ru/companies/ru_mts/articles/1060112/
#agents #llm #rlvr #grpo #harness #synthetic_data #искусственный_интеллект #агентский_режим #ииагенты #ииагенты_для_бизнеса
-
Данные, которые нельзя выдумать: как мы собираем трейсы для тренировки агентских навыков Cotype
Привет, Хабр! На днях мы выкатили третье поколение языковых моделей Cotype — флагманскую Cotype Pro 3 и ее облегченную версию Cotype Light 3 (к слову, лайт мы выкатили чутка раньше, но это не имеет значения). Почитать о характеристиках моделей можно вот в этом посте. Тут же мы расскажем вам о нюансах обучения — точнее о том, как мы учим наши модели выполнять многошаговые сценарии без запинок, что совершенно необходимо, так как они выступают ключевым компонентом наших корпоративных ИИ-агентов и мультиагентных систем. Велком под кат
https://habr.com/ru/companies/ru_mts/articles/1060112/
#agents #llm #rlvr #grpo #harness #synthetic_data #искусственный_интеллект #агентский_режим #ииагенты #ииагенты_для_бизнеса
-
Данные, которые нельзя выдумать: как мы собираем трейсы для тренировки агентских навыков Cotype
Привет, Хабр! На днях мы выкатили третье поколение языковых моделей Cotype — флагманскую Cotype Pro 3 и ее облегченную версию Cotype Light 3 (к слову, лайт мы выкатили чутка раньше, но это не имеет значения). Почитать о характеристиках моделей можно вот в этом посте. Тут же мы расскажем вам о нюансах обучения — точнее о том, как мы учим наши модели выполнять многошаговые сценарии без запинок, что совершенно необходимо, так как они выступают ключевым компонентом наших корпоративных ИИ-агентов и мультиагентных систем. Велком под кат
https://habr.com/ru/companies/ru_mts/articles/1060112/
#agents #llm #rlvr #grpo #harness #synthetic_data #искусственный_интеллект #агентский_режим #ииагенты #ииагенты_для_бизнеса
-
[Перевод] Внутри Claude нашли сознание у моделей. J-пространство в LLM
Привет, Хабр. 6 июля 2026 года Anthropic опубликовала исследование Verbalizable Representations Form a Global Workspace in Language Models . Разбираю технические детали и практические следствия. Пока вы читаете это предложение, ваш мозг делает десятки вещей, которых вы не замечаете: держит осанку, регулирует дыхание, превращает линии на экране в слова. Но часть работы мозга вам всё-таки доступна - всплывший образ, обдуманное решение, куда направить внимание. Нейробиологи называют такую активность сознательно доступной и отличают её от всей остальной, идущей без нашего ведома. У неё особые свойства: её можно описать словами, ею можно управлять и рассуждать -> в отличие от автоматических процессов, которые текут сами собой. В новой работе Anthropic показывает, что похожее разделение появилось и в языковых моделях вроде Claude. У модели обнаружился небольшой набор внутренних нейронных паттернов, которые - на фоне всей остальной обработки играют особую роль.
https://habr.com/ru/articles/1056320/
#claude #anthropic #chatgpt #evals #archive #codex #claudecode #research #antigravity #rlvr
-
[Перевод] Внутри Claude нашли сознание у моделей. J-пространство в LLM
Привет, Хабр. 6 июля 2026 года Anthropic опубликовала исследование Verbalizable Representations Form a Global Workspace in Language Models . Разбираю технические детали и практические следствия. Пока вы читаете это предложение, ваш мозг делает десятки вещей, которых вы не замечаете: держит осанку, регулирует дыхание, превращает линии на экране в слова. Но часть работы мозга вам всё-таки доступна - всплывший образ, обдуманное решение, куда направить внимание. Нейробиологи называют такую активность сознательно доступной и отличают её от всей остальной, идущей без нашего ведома. У неё особые свойства: её можно описать словами, ею можно управлять и рассуждать -> в отличие от автоматических процессов, которые текут сами собой. В новой работе Anthropic показывает, что похожее разделение появилось и в языковых моделях вроде Claude. У модели обнаружился небольшой набор внутренних нейронных паттернов, которые - на фоне всей остальной обработки играют особую роль.
https://habr.com/ru/articles/1056320/
#claude #anthropic #chatgpt #evals #archive #codex #claudecode #research #antigravity #rlvr
-
[Перевод] Внутри Claude нашли сознание у моделей. J-пространство в LLM
Привет, Хабр. 6 июля 2026 года Anthropic опубликовала исследование Verbalizable Representations Form a Global Workspace in Language Models . Разбираю технические детали и практические следствия. Пока вы читаете это предложение, ваш мозг делает десятки вещей, которых вы не замечаете: держит осанку, регулирует дыхание, превращает линии на экране в слова. Но часть работы мозга вам всё-таки доступна - всплывший образ, обдуманное решение, куда направить внимание. Нейробиологи называют такую активность сознательно доступной и отличают её от всей остальной, идущей без нашего ведома. У неё особые свойства: её можно описать словами, ею можно управлять и рассуждать -> в отличие от автоматических процессов, которые текут сами собой. В новой работе Anthropic показывает, что похожее разделение появилось и в языковых моделях вроде Claude. У модели обнаружился небольшой набор внутренних нейронных паттернов, которые - на фоне всей остальной обработки играют особую роль.
https://habr.com/ru/articles/1056320/
#claude #anthropic #chatgpt #evals #archive #codex #claudecode #research #antigravity #rlvr
-
ICLR 2026 tổng hợp: Cộng đồng nghiên cứu tập trung vào GRPO (157 bài) thay vì DPO, ưu tiên RLVR (125 bài) thay vì RLHF, và 202 bài về Mamba/SSMs. Nait (tuning thông minh chỉ 10% dữ liệu) giúp tối ưu hiệu quả. 257 bài về tính toán lúc test, 123 bài về hallucination. Cảnh báo: mô hình tuân thủ tốt dễ bị tấn công injection. #AI #HọcMáy #ICLR2026 #NCKH #DeepLearning #Mamba #RLVR #GRPO #MạngNeural #BảoMậtAI #ViễnTưởngAI
https://www.reddit.com/r/LocalLLaMA/comments/1qsh7dz/analyzed_5357_iclr_2026_acc
-
🧠 Mới! Notebook code RLVR kết hợp GRPO từ đầu, được chia sẻ trong dự án “Reasoning‑from‑Scratch”. Hữu ích cho những ai muốn khám phá mô hình RL và tối ưu hoá trong AI/ML. #AI #MachineLearning #RLVR #GRPO #LậpTrình #MãNguồn
https://www.reddit.com/r/LocalLLaMA/comments/1qgcj8b/rlvr_with_grpo_from_scratch_code_notebook/
-
RLVR promises faster sampling but leaves reasoning untouched—base LLMs still carry the heavy‑lifting of trajectories. The paper (NeurIPS 2025) shows that gains come from smarter teacher‑distillation and minor architectural tweaks, not a new reasoning engine. Curious how sampling efficiency separates from true understanding? Dive into the details. #RLVR #SamplingEfficiency #LLMReasoning #NeurIPS2025
🔗 https://aidailypost.com/news/rlvr-lifts-sampling-efficiency-not-reasoning-base-models-hold
-
2025 saw significant advancements in #LLMs, with #ReinforcementLearning from #VerifiableRewards (#RLVR) emerging as a key stage in training, leading to improved #reasoning capabilities. The industry also began to understand the unique “jagged” intelligence of LLMs, excelling in specific domains but lacking generalisation. https://karpathy.bearblog.dev/year-in-review-2025/?eicker.news #tech #media #news
-
2025 saw significant advancements in #LLMs, with #ReinforcementLearning from #VerifiableRewards (#RLVR) emerging as a key stage in training, leading to improved #reasoning capabilities. The industry also began to understand the unique “jagged” intelligence of LLMs, excelling in specific domains but lacking generalisation. https://karpathy.bearblog.dev/year-in-review-2025/?eicker.news #tech #media #news
-
2025 saw significant advancements in #LLMs, with #ReinforcementLearning from #VerifiableRewards (#RLVR) emerging as a key stage in training, leading to improved #reasoning capabilities. The industry also began to understand the unique “jagged” intelligence of LLMs, excelling in specific domains but lacking generalisation. https://karpathy.bearblog.dev/year-in-review-2025/?eicker.news #tech #media #news
-
2025 saw significant advancements in #LLMs, with #ReinforcementLearning from #VerifiableRewards (#RLVR) emerging as a key stage in training, leading to improved #reasoning capabilities. The industry also began to understand the unique “jagged” intelligence of LLMs, excelling in specific domains but lacking generalisation. https://karpathy.bearblog.dev/year-in-review-2025/?eicker.news #tech #media #news
-
2025 saw significant advancements in #LLMs, with #ReinforcementLearning from #VerifiableRewards (#RLVR) emerging as a key stage in training, leading to improved #reasoning capabilities. The industry also began to understand the unique “jagged” intelligence of LLMs, excelling in specific domains but lacking generalisation. https://karpathy.bearblog.dev/year-in-review-2025/?eicker.news #tech #media #news
-
New research from Tsinghua shows that reasoning‑augmented LLMs solve tasks with fewer calls but don’t surpass raw capability. The study compares chain‑of‑thought prompting, RL‑based RLVR, and pass@1 metrics, highlighting efficiency gains for open‑source models. Worth a read for anyone tracking LLM benchmarks. #LLM #ChainOfThought #RLVR #PassAt1
🔗 https://aidailypost.com/news/study-finds-reasoning-llms-are-more-efficient-not-more-capable
-
New research from Tsinghua shows that reasoning‑augmented LLMs solve tasks with fewer calls but don’t surpass raw capability. The study compares chain‑of‑thought prompting, RL‑based RLVR, and pass@1 metrics, highlighting efficiency gains for open‑source models. Worth a read for anyone tracking LLM benchmarks. #LLM #ChainOfThought #RLVR #PassAt1
🔗 https://aidailypost.com/news/study-finds-reasoning-llms-are-more-efficient-not-more-capable
-
New research from Tsinghua shows that reasoning‑augmented LLMs solve tasks with fewer calls but don’t surpass raw capability. The study compares chain‑of‑thought prompting, RL‑based RLVR, and pass@1 metrics, highlighting efficiency gains for open‑source models. Worth a read for anyone tracking LLM benchmarks. #LLM #ChainOfThought #RLVR #PassAt1
🔗 https://aidailypost.com/news/study-finds-reasoning-llms-are-more-efficient-not-more-capable
-
New research from Tsinghua shows that reasoning‑augmented LLMs solve tasks with fewer calls but don’t surpass raw capability. The study compares chain‑of‑thought prompting, RL‑based RLVR, and pass@1 metrics, highlighting efficiency gains for open‑source models. Worth a read for anyone tracking LLM benchmarks. #LLM #ChainOfThought #RLVR #PassAt1
🔗 https://aidailypost.com/news/study-finds-reasoning-llms-are-more-efficient-not-more-capable
-
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
https://arxiv.org/abs/2507.15855
#HackerNews #Implicit #Actor #Critic #Coupling #Supervised #Learning #RLVR #ReinforcementLearning
-
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
https://arxiv.org/abs/2507.15855
#HackerNews #Implicit #Actor #Critic #Coupling #Supervised #Learning #RLVR #ReinforcementLearning
-
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
https://arxiv.org/abs/2507.15855
#HackerNews #Implicit #Actor #Critic #Coupling #Supervised #Learning #RLVR #ReinforcementLearning
-
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
https://arxiv.org/abs/2507.15855
#HackerNews #Implicit #Actor #Critic #Coupling #Supervised #Learning #RLVR #ReinforcementLearning
-
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
https://arxiv.org/abs/2507.15855
#HackerNews #Implicit #Actor #Critic #Coupling #Supervised #Learning #RLVR #ReinforcementLearning
-
→ Les 4 étapes pour entrainer un LLM
https://scienceetonnante.com/blog/2025/04/25/les-4-etapes-pour-entrainer-un-llm/« Voilà le principe de l'apprentissage par renforcement avec une récompense vérifiable [RLVR], qui permet de se passer d'humains qui doivent juger si la réponse est conforme ou pas. »
-
→ Les 4 étapes pour entrainer un LLM
https://scienceetonnante.com/blog/2025/04/25/les-4-etapes-pour-entrainer-un-llm/« Voilà le principe de l'apprentissage par renforcement avec une récompense vérifiable [RLVR], qui permet de se passer d'humains qui doivent juger si la réponse est conforme ou pas. »
-
→ Les 4 étapes pour entrainer un LLM
https://scienceetonnante.com/blog/2025/04/25/les-4-etapes-pour-entrainer-un-llm/« Voilà le principe de l'apprentissage par renforcement avec une récompense vérifiable [RLVR], qui permet de se passer d'humains qui doivent juger si la réponse est conforme ou pas. »
-
→ Les 4 étapes pour entrainer un LLM
https://scienceetonnante.com/blog/2025/04/25/les-4-etapes-pour-entrainer-un-llm/« Voilà le principe de l'apprentissage par renforcement avec une récompense vérifiable [RLVR], qui permet de se passer d'humains qui doivent juger si la réponse est conforme ou pas. »
-
→ Les 4 étapes pour entrainer un LLM
https://scienceetonnante.com/blog/2025/04/25/les-4-etapes-pour-entrainer-un-llm/« Voilà le principe de l'apprentissage par renforcement avec une récompense vérifiable [RLVR], qui permet de se passer d'humains qui doivent juger si la réponse est conforme ou pas. »
-
New study challenges a key belief about Reinforcement Learning with Verifiable Rewards (RLVR) for #LLMs:
#RLVR boosts efficiency but doesn't create new reasoning skills — #AI base models already had them!
https://arxiv.org/abs/2504.13837 -
New study challenges a key belief about Reinforcement Learning with Verifiable Rewards (RLVR) for #LLMs:
#RLVR boosts efficiency but doesn't create new reasoning skills — #AI base models already had them!
https://arxiv.org/abs/2504.13837 -
New study challenges a key belief about Reinforcement Learning with Verifiable Rewards (RLVR) for #LLMs:
#RLVR boosts efficiency but doesn't create new reasoning skills — #AI base models already had them!
https://arxiv.org/abs/2504.13837 -
New study challenges a key belief about Reinforcement Learning with Verifiable Rewards (RLVR) for #LLMs:
#RLVR boosts efficiency but doesn't create new reasoning skills — #AI base models already had them!
https://arxiv.org/abs/2504.13837 -
New study challenges a key belief about Reinforcement Learning with Verifiable Rewards (RLVR) for #LLMs:
#RLVR boosts efficiency but doesn't create new reasoning skills — #AI base models already had them!
https://arxiv.org/abs/2504.13837 -
Forschende der Tsinghua University und der Shanghai Jiao Tong University zeigen in einer Studie, dass #RLVR zwar die Wahrscheinlichkeit erhöht, beim ersten Versuch eine richtige Antwort zu generieren – das sogenannte pass@1 –, aber keine neuen Problemlösestrategien erschließt.
https://the-decoder.de/forscher-zweifeln-an-reasoning-modellen-effizienter-ja-intelligenter-nein/
#KI -
Forschende der Tsinghua University und der Shanghai Jiao Tong University zeigen in einer Studie, dass #RLVR zwar die Wahrscheinlichkeit erhöht, beim ersten Versuch eine richtige Antwort zu generieren – das sogenannte pass@1 –, aber keine neuen Problemlösestrategien erschließt.
https://the-decoder.de/forscher-zweifeln-an-reasoning-modellen-effizienter-ja-intelligenter-nein/
#KI -
Forschende der Tsinghua University und der Shanghai Jiao Tong University zeigen in einer Studie, dass #RLVR zwar die Wahrscheinlichkeit erhöht, beim ersten Versuch eine richtige Antwort zu generieren – das sogenannte pass@1 –, aber keine neuen Problemlösestrategien erschließt.
https://the-decoder.de/forscher-zweifeln-an-reasoning-modellen-effizienter-ja-intelligenter-nein/
#KI -
Forschende der Tsinghua University und der Shanghai Jiao Tong University zeigen in einer Studie, dass #RLVR zwar die Wahrscheinlichkeit erhöht, beim ersten Versuch eine richtige Antwort zu generieren – das sogenannte pass@1 –, aber keine neuen Problemlösestrategien erschließt.
https://the-decoder.de/forscher-zweifeln-an-reasoning-modellen-effizienter-ja-intelligenter-nein/
#KI -
Forschende der Tsinghua University und der Shanghai Jiao Tong University zeigen in einer Studie, dass #RLVR zwar die Wahrscheinlichkeit erhöht, beim ersten Versuch eine richtige Antwort zu generieren – das sogenannte pass@1 –, aber keine neuen Problemlösestrategien erschließt.
https://the-decoder.de/forscher-zweifeln-an-reasoning-modellen-effizienter-ja-intelligenter-nein/
#KI -
Does RL Incentivize Reasoning in LLMs Beyond the Base Model?
https://limit-of-rlvr.github.io/
#ycombinator #Qwen #Deepseek_R1 #PPO #GRPO #AIME #RLVR #Tsinghua_University -
Does RL Incentivize Reasoning in LLMs Beyond the Base Model?
https://limit-of-rlvr.github.io/
#ycombinator #Qwen #Deepseek_R1 #PPO #GRPO #AIME #RLVR #Tsinghua_University -
Does RL Incentivize Reasoning in LLMs Beyond the Base Model?
https://limit-of-rlvr.github.io/
#ycombinator #Qwen #Deepseek_R1 #PPO #GRPO #AIME #RLVR #Tsinghua_University -
Does RL Incentivize Reasoning in LLMs Beyond the Base Model?
https://limit-of-rlvr.github.io/
#ycombinator #Qwen #Deepseek_R1 #PPO #GRPO #AIME #RLVR #Tsinghua_University -
• 🧠 Advanced post-training with reinforcement learning with verifiable rewards (#RLVR) using Group Relative Policy Optimization
• 🔮 All models available in 7B, 13B, and 32B sizes, can be fine-tuned on a single H100 GPU
-
• 🧠 Advanced post-training with reinforcement learning with verifiable rewards (#RLVR) using Group Relative Policy Optimization
• 🔮 All models available in 7B, 13B, and 32B sizes, can be fine-tuned on a single H100 GPU
-
• 🧠 Advanced post-training with reinforcement learning with verifiable rewards (#RLVR) using Group Relative Policy Optimization
• 🔮 All models available in 7B, 13B, and 32B sizes, can be fine-tuned on a single H100 GPU
-
• 🧠 Advanced post-training with reinforcement learning with verifiable rewards (#RLVR) using Group Relative Policy Optimization
• 🔮 All models available in 7B, 13B, and 32B sizes, can be fine-tuned on a single H100 GPU
-
• 🧠 Advanced post-training with reinforcement learning with verifiable rewards (#RLVR) using Group Relative Policy Optimization
• 🔮 All models available in 7B, 13B, and 32B sizes, can be fine-tuned on a single H100 GPU
-
Alibaba’s R1-Omni AI Model Expands the Frontier of Emotion Recognition
#AI #AlibabaAI #GenAI #R1Omni #EmotionRecognition #China #OpenSourceAI #RLVR #AIModels
-
Alibaba’s R1-Omni AI Model Expands the Frontier of Emotion Recognition
#AI #AlibabaAI #GenAI #R1Omni #EmotionRecognition #China #OpenSourceAI #RLVR #AIModels
-
Alibaba’s R1-Omni AI Model Expands the Frontier of Emotion Recognition
#AI #AlibabaAI #GenAI #R1Omni #EmotionRecognition #China #OpenSourceAI #RLVR #AIModels
-
Alibaba’s R1-Omni AI Model Expands the Frontier of Emotion Recognition
#AI #AlibabaAI #GenAI #R1Omni #EmotionRecognition #China #OpenSourceAI #RLVR #AIModels
-
Alibaba’s R1-Omni AI Model Expands the Frontier of Emotion Recognition
#AI #AlibabaAI #GenAI #R1Omni #EmotionRecognition #China #OpenSourceAI #RLVR #AIModels