home.social

#ppo — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #ppo, aggregated by home.social.

fetched live
  1. Мы обошли фильтр Калмана на одной камере, но сначала три недели измеряли труп

    Цель видна только через камеру, детектор отдает рамку, дальномера нет. В какой-то момент цель уходит из кадра на три четверти секунды, и ее надо продолжать вести вслепую. Классический фильтр Калмана в этой ситуации теряет цель в 84% эпизодов. Наш трекер теряет ее в 18% и возвращает в центр кадра в 80%.

    habr.com/ru/articles/1069900/

    #обучение_с_подкреплением #PPO #фильтр_Калмана #KalmanNet #simtoreal #компьютерное_зрение #трекинг_объектов #LSTM #reward_shaping #мирмодель

  2. Походка за двадцать минут и миллион рублей: что RL сделал с двуногими роботами и во что упёрся их «мозг»

    За один 2026 год двуногие роботы успели пробежать полумарафон быстрее человеческого рекорда, довести публику до того, что CEO пришлось резать роботу ногу ножницами прямо на сцене, и провалить задачу «пройтись ровно» на презентации за миллионы долларов. Разбираю с первоисточниками, почему походка стала дешёвой инженерией — 20 минут на одной видеокарте, — а «живой» робот упирается в ватты, переполняющийся контекст и отсутствие непрерывного обучения.

    habr.com/ru/articles/1057816/

    #reinforcement_learning #RLлокомоция #гуманоидные_роботы #simtoreal #Isaac_Lab #VLAмодели #PPO #world_models #робототехника #KVcache

  3. Продвинутые RL алгоритмы: Normal Policy, TRPO, PPO

    Большой конспект по продвинутым RL алгоритмам: TRPO и PPO. Автор слегка упоролся в формулах, но это из любви к прозрачности алгоритмов.

    habr.com/ru/articles/991622/

    #Policy_gradient_methods #ActorCritic #reinforcementlearning #ppo #trpo

  4. RL (RLM): Разбираемся вместе

    Всем привет! Недавно я познакомился с курсом по глубокому обучению с подкреплением от HuggingFace Deep Reinforcement Learning Course и захотел сделать выжимку самого интересного. Эта статья — своего рода шпаргалка по основам Reinforcement Learning (RL) и одному из ключевых алгоритмов — PPO, который лежит в основе тонкой настройки современных LLM (Large Language Models).

    habr.com/ru/articles/958062/

    #Искуственный_интеллект #Машинное_обучение #Алгоритмы #RLHF #LLM #Большие_языковые_модели #RL #Reinforcement_learning #PPO #Proxi

  5. A Vulnerable Sector Check (VSC) pre-employment screening can take over 3 months because of a backlog at the OPP.

    cbc.ca/news/canada/toronto/opp
    - - -
    La vérification des antécédents en vue d’un travail auprès de personnels vulnérables (VATPV) peut prendre plus de 3 mois à cause de retards chez la PPO.

    // Article en anglais //

    #Ontario #OPP #PPO

  6. A Vulnerable Sector Check (VSC) pre-employment screening can take over 3 months because of a backlog at the OPP.

    cbc.ca/news/canada/toronto/opp
    - - -
    La vérification des antécédents en vue d’un travail auprès de personnels vulnérables (VATPV) peut prendre plus de 3 mois à cause de retards chez la PPO.

    // Article en anglais //

    #Ontario #OPP #PPO

  7. #MedicalInsurance #Medicare #MedicarePlus

    Just received noticed from #BlueShield that #UCSF, my medical provider for the last 15 years, is leaving the #BlueShieldOfCA #PPO medical network as of 7/10/2025. ☹️

    Just started doing some research on which groups are available where I can find a new PCP & all of the "reviews" for all of the medical groups in my area & beyond are dismal. 🤦‍♂️

    That said, I've found that as long as I get a PCP that I get along with & who is responsive to my needs/requests, I'm happy even if the reviews for the group are poor.

    So, I may need to try a couple in various groups before I find the PCP that I like.

    As the member of a PPO, I don't have to worry all about getting referrals for specialized care but the day-to-day medical care -- labs & prescriptions -- is all I generally need & I just need to find another PCP who is on the same page with me for those things.

    Wish me luck! 😉

  8. #MedicalInsurance #Medicare #MedicarePlus

    Just received noticed from #BlueShield that #UCSF, my medical provider for the last 15 years, is leaving the #BlueShieldOfCA #PPO medical network as of 7/10/2025. ☹️

    Just started doing some research on which groups are available where I can find a new PCP & all of the "reviews" for all of the medical groups in my area & beyond are dismal. 🤦‍♂️

    That said, I've found that as long as I get a PCP that I get along with & who is responsive to my needs/requests, I'm happy even if the reviews for the group are poor.

    So, I may need to try a couple in various groups before I find the PCP that I like.

    As the member of a PPO, I don't have to worry all about getting referrals for specialized care but the day-to-day medical care -- labs & prescriptions -- is all I generally need & I just need to find another PCP who is on the same page with me for those things.

    Wish me luck! 😉

  9. Mobile PPO groups in Kherson region hit "shaheeds" at night, destroying 6 enemy drones. #Kherson #PPO

  10. Mobile PPO groups in Kherson region hit "shaheeds" at night, destroying 6 enemy drones. #Kherson #PPO

  11. In Kyiv, PPO@censor_net operates. They are active on social media platforms. #Kyiv #PPO

  12. In Kyiv, PPO@censor_net operates. They are active on social media platforms. #Kyiv #PPO

  13. "PPO is operating loudly in the Dnipro region." #Dnipro #PPO

  14. Using clever change of variables trick #DPO is a more efficient drop-in replacement for #PPO in #RLHF.

    Using DPO with preference labels from #chatbot panel of judges for virtually embodied agents would be a great way to achieve an unambiguous #AGI.

    [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Model arxiv.org/abs/2305.18290

  15. Using clever change of variables trick #DPO is a more efficient drop-in replacement for #PPO in #RLHF.

    Using DPO with preference labels from #chatbot panel of judges for virtually embodied agents would be a great way to achieve an unambiguous #AGI.

    [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Model arxiv.org/abs/2305.18290

  16. Today I’ve started to make an humanoid robot learn to walk by itself. Funny to watch when evaluating the model after a few hours.

    #ai #gym #deeplearning #ppo

  17. Details zur Technik hinter #ChatGPT erklärt mir @ct_Magazin Redakteurin Pina Merkert in diesem c't uplink kompakt

    youtube.com/watch?v=jcrBBxXK36

    Um Anwendung und Auswirkung von #ChatGPT geht es dann am Samstag im #ctuplink 47.0, wo @johoo, @wstieler und Hartmut Gieselmann meine Gäste sind.

    #ChatGPT #GPT3 #OpenAI #AI #ArtificialIntelligence #KI #ML #MachineLearning #Transformer #PPO #NeuronaleNetze #KünstlicheIntelligenz #ctmagazin #uplink #uplinkkompakt