home.social

#reinforcement — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #reinforcement, aggregated by home.social.

fetched live
  1. A quotation from Eric Hoffer

    The truth seems to be that propaganda on its own cannot force its way into unwilling minds; neither can it inculcate something wholly new; nor can it keep people persuaded once they have ceased to believe. It penetrates only into minds already open, and rather than instill opinion it articulates and justifies opinions already present in the minds of its recipients. The gifted propagandist brings to a boil ideas and passions already simmering in the minds of his hearers. he echoes their innermost feelings. Where opinion is not coerced, people can be made to believe only in what they already “know.”

    Eric Hoffer (1902-1983) American writer, philosopher, longshoreman
    True Believer: Thoughts on the Nature of Mass Movements, Part 3, ch. 14, § 83 (3.14.83) (1951)

    More about this quote: wist.info/hoffer-eric/11263/

    #quote #quotes #quotation #qotd #erichoffer #agitprop #belief #confirmation #confirmationbias #conviction #disinformation #fear #ideas #justification #opinion #passion #predisposition #prejudice #propaganda #reinforcement #truebeliever

  2. A quotation from Eric Hoffer

    The truth seems to be that propaganda on its own cannot force its way into unwilling minds; neither can it inculcate something wholly new; nor can it keep people persuaded once they have ceased to believe. It penetrates only into minds already open, and rather than instill opinion it articulates and justifies opinions already present in the minds of its recipients. The gifted propagandist brings to a boil ideas and passions already simmering in the minds of his hearers. he echoes their innermost feelings. Where opinion is not coerced, people can be made to believe only in what they already “know.”

    Eric Hoffer (1902-1983) American writer, philosopher, longshoreman
    True Believer: Thoughts on the Nature of Mass Movements, Part 3, ch. 14, § 83 (3.14.83) (1951)

    More about this quote: wist.info/hoffer-eric/11263/

    #quote #quotes #quotation #qotd #erichoffer #agitprop #belief #confirmation #confirmationbias #conviction #disinformation #fear #ideas #justification #opinion #passion #predisposition #prejudice #propaganda #reinforcement #truebeliever

  3. How to Explore to Scale RL Training of LLMs on Hard Problems? LLM RL typically operates in one of three exploration regimes: sharpening, chaining, or guided exploration; standard RL stays in the fi...

    #machine #learning #reinforcement #learning #Research

    Origin | Interest | Match
  4. How to Explore to Scale RL Training of LLMs on Hard Problems? LLM RL typically operates in one of three exploration regimes: sharpening, chaining, or guided exploration; standard RL stays in the fi...

    #machine #learning #reinforcement #learning #Research

    Origin | Interest | Match
  5. A quotation from Mignon McLaughlin

    Most of us would try to be noble, if we just had a claque we could depend on.

    Mignon McLaughlin (1913-1983) American journalist and author
    The Neurotic’s Notebook, ch. 6 (1963)

    More info about this quote: wist.info/mclaughlin-mignon/80…

    #quote #quotes #quotation #qotd #mignonmclaughlin #admiration #applause #audience #nobility #recognition #reinforcement #virtue #supporters

  6. A quotation from Samuel Johnson

    To be of no church is dangerous. Religion, of which the rewards are distant, and which is animated only by faith and hope, will glide by degrees out of the mind unless it be invigorated and reimpressed by external ordinances, by stated calls to worship, and the salutary influence of example.

    Samuel Johnson (1709-1784) English writer, lexicographer, critic
    Lives of the Most Eminent English Poets, “Milton” (1781)

    Sourcing, notes: wist.info/johnson-samuel/21120…

    #quote #quotes #quotation #qotd #samueljohnson #church #churchgoing #community #example #congregation #organizedreligion #reinforcement #religion #worship #mutualsupport

  7. 👨‍💻🤖 Oh joy, another #GitHub #repository boasting about implementing Sutton and Barto's RL methods! Because who doesn't want to navigate through yet another maze of #AI #buzzwords and autopilot #code writing? 🚀🎉
    github.com/ivanbelenky/RL #AI #Research #Reinforcement #Learning #Tech #Trends #HackerNews #ngated

  8. 👨‍💻🤖 Oh joy, another #GitHub #repository boasting about implementing Sutton and Barto's RL methods! Because who doesn't want to navigate through yet another maze of #AI #buzzwords and autopilot #code writing? 🚀🎉
    github.com/ivanbelenky/RL #AI #Research #Reinforcement #Learning #Tech #Trends #HackerNews #ngated

  9. Plinked out a new song on the #guitar last night.
    Kinda Pink Floyd-ish.. needs a little more work and then to the mixer so maybe on my next days off I'll have another #song to post.

    (this post to make me do it)

    #music
    #reinforcement

  10. Plinked out a new song on the #guitar last night.
    Kinda Pink Floyd-ish.. needs a little more work and then to the mixer so maybe on my next days off I'll have another #song to post.

    (this post to make me do it)

    #music
    #reinforcement

  11. 'Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play', by Zelai Xu, Chao Yu, Yancheng Liang, Yi Wu, Yu Wang.

    jmlr.org/papers/v26/24-1503.ht

    #reinforcement #games #play

  12. 'Learning Global Nash Equilibrium in Team Competitive Games with Generalized Fictitious Cross-Play', by Zelai Xu, Chao Yu, Yancheng Liang, Yi Wu, Yu Wang.

    jmlr.org/papers/v26/24-1503.ht

    #reinforcement #games #play

  13. #Zoomposium with Dr. #Patrick #Krauß: Building instructions for #artificial #consciousness

    Transferring the various stages of Damasio's theory of consciousness 1:1 into concrete #schematics for #deep #learning. To this end, strategies such as #feed-forward connections, #recurrent #connections in the form of #reinforcement learning and #unsupervised learning are used to simulate the #biological #processes of the #neuronal #networks.

    More at: philosophies.de/index.php/2023

    or: youtu.be/rXamzyoggCo

  14. #Zoomposium mit Dr. #Patrick #Krauß: „Bauanleitung #Künstliches #Bewusstsein

    Die verschiedenen Stufen von Damasios Theorie des Bewusstseins 1:1 in konkrete #Schaltpläne für ein #Deep #Learning zu überführen. Hierzu werden Strategien wie #feed-forward connections, #recurrent #connections in Form von #reinforcement learning und #unsupervised learning angewendet, um die #biologischen #Prozesse der #neuronalen #Netze zu simulieren.

    Mehr auf: philosophies.de/index.php/2023

    oder: youtu.be/rXamzyoggCo

  15. 'Memory Gym: Towards Endless Tasks to Benchmark Memory Capabilities of Agents', by Marco Pleines, Matthias Pallasch, Frank Zimmer, Mike Preuss.

    jmlr.org/papers/v26/24-0043.ht

    #memory #reinforcement #recurrent

  16. 'Memory Gym: Towards Endless Tasks to Benchmark Memory Capabilities of Agents', by Marco Pleines, Matthias Pallasch, Frank Zimmer, Mike Preuss.

    jmlr.org/papers/v26/24-0043.ht

    #memory #reinforcement #recurrent

  17. 'A New, Physics-Informed Continuous-Time Reinforcement Learning Algorithm with Performance Guarantees', by Brent A. Wallace, Jennie Si.

    jmlr.org/papers/v25/24-0017.ht

    #control #reinforcement #exploration

  18. 'A New, Physics-Informed Continuous-Time Reinforcement Learning Algorithm with Performance Guarantees', by Brent A. Wallace, Jennie Si.

    jmlr.org/papers/v25/24-0017.ht

    #control #reinforcement #exploration

  19. 'Learning Dynamic Mechanisms in Unknown Environments: A Reinforcement Learning Approach', by Shuang Qiu, Boxiang Lyu, Qinglin Meng, Zhaoran Wang, Zhuoran Yang, Michael I. Jordan.

    jmlr.org/papers/v25/23-0159.ht

    #reinforcement #reward #dynamic

  20. 'Learning Dynamic Mechanisms in Unknown Environments: A Reinforcement Learning Approach', by Shuang Qiu, Boxiang Lyu, Qinglin Meng, Zhaoran Wang, Zhuoran Yang, Michael I. Jordan.

    jmlr.org/papers/v25/23-0159.ht

    #reinforcement #reward #dynamic

  21. 'Learning Regularized Graphon Mean-Field Games with Unknown Graphons', by Fengzhuo Zhang, Vincent Y. F. Tan, Zhaoran Wang, Zhuoran Yang.

    jmlr.org/papers/v25/23-1409.ht

    #graphon #graphons #reinforcement

  22. 'Learning Regularized Graphon Mean-Field Games with Unknown Graphons', by Fengzhuo Zhang, Vincent Y. F. Tan, Zhaoran Wang, Zhuoran Yang.

    jmlr.org/papers/v25/23-1409.ht

    #graphon #graphons #reinforcement

  23. #ITByte: Deep Q-learning is a #Reinforcement #Learning technique that combines Q-Learning and deep neural networks. It aims to help agents learn optimal actions in complex environments.

    Here is a brief overview of Q-Learning and Deep Q-Learning.

    knowledgezone.co.in/posts/Deep

  24. #ITByte: Deep Q-learning is a #Reinforcement #Learning technique that combines Q-Learning and deep neural networks. It aims to help agents learn optimal actions in complex environments.

    Here is a brief overview of Q-Learning and Deep Q-Learning.

    knowledgezone.co.in/posts/Deep

  25. 'Sample Complexity of Variance-Reduced Distributionally Robust Q-Learning', by Shengbo Wang, Nian Si, Jose Blanchet, Zhengyuan Zhou.

    jmlr.org/papers/v25/23-0526.ht

    #robust #efficiently #reinforcement

  26. 'Sample Complexity of Variance-Reduced Distributionally Robust Q-Learning', by Shengbo Wang, Nian Si, Jose Blanchet, Zhengyuan Zhou.

    jmlr.org/papers/v25/23-0526.ht

    #robust #efficiently #reinforcement

  27. 'Empirical Design in Reinforcement Learning', by Andrew Patterson, Samuel Neumann, Martha White, Adam White.

    jmlr.org/papers/v25/23-0183.ht

    #reinforcement #experiments #hyperparameters

  28. 'Empirical Design in Reinforcement Learning', by Andrew Patterson, Samuel Neumann, Martha White, Adam White.

    jmlr.org/papers/v25/23-0183.ht

    #reinforcement #experiments #hyperparameters

  29. 'Instrumental Variable Value Iteration for Causal Offline Reinforcement Learning', by Luofeng Liao, Zuyue Fu, Zhuoran Yang, Yixin Wang, Dingli Ma, Mladen Kolar, Zhaoran Wang.

    jmlr.org/papers/v25/22-0965.ht

    #reinforcement #unobserved #causal

  30. 'Instrumental Variable Value Iteration for Causal Offline Reinforcement Learning', by Luofeng Liao, Zuyue Fu, Zhuoran Yang, Yixin Wang, Dingli Ma, Mladen Kolar, Zhaoran Wang.

    jmlr.org/papers/v25/22-0965.ht

    #reinforcement #unobserved #causal

  31. 'Value-Distributional Model-Based Reinforcement Learning', by Carlos E. Luis, Alessandro G. Bottero, Julia Vinogradska, Felix Berkenkamp, Jan Peters.

    jmlr.org/papers/v25/23-0913.ht

    #reinforcement #quantile #learns

  32. 'Value-Distributional Model-Based Reinforcement Learning', by Carlos E. Luis, Alessandro G. Bottero, Julia Vinogradska, Felix Berkenkamp, Jan Peters.

    jmlr.org/papers/v25/23-0913.ht

    #reinforcement #quantile #learns

  33. 'Pearl: A Production-Ready Reinforcement Learning Agent', by Zheqing Zhu et al.

    jmlr.org/papers/v25/24-0196.ht

    #reinforcement #rl #agent

  34. 'Pearl: A Production-Ready Reinforcement Learning Agent', by Zheqing Zhu et al.

    jmlr.org/papers/v25/24-0196.ht

    #reinforcement #rl #agent

  35. 'Mean-Field Approximation of Cooperative Constrained Multi-Agent Reinforcement Learning (CMARL)', by Washim Uddin Mondal, Vaneet Aggarwal, Satish V. Ukkusuri.

    jmlr.org/papers/v25/22-0956.ht

    #reinforcement #maximization #constraints

  36. 'Mean-Field Approximation of Cooperative Constrained Multi-Agent Reinforcement Learning (CMARL)', by Washim Uddin Mondal, Vaneet Aggarwal, Satish V. Ukkusuri.

    jmlr.org/papers/v25/22-0956.ht

    #reinforcement #maximization #constraints

  37. ✨Meet Kai Cui, a PhD student at #TUDarmstadt. His research focuses on #multi-agent #reinforcement learning, Mean Field Games and their application. At #SoftwareCampus he's collaborating with #Huawei & currently working on mean-field games for solving large-scale route planning problems in the context of #operations #research.

    "Personally, I was always fascinated by the increasing and large-scale automation or optimization processes."

    Click here, to find out more:
    softwarecampus.de/teilnehmer/k

  38. 'Distributionally Robust Model-Based Offline Reinforcement Learning with Near-Optimal Sample Complexity', by Laixi Shi, Yuejie Chi.

    jmlr.org/papers/v25/22-1482.ht

    #robustness #reinforcement #uncertainty

  39. 'Robust Black-Box Optimization for Stochastic Search and Episodic Reinforcement Learning', by Maximilian Hüttenrauch, Gerhard Neumann.

    jmlr.org/papers/v25/22-0564.ht

    #reinforcement #optimizers #optimizes

  40. 'Policy Gradient Methods in the Presence of Symmetries and State Abstractions', by Prakash Panangaden, Sahand Rezaei-Shoshtari, Rosie Zhao, David Meger, Doina Precup.

    jmlr.org/papers/v25/23-1415.ht

    #reinforcement #abstraction #abstractions

  41. 'Off-Policy Action Anticipation in Multi-Agent Reinforcement Learning', by Ariyan Bighashdel, Daan de Geus, Pavol Jancura, Gijs Dubbelman.

    jmlr.org/papers/v25/23-0413.ht

    #reinforcement #games #anticipation

  42. 'On the Sample Complexity and Metastability of Heavy-tailed Policy Search in Continuous Control', by Amrit Singh Bedi, Anjaly Parayil, Junyu Zhang, Mengdi Wang, Alec Koppel.

    jmlr.org/papers/v25/21-1343.ht

    #reinforcement #optimality #exploration

  43. 'Heterogeneous-Agent Reinforcement Learning', by Yifan Zhong, Jakub Grudzien Kuba, Xidong Feng, Siyi Hu, Jiaming Ji, Yaodong Yang.

    jmlr.org/papers/v25/23-0488.ht

    #reinforcement #agents #mirror