#rewardmodelling — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #rewardmodelling, aggregated by home.social.
-
📜 Paper: https://arxiv.org/abs/2506.15498
🤖 Models: https://huggingface.co/collections/UKPLab/spare-prm
💻 Code: https://github.com/UKPLab/aaai2026-spare-prmFollow the authors Imbesat Hassan Rizvi and Iryna Gurevych from the Ubiquitous Knowledge Processing Lab (UKP Lab), Technische Universität Darmstadt and Xiaodan Zhu from the Department of Electrical and Computer Engineering, Smith Engineering and Ingenuity Labs Research Institute at Queen's University.
#AAAI2026 #ProcessSupervision #Reasoning #RewardModelling #ReferenceGuidedEvaluation
-
📜 Paper: https://arxiv.org/abs/2506.15498
🤖 Models: https://huggingface.co/collections/UKPLab/spare-prm
💻 Code: https://github.com/UKPLab/aaai2026-spare-prmFollow the authors Imbesat Hassan Rizvi and Iryna Gurevych from the Ubiquitous Knowledge Processing Lab (UKP Lab), Technische Universität Darmstadt and Xiaodan Zhu from the Department of Electrical and Computer Engineering, Smith Engineering and Ingenuity Labs Research Institute at Queen's University.
#AAAI2026 #ProcessSupervision #Reasoning #RewardModelling #ReferenceGuidedEvaluation
-
📜 Paper: https://arxiv.org/abs/2506.15498
🤖 Models: https://huggingface.co/collections/UKPLab/spare-prm
💻 Code: https://github.com/UKPLab/aaai2026-spare-prmFollow the authors Imbesat Hassan Rizvi and Iryna Gurevych from the Ubiquitous Knowledge Processing Lab (UKP Lab), Technische Universität Darmstadt and Xiaodan Zhu from the Department of Electrical and Computer Engineering, Smith Engineering and Ingenuity Labs Research Institute at Queen's University.
#AAAI2026 #ProcessSupervision #Reasoning #RewardModelling #ReferenceGuidedEvaluation
-
📜 Paper: https://arxiv.org/abs/2506.15498
🤖 Models: https://huggingface.co/collections/UKPLab/spare-prm
💻 Code: https://github.com/UKPLab/aaai2026-spare-prmFollow the authors Imbesat Hassan Rizvi and Iryna Gurevych from the Ubiquitous Knowledge Processing Lab (UKP Lab), Technische Universität Darmstadt and Xiaodan Zhu from the Department of Electrical and Computer Engineering, Smith Engineering and Ingenuity Labs Research Institute at Queen's University.
#AAAI2026 #ProcessSupervision #Reasoning #RewardModelling #ReferenceGuidedEvaluation