home.social

#finetuningllms — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #finetuningllms, aggregated by home.social.

fetched live
  1. A practical map of LLM post-training: how SFT, reward models, RL (PPO, GRPO), DPO, and RLVR fit together, and why a reward model is not RL. hackernoon.com/how-llms-are-tr #finetuningllms