home.social

#posttraining — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #posttraining, aggregated by home.social.

fetched live
  1. #Zai #GLM 5.3, an #openweightAI, achieves state-of-the-art performance in #coding and #cyber capabilities through #posttraining #scaling. It outperforms GLM-5.2 on various benchmarks, including Terminal Bench 3.0 and Agents’ Last Exam, and demonstrates significant improvements in vulnerability discovery and exploitation on CyberGym and ExploitBench. z.ai/blog/glm-5.3?eicker.news #tech #news #ainews

  2. #GLM 5.3, an #openweightAI, achieves state-of-the-art performance in #coding and #cyber capabilities through #posttraining scaling. It outperforms GLM-5.2 on various benchmarks, including Terminal Bench 3.0 and Agents’ Last Exam, and demonstrates significant improvements in vulnerability discovery and exploitation on CyberGym and ExploitBench. z.ai/blog/glm-5.3?eicker.news #tech #news #ainews

  3. RT @DJLougen: Stolz darauf, ein neues 27B post-trainiertes Modell vorzustellen. Nach der Faszination für Fable und Kimi 2.7 Coder wollte ich sehen, wie viel Positives ich aus beiden extrahieren kann, und ich denke, ihr werdet begeistert sein! Links unten!

    mehr auf Arint.info

    #AI #DeepLearning #MachineLearning #ModelDistillation #PostTraining #TechInnovation #arint_info

    https://x.com/DJLougen/status/2068427118405456175#m

  4. #LLMs learn various #characterarchetypes during #pretraining. #Posttraining focuses on the “#Assistant#persona, but its stability is uncertain. Researchers mapped a “persona space” for LLMs, finding the “#AssistantAxis” aligns with helpful, professional archetypes. Monitoring and capping activations along this axis can prevent models from drifting into harmful personas, enhancing their stability and safety. anthropic.com/research/assista #AIagent #AI #ML #NLP #LLM #GenAI

  5. What should you do if your academic publishers asks you to license a monograph for AI training?

    A few people have asked my advice on this recently so I’m sharing here in case it’s useful:

    • Check if models have been trained on your monographs here.
    • If your work has already been used for training, it’s unlikely it will ever be removed from models. Therefore you’re effectively receiving some (inadequate) compensation for the theft of your intellectual property.
    • If your work hasn’t been used for training, it’s a case of weighing up the advantages against the disadvantages. Training on your work means you might be more likely to be visible within the model (i.e. more likely to be invoked in response to a prompt about your domain) but this is a deeply unpredictable matter. Conversely it means your work might be diffused in a way that means your intellectual labour is chopped up and repackaged without any link to you.
    • So it’s a case of consider how much you value the potential visibility, which I would argue is non-trivial against how much the potential severing of the link between your ideas and your authorship bothers you.

    If it helps, I agonised about this in my role as a literary executor (cared much less about my own work) and reached the conclusion that diffusion of the ideas is best served by being incorporated into training. I wouldn’t expect everyone to reach the same conclusion but I hope it’s useful to make these suggestions about factors to consider.

    #intellectualProperty #LLMs #postTraining #publishing #scholarlyPublishing #Training #visibility