home.social

#slm — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #slm, aggregated by home.social.

fetched live
  1. A Small language model by a friend of mine and old style hacker:
    'Yachay #SLM"
    github.com/unimauro/yachay-slm

    Documentation in Spanish but code is in python and rust :-)

    Created to work in regular hardware

    #AI #Opensource #Python #rustlang

  2. Testing a tiny 60M commit-message model. Wondering if it could make @silex new git integration more useful

    Based on T5, not a chatbot, a model that you pretrain once, then specialize
    huggingface.co/SEBIS/code_tran

    Interesting architecture :)

    Plus T5 is fully open, like #openllm / @LINAGORA 's Luciole :)

    It's an open academic work

    Here some results on my machine:
    * 0.2s/commit message on CPU
    * 11s load time, 0 VRAM
    * on this narrow task, the 60M is matching Luciole-1B

    #SLM #LocalAI #OpenSource #AI

  3. Testing SmolLM2-135M on my phone

    Fascinating failure mode: it understands the question perfectly and even starts with the right answer (324m), then completely hallucinates the rest 😅

    At just 135M params, language understanding seems still strong!

    Does someone know how low we can go before language itself breaks?

    For humans I know, it's mesured in shots 🥃

    huggingface.co/HuggingFaceTB/S

    #AI #LLM #LocalAI #MachineLearning #SLM

  4. How many parameters does it take to understand English? 🤓

    This paper compresses BERT down to 1.2M parameters and still gets 98.9% intent accuracy on spoken language understanding
    arxiv.org/pdf/1909.11687

    Now I want to know what happens at 500K. 100K?

    #llm #slm #ia #genai

  5. Just when you think Zucky @zuck is finally out of ideas - to steal

    It starts to feel like even a weak #LLM like #MetaAI is more creative than him.

    Maybe we already have created truly artificial intelligence because he is more of an #SLM: a Small Language Model?

    futurism.com/future-socie... #News

    Jealously Watching OpenAI and ...

  6. Just when you think Zucky @zuck is finally out of ideas - to steal.

    It starts to feel like even a weak #LLM like #MetaAI is more creative than him.

    Maybe we already have created truly artificial intelligence because he is more of an #SLM: a Small Language Model?

    futurism.com/future-society/je

    #toxic #humor #Meta #News #Facebook #Zuckerberg #MarkZuckerberg #ai #artificialintelligence #creativity #robot #sarcasm

  7. Your production agents are already generating high-value training data. It's in your OTEL traces.

    Learn how to turn production telemetry into datasets for fine-tuning SLMs, then safely roll them out with evaluation gates and canary releases.

    🎬 Watch now: bit.ly/4pGpCCx

    #AI #OTEL #SLM #InfoQ

  8. I'm running Qwen3.6-35B Q4_K_M on my phone!!!! (a total random android phone)

    The response is streaming at 1.38 tok/s

    github.com/Helldez/BigMoeOnEdg

    A dev is messing with loading and caching parts of the models to run them on smaller ram that their size

    #SLM #LLM #ia #OpenSource

  9. Neat work from my previous team. New open weight #SLM Antares optimized not for #vulnerability discovery à la #Mythos but for vuln *localization* (find instances of a known vuln in your code) which is what orgs really need.

    Their 3B param model is on par with GPT-5.5 but without the cost or the cloud dependency.

    blogs.cisco.com/ai/introducing

  10. 🧪 Ever experimented with what's behind a #LLM provider's interface ?

    I tried simulating #Ollama's API with #FastAPI in my IDE, on local #SLM only, over a local #RAG. 🏗️

    🔮 Spoiler: on a laptop it's not just an experiment, some results are usable daily 🚀

    In the article I walk through the technologies, complicating the system one benchmark at a time 😄

    alessandra.bilardi.net/diary/a

    #DiaryOfALazyDeveloper

  11. А мы и не заметили, или о том, какие ИИ живут в вашем кармане

    Один разработчик запустил 400B-модель на iPhone 17 Pro. Еще раз, 400 млрд параметров, правда с оговоркой, сделал он это через стриминг весов с SSD прямо в обход оперативки. И она работала. Выдавала 0,6 токена в секунду — это примерно слово в пару секунд, с задумчивыми паузами, как диалап в 2003. Технически бесполезно, на практике возможно (но ведь запустилась же). И вот еще из неочевидного. Прямо сейчас, пока вы читаете это в Chrome на десктопе, в ваш браузер уже скачана языковая модель на несколько миллиардов параметров, и любая открытая вкладка может ее дернуть без спроса. В вашем телефоне крутится Apple Intelligence? Поздравляю, вы обладатель ИИ в вашем смартфоне. В общем, маленькие языковые модели уже среди нас. Так давайте разберемся, что это за модели на 1–7 млрд параметров, кто их делает, как они запускаются на телефоне, в браузере, и где вас ждет подвох, которого нет ни в одной таблице с бенчмарками.

    habr.com/ru/companies/selectel

    #малые_языковые_модели #SLM #локальные_модели #квантизация #Gemma #Phi4mini #Qwen #Gemini_Nano #пропускная_способность_памяти #троттлинг

  12. I bet that in 2030 we'll be using claude or gpt only to setup local models with llamafile, optimized for specific tasks on a specific laptop or phone

    > The only thing Internet Explorer is good for is downloading Firefox

    @mozilla #llamafile is a #foss alternative to VC funded #ollama

    #ai #llm #slm #fantasticfour #enshitification #OpenSource

  13. Kennt ihr #SLM für Alt-Textgeneratoren?🔎

    Oder #LLM die nicht #Azure von #MS nutzen (wie zB der von barrierefreies.design) sondern wirklich #datenschutzkonform sind? 🤖

    Und auf einem nachhaltigen Rechenzentrum betrieben?💚

    Freu mich über Empfehlungen!

    #AI #KI #Alttext #Barrierefreiheit

  14. OCC-RAG: компактные модели, которые отвечают только по источникам

    Привет, Хабр! На связи команда Optimal Cognitive Core (OCC) из AIRI. Развитие языковых моделей в последние годы определяется масштабом: каждое новое поколение вмещает в веса всё больше знаний о мире. Но огромная доля практических задач выигрывает тогда, когда модель демонстрирует не свою энциклопедичность, а способность рассуждать и анализировать предоставленный контекст. Из этого наблюдения и выросло OCC — наше семейство компактных языковых моделей (SLM), которые имеют сильные когнитивные способности, не обладая при этом большим багажом «вызубренной» информации. В этой статье расскажем о первой модели нашего семейства, OCC‑RAG, которая оптимизирована под задачу контекстного Q&A. Мы выложили два чекпойнта, OCC‑RAG-0.6B и OCC‑RAG-1.7B (плюс ONNX‑ и GGUF‑сборки). При размере 0.6 и 1.7 млрд. параметров, соответственно, они отвечают на равных или лучше моделей общего назначения, которые в 2–6 раз больше, а по верности контексту показывают лучший результат среди моделей до 32B. Внутри — как устроена модель, как мы её обучили и что в итоге получилось.

    habr.com/ru/companies/airi/art

    #ai #small_language_model #slm #Context_QA

  15. I submitted a conference proposal to the #OpenSourceSummit to talk about fine tuned SLM for a specific task, full open source models (hey @LINAGORA and #LLMFrance), and @silex for your web design vibe coding

    events.linuxfoundation.org/ope

    What do you think of the title?

    > Local SLMs + No-Code to build websites: where we are, and how to contribute

    #foss #openSource #nocode #vibeCoding #llm #slm

  16. Microsoft、新しいオンデバイスモデル「Aion 1.0 Instruct」「Aion 1.0 Plan」を発表/「Instruct」はCPUだけでも動作。「Plan」は完全なローカルエージェント機能を備える
    forest.watch.impress.co.jp/doc

    #forest_watch_impress #Microsoft_Edge #SLM #Phi_4_mini #Aion_1_0_Instruct #Aion_1_0_Plan #システム_ファイル #システム #Windows

  17. Die neue Podcast-Staffel von "Medien - aber richtig!" ist erschienen 🎉 .

    In drei Folgen spreche ich mit MDR-Journalist Thomas Lopau über das Buzzword "digitale Souveränität".

    1️⃣ In Folge 1 geht es um den Begriff selbst und warum uns das Thema alle betrifft.
    2️⃣ In Folge 2 schauen wir, wo wir aktuell eigentlich stehen und was die Rolle der großen Tech-Konzerne dabei ist.
    3️⃣ In Folge 3 geht es dann um Lösungen: wo sind die großen und kleinen Hebel, was kann jeder tun?

    Aufraggeber ist die #VHS Sächsische Schweiz-Osterzgebirge, gefördert wird das Projekt von der Sächsischen Landesmedienanstalt (#SLM).

    Hört gerne rein! 📻

    vhs-ssoe.de/medien-aber-richti

    #Podcast #Medienbildung #digitaleSouveränität #UnplugBigTech #DiDay #Sachsen #Medienkompetenz #SaechsischeSchweiz #Pirna

  18. 30 лет мы внедряли в России Ansys. А потом он ушёл — и пришлось садиться писать собственный CAE для аддитивной печати

    Если коротко - речь про софт, который моделирует, что произойдёт с титановой лопаткой, пока её печатает SLM-принтер. И почему до 2022 года эту задачу в авиастроении в России решали в Ansys Additive (реже - в Simufact Additive), а теперь приходится решать чем-то ещё. В этом «чем-то ещё» мы копаемся последние несколько лет. Ниже - про текущее состояние: что работает, что мы пока не умеем, и почему 3D-печать металла в авиации - это, мягко говоря, не «нажал кнопку - получил деталь».

    habr.com/ru/articles/1039514/

    #SIMMAXADDITIVE #CAE #SLM #аддитивные_технологии #Ansys #Импортозамещение #авиастроение #инженерный_анализ #российское_ПО #метод_конечных_элементов

  19. @sjvn @ZDNet Lets not forget #Microsoft also had its own #UNIX distribution called #XENIX which was used with their internal version control system #SLM as the repository for their proprietary platforms source code such as #Office and #Windows. They also had one of the largest SUN Solaris installations for handling internal mail before Exchange also for running #Hotmail.