#slm — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #slm, aggregated by home.social.
-
A Small language model by a friend of mine and old style hacker:
'Yachay #SLM"
https://github.com/unimauro/yachay-slmDocumentation in Spanish but code is in python and rust :-)
Created to work in regular hardware
-
Testing a tiny 60M commit-message model. Wondering if it could make @silex new git integration more useful
Based on T5, not a chatbot, a model that you pretrain once, then specialize
https://huggingface.co/SEBIS/code_trans_t5_small_commit_generation_transfer_learning_finetuneInteresting architecture :)
Plus T5 is fully open, like #openllm / @LINAGORA 's Luciole :)
It's an open academic work
Here some results on my machine:
* 0.2s/commit message on CPU
* 11s load time, 0 VRAM
* on this narrow task, the 60M is matching Luciole-1B -
Testing SmolLM2-135M on my phone
Fascinating failure mode: it understands the question perfectly and even starts with the right answer (324m), then completely hallucinates the rest 😅
At just 135M params, language understanding seems still strong!
Does someone know how low we can go before language itself breaks?
For humans I know, it's mesured in shots 🥃
-
How many parameters does it take to understand English? 🤓
This paper compresses BERT down to 1.2M parameters and still gets 98.9% intent accuracy on spoken language understanding
https://arxiv.org/pdf/1909.11687Now I want to know what happens at 500K. 100K?
-
Just when you think Zucky @zuck is finally out of ideas - to steal
It starts to feel like even a weak #LLM like #MetaAI is more creative than him.
Maybe we already have created truly artificial intelligence because he is more of an #SLM: a Small Language Model?
futurism.com/future-socie... #News
Jealously Watching OpenAI and ... -
Just when you think Zucky @zuck is finally out of ideas - to steal.
It starts to feel like even a weak #LLM like #MetaAI is more creative than him.
Maybe we already have created truly artificial intelligence because he is more of an #SLM: a Small Language Model?
https://futurism.com/future-society/jealous-meta-claims-ai-went-hacking-too
#toxic #humor #Meta #News #Facebook #Zuckerberg #MarkZuckerberg #ai #artificialintelligence #creativity #robot #sarcasm
-
Your production agents are already generating high-value training data. It's in your OTEL traces.
Learn how to turn production telemetry into datasets for fine-tuning SLMs, then safely roll them out with evaluation gates and canary releases.
🎬 Watch now: https://bit.ly/4pGpCCx
-
I'm running Qwen3.6-35B Q4_K_M on my phone!!!! (a total random android phone)
The response is streaming at 1.38 tok/s
https://github.com/Helldez/BigMoeOnEdge/releases
A dev is messing with loading and caching parts of the models to run them on smaller ram that their size
-
Neat work from my previous team. New open weight #SLM Antares optimized not for #vulnerability discovery à la #Mythos but for vuln *localization* (find instances of a known vuln in your code) which is what orgs really need.
Their 3B param model is on par with GPT-5.5 but without the cost or the cloud dependency.
-
🧪 Ever experimented with what's behind a #LLM provider's interface ?
I tried simulating #Ollama's API with #FastAPI in my IDE, on local #SLM only, over a local #RAG. 🏗️
🔮 Spoiler: on a laptop it's not just an experiment, some results are usable daily 🚀
In the article I walk through the technologies, complicating the system one benchmark at a time 😄
-
А мы и не заметили, или о том, какие ИИ живут в вашем кармане
Один разработчик запустил 400B-модель на iPhone 17 Pro. Еще раз, 400 млрд параметров, правда с оговоркой, сделал он это через стриминг весов с SSD прямо в обход оперативки. И она работала. Выдавала 0,6 токена в секунду — это примерно слово в пару секунд, с задумчивыми паузами, как диалап в 2003. Технически бесполезно, на практике возможно (но ведь запустилась же). И вот еще из неочевидного. Прямо сейчас, пока вы читаете это в Chrome на десктопе, в ваш браузер уже скачана языковая модель на несколько миллиардов параметров, и любая открытая вкладка может ее дернуть без спроса. В вашем телефоне крутится Apple Intelligence? Поздравляю, вы обладатель ИИ в вашем смартфоне. В общем, маленькие языковые модели уже среди нас. Так давайте разберемся, что это за модели на 1–7 млрд параметров, кто их делает, как они запускаются на телефоне, в браузере, и где вас ждет подвох, которого нет ни в одной таблице с бенчмарками.
https://habr.com/ru/companies/selectel/articles/1059128/
#малые_языковые_модели #SLM #локальные_модели #квантизация #Gemma #Phi4mini #Qwen #Gemini_Nano #пропускная_способность_памяти #троттлинг
-
I bet that in 2030 we'll be using claude or gpt only to setup local models with llamafile, optimized for specific tasks on a specific laptop or phone
> The only thing Internet Explorer is good for is downloading Firefox
@mozilla #llamafile is a #foss alternative to VC funded #ollama
-
How to choose between small and frontier models - The rise of small language models https://towardsdatascience.com/how-to-choose-between-small-and-frontier-models/ #AI #GenAI #LLM #SLM
-
Kennt ihr #SLM für Alt-Textgeneratoren?🔎
Oder #LLM die nicht #Azure von #MS nutzen (wie zB der von barrierefreies.design) sondern wirklich #datenschutzkonform sind? 🤖
Und auf einem nachhaltigen Rechenzentrum betrieben?💚
Freu mich über Empfehlungen!
-
OCC-RAG: компактные модели, которые отвечают только по источникам
Привет, Хабр! На связи команда Optimal Cognitive Core (OCC) из AIRI. Развитие языковых моделей в последние годы определяется масштабом: каждое новое поколение вмещает в веса всё больше знаний о мире. Но огромная доля практических задач выигрывает тогда, когда модель демонстрирует не свою энциклопедичность, а способность рассуждать и анализировать предоставленный контекст. Из этого наблюдения и выросло OCC — наше семейство компактных языковых моделей (SLM), которые имеют сильные когнитивные способности, не обладая при этом большим багажом «вызубренной» информации. В этой статье расскажем о первой модели нашего семейства, OCC‑RAG, которая оптимизирована под задачу контекстного Q&A. Мы выложили два чекпойнта, OCC‑RAG-0.6B и OCC‑RAG-1.7B (плюс ONNX‑ и GGUF‑сборки). При размере 0.6 и 1.7 млрд. параметров, соответственно, они отвечают на равных или лучше моделей общего назначения, которые в 2–6 раз больше, а по верности контексту показывают лучший результат среди моделей до 32B. Внутри — как устроена модель, как мы её обучили и что в итоге получилось.
-
Amazing talk about #AI frugality at #nocodeweek
Concepts I think are important:
Small language models #slm vs #llm
https://en.wikipedia.org/wiki/Small_language_modelPruning
https://en.wikipedia.org/wiki/Pruning_(artificial_neural_network)Quantization
https://en.wikipedia.org/wiki/Quantization_(machine_learning)EcoLogits Python library
https://pypi.org/project/ecologits/ -
I submitted a conference proposal to the #OpenSourceSummit to talk about fine tuned SLM for a specific task, full open source models (hey @LINAGORA and #LLMFrance), and @silex for your web design vibe coding
https://events.linuxfoundation.org/open-source-summit-europe/
What do you think of the title?
> Local SLMs + No-Code to build websites: where we are, and how to contribute
-
Microsoft、新しいオンデバイスモデル「Aion 1.0 Instruct」「Aion 1.0 Plan」を発表/「Instruct」はCPUだけでも動作。「Plan」は完全なローカルエージェント機能を備える
https://forest.watch.impress.co.jp/docs/news/2113916.html#forest_watch_impress #Microsoft_Edge #SLM #Phi_4_mini #Aion_1_0_Instruct #Aion_1_0_Plan #システム_ファイル #システム #Windows
-
Die neue Podcast-Staffel von "Medien - aber richtig!" ist erschienen 🎉 .
In drei Folgen spreche ich mit MDR-Journalist Thomas Lopau über das Buzzword "digitale Souveränität".
1️⃣ In Folge 1 geht es um den Begriff selbst und warum uns das Thema alle betrifft.
2️⃣ In Folge 2 schauen wir, wo wir aktuell eigentlich stehen und was die Rolle der großen Tech-Konzerne dabei ist.
3️⃣ In Folge 3 geht es dann um Lösungen: wo sind die großen und kleinen Hebel, was kann jeder tun?Aufraggeber ist die #VHS Sächsische Schweiz-Osterzgebirge, gefördert wird das Projekt von der Sächsischen Landesmedienanstalt (#SLM).
Hört gerne rein! 📻
https://www.vhs-ssoe.de/medien-aber-richtig/podcast
#Podcast #Medienbildung #digitaleSouveränität #UnplugBigTech #DiDay #Sachsen #Medienkompetenz #SaechsischeSchweiz #Pirna
-
30 лет мы внедряли в России Ansys. А потом он ушёл — и пришлось садиться писать собственный CAE для аддитивной печати
Если коротко - речь про софт, который моделирует, что произойдёт с титановой лопаткой, пока её печатает SLM-принтер. И почему до 2022 года эту задачу в авиастроении в России решали в Ansys Additive (реже - в Simufact Additive), а теперь приходится решать чем-то ещё. В этом «чем-то ещё» мы копаемся последние несколько лет. Ниже - про текущее состояние: что работает, что мы пока не умеем, и почему 3D-печать металла в авиации - это, мягко говоря, не «нажал кнопку - получил деталь».
https://habr.com/ru/articles/1039514/
#SIMMAXADDITIVE #CAE #SLM #аддитивные_технологии #Ansys #Импортозамещение #авиастроение #инженерный_анализ #российское_ПО #метод_конечных_элементов
-
@sjvn @ZDNet Lets not forget #Microsoft also had its own #UNIX distribution called #XENIX which was used with their internal version control system #SLM as the repository for their proprietary platforms source code such as #Office and #Windows. They also had one of the largest SUN Solaris installations for handling internal mail before Exchange also for running #Hotmail.