#sglang — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #sglang, aggregated by home.social.
-
Инференс LLM: от KV-кэша до продакшен-деплоя
Привет! Я Саша Рыжов, MLOps-инженер в hh.ru , уже три года занимаюсь развитием инфраструктуры для искусственного интеллекта. Компании, которые развивают GenAI, рано или поздно приходят к задачам по запуску LLM на собственном железе. В статье я расскажу, как обстоят дела с движками инференса в 2026 году и как запустить on‑prem-прод и не изобрести при этом велосипед.
-
CVE-2026-14890 is an unpatched SGLang vulnerability enabling unauthenticated remote code execution through pickle deserialization. Restrict access now.
#SGLang #CVE202614890 #RCE #Pickle #AI #LLM #CERTCC #Vulnerability
https://securityonline.info/sglang-rce-cve-2026-14890/?utm_source=mastodon&utm_medium=jetpack_social
-
CVE-2026-14890 is an unpatched SGLang vulnerability enabling unauthenticated remote code execution through pickle deserialization. Restrict access now.
#SGLang #CVE202614890 #RCE #Pickle #AI #LLM #CERTCC #Vulnerability
https://securityonline.info/sglang-rce-cve-2026-14890/?utm_source=mastodon&utm_medium=jetpack_social
-
RT @h100envy: Der ehemalige Berkeley-PhD, der SGLang bei xAI leitet, erklärte, wie sie Grok auf 100.000 GPUs in 23 Minuten bedienen – besser als $2000 Inference-at-Scale-Kurse. Aufteilung von Prefill und Decode –> Sharding von Experten über GPUs –> Token-Routing pro Experte –> Überlappung von Kommunikation und Berechnung –> Bedienung zu DeepSeek-API-zerstörenden Preisen. Diese Schleife ist der Grund, warum xAI Grok auf SGLang betreibt und Dritte DeepSeek's eigene API um den Faktor 5 bei den Kosten schlagen. SGLang + Prefill-Decode-Disaggregation + Experten-Parallelität + AMD MI300 – das ist der Stack. Schau dir das Video an und speichere es, dann lies den Artikel unten. Video h100envy (@h100envy) Artikel Loop Engineering: Eine Schleife, die RAG automatisch auf ein Ziel-Recall einstellt. Ein konkreter Fall von Loop-Engineering. Nicht „RAG von Hand einstellen“, sondern eine Schleife bauen, die Konfigurationen selbst sucht, Recall auf einem Eval misst und stoppt, wenn das Ziel erreicht ist. Mit vollständigem Code. RAG-Einstellung – https://nitter.net/h100envy/status/2070852290878009586#m
mehr auf Arint.info
#DeepSeek #Grok #LoopEngineering #RAG #SGLang #xAI #arint_info
-
RT @PavloMolchanov: 🚀 Selbst-Spekulation ermöglicht eine 6,75-fache echte Beschleunigung der LLM-Generierung mit SGLang-Inference!
mehr auf Arint.info
#AI #Diffusion #LLM #MachineLearning #Nemotron #SGLang #arint_info
-
RT @PavloMolchanov: 🚀 Selbst-Spekulation ermöglicht eine 6,75-fache echte Beschleunigung der LLM-Generierung mit SGLang-Inference!
mehr auf Arint.info
#AI #Diffusion #LLM #MachineLearning #Nemotron #SGLang #arint_info
-
Архитектура AI-сервисов: почему монолит убивает latency и GPU
Ваш AI‑чат или автокомплит тормозит при 50 запросах в секунду? Монолит убивает GPU и латенси? В этом туториале — реальная архитектура low‑latency инференса на high‑load: почему изолированный inference‑bundle вместо монолита, как выбрать между vLLM и SGLang без маркетинга, зачем нужны continuous batching и admission control. Читать разбор
https://habr.com/ru/companies/otus/articles/1031286/
#AIсервисы #LLM #инференс #highload #latency #GPU #vLLM #SGLang #continuous_batching #admission_control
-
SGLang Flaw Enables Remote Code Execution via Malicious Model Files
A single malicious file can become a powerful gateway for attackers to run arbitrary commands on vulnerable machines - and a newly disclosed flaw in SGLang, CVE-2026-5760, reveals just how easily this can happen through specially crafted GGUF model files. This highly severe vulnerability, scoring 9.8 out of 10.0, enables remote code…
#RemoteCodeExecution #Cve20265760 #CommandInjection #Gguf #Sglang
-
Install SGLang with uv, pip, or Docker; configure YAML and server flags; then serve Hugging Face LLMs with an OpenAI-compatible API plus native /generate and offline Engine examples.
#Cheatsheet #Self-Hosting #LLM #AI #AI Coding #DevOps #Docker #sglang #openai #SelfHosting
-
Install SGLang with uv, pip, or Docker; configure YAML and server flags; then serve Hugging Face LLMs with an OpenAI-compatible API plus native /generate and offline Engine examples.
#Cheatsheet #Self-Hosting #LLM #AI #AI Coding #DevOps #Docker #sglang #openai #SelfHosting
-
SGLang and vLLM Workshops Coming to GOSIM Paris 2026!
The GOSIM Workshops have long been known for their diversity, hands-on learning, and interactivity, making them one of the most popular segments of the conference.
This May, the SGLang Workshop and vLLM Workshop will arrive at GOSIM Paris 2026, bringing together AI infrastructure developers from around the world to explore the latest advances in LLM inference systems.
Ticket purchase link:
https://eventbrite.com/e/gosim-paris-2026-tickets-1984013840806?aff=oddtdtcreator -
SGLang and vLLM Workshops Coming to GOSIM Paris 2026!
The GOSIM Workshops have long been known for their diversity, hands-on learning, and interactivity, making them one of the most popular segments of the conference.
This May, the SGLang Workshop and vLLM Workshop will arrive at GOSIM Paris 2026, bringing together AI infrastructure developers from around the world to explore the latest advances in LLM inference systems.
Ticket purchase link:
https://eventbrite.com/e/gosim-paris-2026-tickets-1984013840806?aff=oddtdtcreator -
🚀 Big news!
The SGLang Workshop & vLLM Workshop are coming to GOSIM Paris 2026! 🎉
🌐 A must-attend event for AI developers and open-source contributors worldwide
💡 Dive into cutting-edge topics: large model inference, agentic AI, and more
🎓 Hands-on sessions and discussions to bring high-value learning and networkingGet your early bird tickets now and enjoy the discount: https://eventbrite.com/e/gosim-paris-2026-tickets-1984013840806?aff=oddtdtcreator 🚀
-
🚀 Big news!
The SGLang Workshop & vLLM Workshop are coming to GOSIM Paris 2026! 🎉
🌐 A must-attend event for AI developers and open-source contributors worldwide
💡 Dive into cutting-edge topics: large model inference, agentic AI, and more
🎓 Hands-on sessions and discussions to bring high-value learning and networkingGet your early bird tickets now and enjoy the discount: https://eventbrite.com/e/gosim-paris-2026-tickets-1984013840806?aff=oddtdtcreator 🚀
-
Indian developers report that AI‑driven assistants let them spend less time typing and more time learning. From SGLang to Tensoic AI, interns say generative tools boost productivity and skill growth. Curious how the balance is shifting? Read the full story. #AIDrivenAssistants #GenerativeTools #SGLang #TensoicAI
🔗 https://aidailypost.com/news/indian-developers-say-ai-helps-them-learn-more-they-code-less
-
🎉 sglang-docs-l10n is published!
🚀 Preview:
https://projects.localizethedocs.org/sglang-docs-l10n
🌐 Crowdin:
https://localizethedocs.crowdin.com/sglang-docs-l10n
🐙 GitHub:
-
📦 Model weights for base and chat variants coming soon to #HuggingFace and #ModelScope with support for #vLLM and #SGLang inference frameworks
📋 Complete evaluation details and trajectory data publicly available for community research on HuggingFace datasets
-
📦 Model weights for base and chat variants coming soon to #HuggingFace and #ModelScope with support for #vLLM and #SGLang inference frameworks
📋 Complete evaluation details and trajectory data publicly available for community research on HuggingFace datasets
-
Как запустить свою LLM для инференса. Руководство по запуску: Ollama, vLLM, Triton, LM Studio, llama.cpp, SGLang
В этой статье будет приведено практическое руководство по базовой настройке и запуску следующих инструментов для работы с LLM: Ollama, LM Studio, vLLM, Triton, llama.cpp, SGLang. 🔥 Начинаем? 🔥
https://habr.com/ru/articles/948934/
#ollama #vllm #triton #lm_studio #llamacpp #sglang #запуск_llm
-
🤖 Oh joy, another thrilling journey through the riveting world of Flash Attention in SGLang! 🌟 Because clearly, the universe was desperately yearning for a detailed breakdown of yet another backend implementation. 🤯 Guess #SGLang 0.4.6 just wouldn’t be the same without it! 🥳
https://hebiao064.github.io/fa3-attn-backend-basic #FlashAttention #BackendImplementation #TechNews #Innovation #Excitement #HackerNews #ngated -
🤖 Oh joy, another thrilling journey through the riveting world of Flash Attention in SGLang! 🌟 Because clearly, the universe was desperately yearning for a detailed breakdown of yet another backend implementation. 🤯 Guess #SGLang 0.4.6 just wouldn’t be the same without it! 🥳
https://hebiao064.github.io/fa3-attn-backend-basic #FlashAttention #BackendImplementation #TechNews #Innovation #Excitement #HackerNews #ngated -
Implement Flash Attention Back End in SGLang – Basics and KV Cache
https://hebiao064.github.io/fa3-attn-backend-basic
#HackerNews #ImplementFlashAttention #SGLang #KVCache #BackEnd #AIResearch #TechTutorial
-
Implement Flash Attention Back End in SGLang – Basics and KV Cache
https://hebiao064.github.io/fa3-attn-backend-basic
#HackerNews #ImplementFlashAttention #SGLang #KVCache #BackEnd #AIResearch #TechTutorial