home.social

#modelcompression — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #modelcompression, aggregated by home.social.

fetched live
  1. Authors: Federico Marcuzzi (INSAIT - Institute for Computer Science, Artificial Intelligence and Technology), Xuefei Ning (Tsinghua University), Roy Schwartz (The Hebrew University of Jerusalem), and Iryna Gurevych (UKP Lab, Technische Universität Darmstadt and ATHENE Center).

    See you at #EACL2026 in Rabat 🕌!

    #UKPLab #NLProc #ResponsibleAI #Quantization #MLSafety #Fairness #TrustworthyAI #ModelCompression #LLMSafety #EthicalAI #NLP #AIResearch

  2. New research shows KV‑cache compaction can slash LLM memory usage by up to 50× while preserving quality. With chunked processing and attention‑matching tricks, models like Llama 3.1 and Qwen‑3 handle far longer contexts—great news for open‑source and enterprise workloads. Dive into the benchmarks! #KVCaching #LLMMemory #LongContexts #ModelCompression

    🔗 aidailypost.com/news/kv-cache-

  3. Here is what I've been reading this week (btw, if the authors are on Mastodon, please let me know their handles). It mostly deals with #modelcompression and #gpu programming, two problems that have become very interesting to me recently.