#airllm — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #airllm, aggregated by home.social.
-
AirLLM 70B inference with single 4GB GPU
https://github.com/lyogavin/airllm
Comments: https://news.ycombinator.com/item?id=49154228
#HackerNews #AirLLM #70B #GPU #AI #inference #4GB #deep #learning
-
AirLLM 70B inference with single 4GB GPU
https://github.com/lyogavin/airllm
Comments: https://news.ycombinator.com/item?id=49154228
#HackerNews #AirLLM #70B #GPU #AI #inference #4GB #deep #learning
-
🎉 Efficient #LLM inference: #AirLLM enables running #Llama3.1 models up to 70B on 4GB VRAM, and up to 405B on 8GB.
💾 Memory optimization: Runs without needing quantization or distillation, saving resources.
🚀 Compression for speed: 4-bit and 8-bit compression options provide up to 3x speed boost with minimal accuracy loss.
🧠 Broad support: Compatible with various models like ChatGLM, QWen, and more.
🔗 Platform-ready: Runs seamlessly on Linux, MacOS, and low-end GPUs.
https://github.com/lyogavin/airllm -
🎉 Efficient #LLM inference: #AirLLM enables running #Llama3.1 models up to 70B on 4GB VRAM, and up to 405B on 8GB.
💾 Memory optimization: Runs without needing quantization or distillation, saving resources.
🚀 Compression for speed: 4-bit and 8-bit compression options provide up to 3x speed boost with minimal accuracy loss.
🧠 Broad support: Compatible with various models like ChatGLM, QWen, and more.
🔗 Platform-ready: Runs seamlessly on Linux, MacOS, and low-end GPUs.
https://github.com/lyogavin/airllm