home.social

#airllm — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #airllm, aggregated by home.social.

fetched live
  1. 🎉 Efficient #LLM inference: #AirLLM enables running #Llama3.1 models up to 70B on 4GB VRAM, and up to 405B on 8GB.

    💾 Memory optimization: Runs without needing quantization or distillation, saving resources.

    🚀 Compression for speed: 4-bit and 8-bit compression options provide up to 3x speed boost with minimal accuracy loss.

    🧠 Broad support: Compatible with various models like ChatGLM, QWen, and more.

    🔗 Platform-ready: Runs seamlessly on Linux, MacOS, and low-end GPUs.
    github.com/lyogavin/airllm

  2. 🎉 Efficient #LLM inference: #AirLLM enables running #Llama3.1 models up to 70B on 4GB VRAM, and up to 405B on 8GB.

    💾 Memory optimization: Runs without needing quantization or distillation, saving resources.

    🚀 Compression for speed: 4-bit and 8-bit compression options provide up to 3x speed boost with minimal accuracy loss.

    🧠 Broad support: Compatible with various models like ChatGLM, QWen, and more.

    🔗 Platform-ready: Runs seamlessly on Linux, MacOS, and low-end GPUs.
    github.com/lyogavin/airllm