home.social

#llmd — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #llmd, aggregated by home.social.

fetched live
  1. Sticky Until Saturated: Token-Aware Routing in llm-d #llmd #ai twp.ai/E5EEOj

  2. Sticky Until Saturated: Token-Aware Routing in llm-d #llmd #ai twp.ai/E5EEOj

  3. Sticky Until Saturated: Token-Aware Routing in llm-d #llmd #ai twp.ai/E5EEOj

  4. Sticky Until Saturated: Token-Aware Routing in llm-d #llmd #ai twp.ai/E5EEOj

  5. Sticky Until Saturated: Token-Aware Routing in llm-d #llmd #ai twp.ai/E5EEOj

  6. Scaling Vision-Heavy Kimi-VL with Heterogeneous E/PD on llm-d and SGLang #llmd #ai twp.ai/E5EtQE

  7. Scaling Vision-Heavy Kimi-VL with Heterogeneous E/PD on llm-d and SGLang #llmd #ai twp.ai/E5EtQE

  8. Scaling Vision-Heavy Kimi-VL with Heterogeneous E/PD on llm-d and SGLang #llmd #ai twp.ai/E5EtQE

  9. Scaling Vision-Heavy Kimi-VL with Heterogeneous E/PD on llm-d and SGLang #llmd #ai twp.ai/E5EtQE

  10. Scaling Vision-Heavy Kimi-VL with Heterogeneous E/PD on llm-d and SGLang #llmd #ai twp.ai/E5EtQE

  11. llm-d 0.4: Achieve SOTA Performance Across Accelerators #llmd #ai twp.ai/E5EnOb

  12. llm-d 0.4: Achieve SOTA Performance Across Accelerators #llmd #ai twp.ai/E5EnOb

  13. llm-d 0.4: Achieve SOTA Performance Across Accelerators #llmd #ai twp.ai/E5EnOb

  14. llm-d 0.4: Achieve SOTA Performance Across Accelerators #llmd #ai twp.ai/E5EnOb

  15. llm-d 0.4: Achieve SOTA Performance Across Accelerators #llmd #ai twp.ai/E5EnOb

  16. KV-Cache Wins You Can See: From Prefix Caching in vLLM to Distributed Scheduling with llm-d #llmd #ai twp.ai/E5EnOZ

  17. KV-Cache Wins You Can See: From Prefix Caching in vLLM to Distributed Scheduling with llm-d #llmd #ai twp.ai/E5EnOZ

  18. KV-Cache Wins You Can See: From Prefix Caching in vLLM to Distributed Scheduling with llm-d #llmd #ai twp.ai/E5EnOZ

  19. KV-Cache Wins You Can See: From Prefix Caching in vLLM to Distributed Scheduling with llm-d #llmd #ai twp.ai/E5EnOZ

  20. KV-Cache Wins You Can See: From Prefix Caching in vLLM to Distributed Scheduling with llm-d #llmd #ai twp.ai/E5EnOZ

  21. llm-d 0.2: Our first well-lit paths (mind the tree roots!) #llmd #ai twp.ai/E5EnOf

  22. llm-d 0.2: Our first well-lit paths (mind the tree roots!) #llmd #ai twp.ai/E5EnOf

  23. llm-d 0.2: Our first well-lit paths (mind the tree roots!) #llmd #ai twp.ai/E5EnOf

  24. llm-d 0.2: Our first well-lit paths (mind the tree roots!) #llmd #ai twp.ai/E5EnOf

  25. llm-d 0.2: Our first well-lit paths (mind the tree roots!) #llmd #ai twp.ai/E5EnOf

  26. llm-d 0.3: Wider Well-Lit Paths for Scalable Inference #llmd #ai twp.ai/E5EnOa

  27. llm-d 0.3: Wider Well-Lit Paths for Scalable Inference #llmd #ai twp.ai/E5EnOa

  28. llm-d 0.3: Wider Well-Lit Paths for Scalable Inference #llmd #ai twp.ai/E5EnOa

  29. llm-d 0.3: Wider Well-Lit Paths for Scalable Inference #llmd #ai twp.ai/E5EnOa

  30. llm-d 0.3: Wider Well-Lit Paths for Scalable Inference #llmd #ai twp.ai/E5EnOa