home.social

#ondeviceai — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #ondeviceai, aggregated by home.social.

  1. On-device AI latency often comes from token-by-token generation, not model size. Speculative decoding attacks that bottleneck directly. hackernoon.com/speculative-dec #ondeviceai

  2. On-device AI latency often comes from token-by-token generation, not model size. Speculative decoding attacks that bottleneck directly. hackernoon.com/speculative-dec #ondeviceai

  3. On-device AI latency often comes from token-by-token generation, not model size. Speculative decoding attacks that bottleneck directly. hackernoon.com/speculative-dec #ondeviceai

  4. On-device AI latency often comes from token-by-token generation, not model size. Speculative decoding attacks that bottleneck directly. hackernoon.com/speculative-dec

  5. On-device AI latency often comes from token-by-token generation, not model size. Speculative decoding attacks that bottleneck directly. hackernoon.com/speculative-dec #ondeviceai