#ondeviceai — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #ondeviceai, aggregated by home.social.
-
On-device AI latency often comes from token-by-token generation, not model size. Speculative decoding attacks that bottleneck directly. https://hackernoon.com/speculative-decoding-and-the-latency-ceiling-for-on-device-llms #ondeviceai
-
On-device AI latency often comes from token-by-token generation, not model size. Speculative decoding attacks that bottleneck directly. https://hackernoon.com/speculative-decoding-and-the-latency-ceiling-for-on-device-llms #ondeviceai
-
On-device AI latency often comes from token-by-token generation, not model size. Speculative decoding attacks that bottleneck directly. https://hackernoon.com/speculative-decoding-and-the-latency-ceiling-for-on-device-llms #ondeviceai
-
On-device AI latency often comes from token-by-token generation, not model size. Speculative decoding attacks that bottleneck directly. https://hackernoon.com/speculative-decoding-and-the-latency-ceiling-for-on-device-llms #ondeviceai
-
On-device AI latency often comes from token-by-token generation, not model size. Speculative decoding attacks that bottleneck directly. https://hackernoon.com/speculative-decoding-and-the-latency-ceiling-for-on-device-llms #ondeviceai