#distributedtraining — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #distributedtraining, aggregated by home.social.
-
Bringing PyTorch Monarch to AMD GPUs
Comments: https://news.ycombinator.com/item?id=49048689
#HackerNews #PyTorch #AMD #GPUs #Monarch #ROCm #DistributedTraining
-
Bringing PyTorch Monarch to AMD GPUs
Comments: https://news.ycombinator.com/item?id=49048689
#HackerNews #PyTorch #AMD #GPUs #Monarch #ROCm #DistributedTraining
-
RT @Alacritic_Super: Als KI-Infrastruktur-Ingenieur: Bitte lerne folgende Themen: GPU-Architektur, VRAM-Grundlagen, CUDA und Speicherverteilung; Quantisierung (INT8/FP8/4-bit), Batching und kontinuierliches Batching; vLLM, TensorRT-LLM, SGLang, llama.cpp und Inference-Optimierung; KV-Caching, spekulatives Decoding, Prefix-Caching und Token-Durchsatz; verteiltes Training (DDP, FSDP, DeepSpeed, ZeRO); Modellauslieferung (Triton, vLLM, KServe, Ray Serve, SGLang); Kubernetes, Docker und GPU-Orchestrierung; NCCL, InfiniBand und Hochgeschwindigkeitsnetzwerke; Multi-GPU- und Multi-Node-Inference; LoRA, QLoRA, PEFT und Fine-Tuning-Pipelines; Vektordatenbanken, Embeddings und RAG-Pipelines; Prompt-Caching, semantisches Caching und Kostenoptimierung; Observability (OpenTelemetry, Prometheus, Grafana, Langfuse, Phoenix); LLM-Evaluation, Benchmarking und A/B-Tests; Model Routing und Fallback-Strategien; MCP, KI-Agenten und Workflow-Orchestrierung; Datenpipelines, Kafka und Streaming-Inference; Sicherheit, Guardrails und Abwehr gegen Prompt-Injection; CI/CD für ML (MLOps), MLflow und Model-Registries; Linux, Netzwerk- und Speichergrundlagen; PyTorch-Internals, CUDA-Profilierung und Kernel-Optimierung.
mehr auf Arint.info
#AIEngineering #DistributedTraining #GPUComputing #KIInfrastruktur #LLMOptimierung #MLOps #arint_info
-
Decoupled DiLoCo: Resilient, Distributed AI Training at Scale
https://deepmind.google/blog/decoupled-diloco/
#HackerNews #DecoupledDiLoCo #ResilientAI #DistributedTraining #AIatScale #DeepMind
-
Decoupled DiLoCo: Resilient, Distributed AI Training at Scale
https://deepmind.google/blog/decoupled-diloco/
#HackerNews #DecoupledDiLoCo #ResilientAI #DistributedTraining #AIatScale #DeepMind
-
These Startups Are Building Advanced AI Models Without Data Centers https://www.wired.com/story/these-startups-are-building-advanced-ai-models-over-the-internet-with-untapped-data/ #AI #DistributedTraining
-
These Startups Are Building Advanced AI Models Without Data Centers https://www.wired.com/story/these-startups-are-building-advanced-ai-models-over-the-internet-with-untapped-data/ #AI #DistributedTraining
-
Import AI 409: Huawei trains a model on 8,000+ Ascend chips; 32B decentralized training run; and the era of experience and superintelligence https://importai.substack.com/p/import-ai-409-huawei-trains-a-model #AI #DistributedTraining
-
Import AI 409: Huawei trains a model on 8,000+ Ascend chips; 32B decentralized training run; and the era of experience and superintelligence https://importai.substack.com/p/import-ai-409-huawei-trains-a-model #AI #DistributedTraining
-
Import AI 404: Scaling laws for distributed training; misalignment predictions made real; and Alibaba's good translation model https://importai.substack.com/p/import-ai-404-scaling-laws-for-distributed #AI #DistributedTraining
-
Import AI 404: Scaling laws for distributed training; misalignment predictions made real; and Alibaba's good translation model https://importai.substack.com/p/import-ai-404-scaling-laws-for-distributed #AI #DistributedTraining