home.social

#cxl — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #cxl, aggregated by home.social.

  1. Meta measures 43.7% of its servers as memory-capacity bound. They run out of RAM before they run out of cores, network, or storage.

    CXL is the standard meant to fix that, and it is further along than most coverage suggests in one direction and further behind in another. Expansion runs in production today. Pooling works on real hardware.

    #CXL #ComputeExpressLink #MemoryPooling #ServerMemory #DataCenter #Interconnect #DRAM #DDR4 #DDR5 #AIInfrastructure #tech

    buysellram.com/blog/cxl-memory

  2. Dentro la gerarchia della memoria delle GPU: come i server AI spostano i dati dagli SSD alla HBM

    Ti sei mai chiesto come i modelli di intelligenza artificiale trasferiscono i dati alla GPU? Questo articolo spiega in modo semplice i diversi livelli di memoria all'interno di un server AI, perché la memoria della GPU è diventata un collo di bottiglia e quali nuove tecnologie potrebbero migliorare le prestazioni dell'intelligenza artificiale.

    buysellram.com/blog/inside-the…

    #HBM #HBM4 #GPUMemory #AIInfrastructure #DataCenter #CXL #HighBandwidthFlash #AIHardware #MemoryHierarchy #AIInference #NVMe #ITAD

  3. Why can't a GPU just carry more HBM? Interposers max out in size, stacking 12 to 16 DRAM dies compounds yield losses, and every stack sits beside a kilowatt-class package that hates sharing heat. So capacity climbs in careful steps — 80 GB on the H100, 141 on the H200, 192 on the B200, 288 on Blackwell Ultra — while KV caches for long-context inference balloon past 40 GB per request.

    That gap between what models demand and what packaging permits is reshaping server design. NVIDIA's Rubin platform treats CPU memory and HBM as one coherent pool. SanDisk and SK hynix are standardizing High Bandwidth Flash as a capacity tier under HBM, with first samples due this half. CXL 4.0 pooling hardware is landing in racks now.

    This article walks the whole memory hierarchy, from on-chip SRAM to NVMe, and explains what each emerging technology actually solves — and what it doesn't.

    buysellram.com/blog/inside-the

    #HBM #HBM4 #GPUMemory #AIInfrastructure #DataCenter #CXL #HighBandwidthFlash #AIHardware #MemoryHierarchy #AIInference #NVMe #ITAD #technology

  4. A modern AI server runs five layers of memory, from nanosecond on-chip SRAM to petabyte-scale SSDs, and keeping the GPU fed is the whole engineering game. This piece walks the full hierarchy: why HBM became the bottleneck, why adding more isn't simple, and where HBF, CXL, and PIM fit into the next generation.

    buysellram.com/blog/inside-the

    #HBM #HBM4 #GPUMemory #AIInfrastructure #DataCenter #CXL #HighBandwidthFlash #AIHardware #MemoryHierarchy #AIInference #NVMe #tech

  5. Ein weiterer Interconnect zur Anbindung von Speicher und Beschleunigern an CPUs wirft das Handtuch. IBMs OpenCAPI ist praktisch Geschichte.
    Neuer Branchenstandard: Intels CXL-Interconnect schluckt OpenCAPI
  6. heise+ | PCIe und CXL: Wie schnelle Schnittstellen die Server-Architektur verändern

    Compute Express Link und PCIe 5.0 und 6.0 legen das Fundament für neue Serverkonzepte. Die Hardware solcher Maschinen lässt sich per Software zusammenschalten.
    PCIe und CXL: Wie schnelle Schnittstellen die Server-Architektur verändern
  7. Durch hohe Transferraten, Switching, Cache-Kohärenz und weitere Funktionen ermöglichen PCIe 6.0 und CXL 2.0 neue Architekturen.
    PCI Express 6.0 und CXL 2.0 sollen Server umkrempeln
  8. heise+ | Compute Express Link: Der Interconnect erklärt

    Memory- und Rechenressourcen aller Couleur will der neue Cache-kohärente Highspeed-Interconnect CXL quer über alle Server eines Racks verteilen.
    Compute Express Link: Der Interconnect erklärt