#mi300x — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #mi300x, aggregated by home.social.
-
🚀 Behold: a stunning display of #techno-babble that tries to convince you that flashing an #AMD #MI300X with #DeepSeek V4 is the pinnacle of human achievement. 🌟 Spoiler alert: it's yet another #GitHub page promising to revolutionize your life with #code #soup, while you just wanted to know what 'DeepSeek' even means. 😂
https://github.com/ryanzhou/deepseek-v4-flash-mi300x #HackerNews #ngated -
🚀 Behold: a stunning display of #techno-babble that tries to convince you that flashing an #AMD #MI300X with #DeepSeek V4 is the pinnacle of human achievement. 🌟 Spoiler alert: it's yet another #GitHub page promising to revolutionize your life with #code #soup, while you just wanted to know what 'DeepSeek' even means. 😂
https://github.com/ryanzhou/deepseek-v4-flash-mi300x #HackerNews #ngated -
DeepSeek V4 Flash on a Single AMD MI300X
https://github.com/ryanzhou/deepseek-v4-flash-mi300x
Comments: https://news.ycombinator.com/item?id=49166386
#HackerNews #DeepSeek #V4 #Flash #AMD #MI300X #AI #Technology #MachineLearning #GPU
-
DeepSeek V4 Flash on a Single AMD MI300X
https://github.com/ryanzhou/deepseek-v4-flash-mi300x
Comments: https://news.ycombinator.com/item?id=49166386
#HackerNews #DeepSeek #V4 #Flash #AMD #MI300X #AI #Technology #MachineLearning #GPU
-
#AMD threatens to go medieval on Nvidia with #Epyc and #Instinct: What we know so far
AMD teased its next-generation #AI accelerators at #CES2026, with CEO Lisa Su boasting the #MI500-series will deliver a 1,000x uplift in performance over its two-year-old #MI300X #GPU.
if AMD wants to stay competitive with Nvidia, MI500-series will need to deliver performance on par with if not better than Rubin Ultra Kyber racks.
AMD joins the #rackscale race with #MI455X #Helios racks.
https://www.theregister.com/2026/01/07/mi500x_amd_ai/ -
#AMD threatens to go medieval on Nvidia with #Epyc and #Instinct: What we know so far
AMD teased its next-generation #AI accelerators at #CES2026, with CEO Lisa Su boasting the #MI500-series will deliver a 1,000x uplift in performance over its two-year-old #MI300X #GPU.
if AMD wants to stay competitive with Nvidia, MI500-series will need to deliver performance on par with if not better than Rubin Ultra Kyber racks.
AMD joins the #rackscale race with #MI455X #Helios racks.
https://www.theregister.com/2026/01/07/mi500x_amd_ai/ -
AMD and OpenAI team up for massive 6 gigawatt GPU partnership
-
Battle of the giants: Nvidia #Blackwell B200 takes the lead in FluidX3D CFD performance
#Nvidia #B200 just launched, and I'm one of the first people to benchmark 8x B200 via Shadeform, in a WhiteFiber server with 2x #Intel #Xeon6 6960P 72-core CPUs. 🖖😋
8x Nvidia B200 go head-to-head with 8x #AMD #MI300X in the #FluidX3D #CFD benchmark, winning overall (with FP16S storage) at 219300 MLUPs/s (~17TB/s combined VRAM bandwidth), but losing in FP32 & FP16C storage. 8x MI300X achieve 204924 MLUPs/s.
-
Battle of the giants: Nvidia #Blackwell B200 takes the lead in FluidX3D CFD performance
#Nvidia #B200 just launched, and I'm one of the first people to benchmark 8x B200 via Shadeform, in a WhiteFiber server with 2x #Intel #Xeon6 6960P 72-core CPUs. 🖖😋
8x Nvidia B200 go head-to-head with 8x #AMD #MI300X in the #FluidX3D #CFD benchmark, winning overall (with FP16S storage) at 219300 MLUPs/s (~17TB/s combined VRAM bandwidth), but losing in FP32 & FP16C storage. 8x MI300X achieve 204924 MLUPs/s.
-
Instella: New Open 3B Language Models
https://rocm.blogs.amd.com/artificial-intelligence/introducing-instella-3B/README.html
#ycombinator #Instella #MI300X #LLMs #ROCm -
Instella: New Open 3B Language Models
https://rocm.blogs.amd.com/artificial-intelligence/introducing-instella-3B/README.html
#ycombinator #Instella #MI300X #LLMs #ROCm -
#AMD Announces "#Instella" Fully #OpenSource 3B Language Models
AMD Instella represents "fully open state-of-the-art 3-billion-parameter language models (LMs)." These models were trained on AMD Instinct #MI300X #GPU and according to AMD's published data delivers competitive performance to the likes of Llama 3.2 3B, Gemma-2 2B, and Qwen 2.5 3B.
https://www.phoronix.com/news/AMD-Intella-Open-Source-LM -
#AMD Announces "#Instella" Fully #OpenSource 3B Language Models
AMD Instella represents "fully open state-of-the-art 3-billion-parameter language models (LMs)." These models were trained on AMD Instinct #MI300X #GPU and according to AMD's published data delivers competitive performance to the likes of Llama 3.2 3B, Gemma-2 2B, and Qwen 2.5 3B.
https://www.phoronix.com/news/AMD-Intella-Open-Source-LM -
Hot Aisle's 8x AMD #MI300X server is the fastest computer I've ever tested in #FluidX3D #CFD, achieving a peak #LBM performance of 205 GLUPs/s, and a combined VRAM bandwidth of 23 TB/s. 🖖🤯
The #RTX 5090 looks like a toy in comparison.MI300X beats even Nvidia's GH200 94GB. This marks a very fascinating inflection point in #GPGPU: #CUDA is not the performance leader anymore. 🖖😛
You need a cross-vendor language like #OpenCL to leverage its power.FluidX3D on #GitHub: https://github.com/ProjectPhysX/FluidX3D
-
Hot Aisle's 8x AMD #MI300X server is the fastest computer I've ever tested in #FluidX3D #CFD, achieving a peak #LBM performance of 205 GLUPs/s, and a combined VRAM bandwidth of 23 TB/s. 🖖🤯
The #RTX 5090 looks like a toy in comparison.MI300X beats even Nvidia's GH200 94GB. This marks a very fascinating inflection point in #GPGPU: #CUDA is not the performance leader anymore. 🖖😛
You need a cross-vendor language like #OpenCL to leverage its power.FluidX3D on #GitHub: https://github.com/ProjectPhysX/FluidX3D
-
Sizing up #MI300A’s #GPU
It’s well ahead of #Nvidia’s #H100 PCIe for just about every major category of 32- or 64-bit operations. MI300A can achieve 113.2 TFLOPS of #FP32 throughput, with each FMA counting as two floating point operations. For comparison, H100 PCIe achieved 49.3 TFLOPS in same test.
#AMD cut down #MI300X’s GPU to create MI300A. 24 #Zen4 cores is a lot of #CPU power, and occupies one quadrant on the MI300 chip. But MI300’s main attraction is still the GPU.
https://chipsandcheese.com/p/sizing-up-mi300as-gpu -
Sizing up #MI300A’s #GPU
It’s well ahead of #Nvidia’s #H100 PCIe for just about every major category of 32- or 64-bit operations. MI300A can achieve 113.2 TFLOPS of #FP32 throughput, with each FMA counting as two floating point operations. For comparison, H100 PCIe achieved 49.3 TFLOPS in same test.
#AMD cut down #MI300X’s GPU to create MI300A. 24 #Zen4 cores is a lot of #CPU power, and occupies one quadrant on the MI300 chip. But MI300’s main attraction is still the GPU.
https://chipsandcheese.com/p/sizing-up-mi300as-gpu -
MI300X vs H100 vs H200 Benchmark Part 1: Training – CUDA Moat Still Alive – SemiAnalysis
Link
📌 Summary: 本文深入比較了AMD的MI300X與Nvidia的H100和H200在訓練性能、用戶體驗和總擁有成本等方面的優劣。儘管MI300X在規格上似乎優於競爭對手,實際性能卻未達預期,主要原因是AMD的公共軟件堆棧存在多重漏洞,導致用戶的初始體驗不佳。AMD必須改進軟件質量和測試過程,並提供更良好的出廠體驗纔能有效競爭。此文亦提供了對AMD的具體建議,助其在AI訓練工作負載中成為更強的競爭者。
🎯 Key Points:
- 性能比較:MI300X在矩陣乘法(GEMM)性能上普遍低於H100/H200。
- 用戶體驗:MI300X的公共穩定版本在出廠時存在多數Bug,影響用戶的使用體驗。
- 總擁有成本(TCO):雖然MI300X的TCO較低,但在公開穩定版本上,其訓練性能表現卻不佳。
- 建議改進:AMD需提高軟件開發資源、改善自家開發流程,並加強自動化測試來提升產品質量。
- 軟件支持:AMD應提交MLPerf訓練結果,以提升其市場競爭力和透明度。
🔖 Keywords: #MI300X #H100 #H200 #訓練性能 #用戶體驗 -
First #AI #Benchmarks Pitting #AMD Against #Nvidia
Results are good in that they show #MI300X is absolutely competitive with H100 #GPU on one set of AI inference benchmarks, and based on our estimates of GPU and total system costs can be competitive with Nvidia’s H100 and #H200. But, tests only done for #Llama2 #LLM model from Meta with 70 billion parameters.
A lot will depend, on how AMD prices #MI325 later this year and how many AMD can get its partners to manufacture.
https://www.nextplatform.com/2024/09/03/the-first-ai-benchmarks-pitting-amd-against-nvidia/ -
First #AI #Benchmarks Pitting #AMD Against #Nvidia
Results are good in that they show #MI300X is absolutely competitive with H100 #GPU on one set of AI inference benchmarks, and based on our estimates of GPU and total system costs can be competitive with Nvidia’s H100 and #H200. But, tests only done for #Llama2 #LLM model from Meta with 70 billion parameters.
A lot will depend, on how AMD prices #MI325 later this year and how many AMD can get its partners to manufacture.
https://www.nextplatform.com/2024/09/03/the-first-ai-benchmarks-pitting-amd-against-nvidia/ -
#AMD posts first Instinct #MI300X #MLPerf #benchmark results — roughly in line with #Nvidia #H100 (But only in #Llama2 70B).
Based on data AMD shared, 8x MI300X processors only slightly slower (23,512 TOPS) than 8x H100 SXM3 (24,323 TOPS), which can probably be called 'competitive' given how well Nvidia's software stack is optimized for popular #LLM like Llama 2. AMD MI300X system is slightly faster than the H100 machine in more or less real-world server benchmarks.
https://www.tomshardware.com/tech-industry/artificial-intelligence/amd-posts-first-instinct-mi300x-mlperf-benchmark-results-roughly-in-line-with-nvidia-h100-performance -
#AMD posts first Instinct #MI300X #MLPerf #benchmark results — roughly in line with #Nvidia #H100 (But only in #Llama2 70B).
Based on data AMD shared, 8x MI300X processors only slightly slower (23,512 TOPS) than 8x H100 SXM3 (24,323 TOPS), which can probably be called 'competitive' given how well Nvidia's software stack is optimized for popular #LLM like Llama 2. AMD MI300X system is slightly faster than the H100 machine in more or less real-world server benchmarks.
https://www.tomshardware.com/tech-industry/artificial-intelligence/amd-posts-first-instinct-mi300x-mlperf-benchmark-results-roughly-in-line-with-nvidia-h100-performance -
I found a way to render bar charts on #GitHub directly from #markdown, by hacking mermaid gantt chart. Not perfect but good enough, with colored bars for AMD/Intel/Nvidia/ARM. This puts #FluidX3D #GPU/#CPU benchmarks into perspective. AMD is pulling a hockey stick with #MI300X, Nvidia follows up with #H100 94GB. The RTX 4090 suddenly doesn't look that fast anymore. 🖖🥺
https://github.com/ProjectPhysX/FluidX3D?tab=readme-ov-file#single-gpucpu-benchmarks -
I found a way to render bar charts on #GitHub directly from #markdown, by hacking mermaid gantt chart. Not perfect but good enough, with colored bars for AMD/Intel/Nvidia/ARM. This puts #FluidX3D #GPU/#CPU benchmarks into perspective. AMD is pulling a hockey stick with #MI300X, Nvidia follows up with #H100 94GB. The RTX 4090 suddenly doesn't look that fast anymore. 🖖🥺
https://github.com/ProjectPhysX/FluidX3D?tab=readme-ov-file#single-gpucpu-benchmarks -
Nvidia #H100 NVL 94GB benchmarks for #FluidX3D #CFD are in, and it's unfathomably quick - 2x faster than H100 80GB. It's the fastest PCIe-form-factor #HPC #GPU ever by a long shot with 3938 GB/s. 🖖🤯
Only AMD #MI300X is still above.
https://github.com/ProjectPhysX/FluidX3D?tab=readme-ov-file#single-gpucpu-benchmarks -
Nvidia #H100 NVL 94GB benchmarks for #FluidX3D #CFD are in, and it's unfathomably quick - 2x faster than H100 80GB. It's the fastest PCIe-form-factor #HPC #GPU ever by a long shot with 3938 GB/s. 🖖🤯
Only AMD #MI300X is still above.
https://github.com/ProjectPhysX/FluidX3D?tab=readme-ov-file#single-gpucpu-benchmarks -
@chipsandcheese damn, #MI300X decimates #H100 in #FluidX3D. Delivers only ~60% of the 5.3TB/s spec sheet VRAM bandwidth, similar to MI100 and MI200, but the 8 HBM3 stacks are a monumental brute force upgrade.
-
@chipsandcheese damn, #MI300X decimates #H100 in #FluidX3D. Delivers only ~60% of the 5.3TB/s spec sheet VRAM bandwidth, similar to MI100 and MI200, but the 8 HBM3 stacks are a monumental brute force upgrade.
-
#Lenovo Shows Huge Optimism Towards #AMD’s Instinct #MI300X #AI Accelerators demand is at a record high — plans to offer AI solutions from all important hardware vendors
#Nvidia of course claims its #GPU are still faster than everything else, but with the right combination of price and performance, it looks like AMD and Intel can win some lucrative contracts. Nobody likes#Nvidia's dominance, it seems.
https://www.tomshardware.com/tech-industry/artificial-intelligence/lenovo-says-demand-for-amds-instinct-mi300-is-record-high-plans-to-offer-ai-solutions-from-all-important-hardware-vendors -
#Lenovo Shows Huge Optimism Towards #AMD’s Instinct #MI300X #AI Accelerators demand is at a record high — plans to offer AI solutions from all important hardware vendors
#Nvidia of course claims its #GPU are still faster than everything else, but with the right combination of price and performance, it looks like AMD and Intel can win some lucrative contracts. Nobody likes#Nvidia's dominance, it seems.
https://www.tomshardware.com/tech-industry/artificial-intelligence/lenovo-says-demand-for-amds-instinct-mi300-is-record-high-plans-to-offer-ai-solutions-from-all-important-hardware-vendors -
#Nvidia's #H100 #AI #GPU cost up to four times more than #AMD's competing #MI300X — AMD's chips cost $10K to $15K apiece; Nvidia's H100 has peaked beyond $40,000: Report https://www.tomshardware.com/tech-industry/artificial-intelligence/nvidias-h100-ai-gpus-cost-up-to-four-times-more-than-amds-competing-mi300x-amds-chips-cost-dollar10-to-dollar15k-apiece-nvidias-h100-has-peaked-beyond-dollar40000
-
#Nvidia's #H100 #AI #GPU cost up to four times more than #AMD's competing #MI300X — AMD's chips cost $10K to $15K apiece; Nvidia's H100 has peaked beyond $40,000: Report https://www.tomshardware.com/tech-industry/artificial-intelligence/nvidias-h100-ai-gpus-cost-up-to-four-times-more-than-amds-competing-mi300x-amds-chips-cost-dollar10-to-dollar15k-apiece-nvidias-h100-has-peaked-beyond-dollar40000
-
💡AMD acquisirà la startup IA Nod.ai
AMD acquisirà la startup di software Nod.ai per avere modelli IA ottimizzati per i suoi chip di AMD e contrastare NvidiaNod.ai https://gomoot.com/amd-acquisira-lo-startup-ia-nod-ai/
#a100 #AI #ia #intelligenzaartificiale #AMD #MI300 #mi300x #chip #gpu #cuda #h100 #shark #blender #GenAI #pytorch #nod_dot_ai #MLIR
-
#AMD Has a #GPU to Rival #Nvidia’s #H100
#MI300X is a GPU-only version of previously announced #MI300A supercomputing chip, which includes a #CPU and #GPU. The MI300A will be in El Capitan, a supercomputer coming next year to the #LosAlamos #NationalLaboratory. El Capitan is expected to surpass 2 exaflops of performance. The MI300X has 192GB of #HBM3, which Su said was 2.4 times more memory density than Nvidia’s H100. The SXM and PCIe versions of H100 have 80GB of HBM3.
https://www.hpcwire.com/2023/06/13/amd-has-a-gpu-to-rival-nvidias-h100/ -
#AMD Has a #GPU to Rival #Nvidia’s #H100
#MI300X is a GPU-only version of previously announced #MI300A supercomputing chip, which includes a #CPU and #GPU. The MI300A will be in El Capitan, a supercomputer coming next year to the #LosAlamos #NationalLaboratory. El Capitan is expected to surpass 2 exaflops of performance. The MI300X has 192GB of #HBM3, which Su said was 2.4 times more memory density than Nvidia’s H100. The SXM and PCIe versions of H100 have 80GB of HBM3.
https://www.hpcwire.com/2023/06/13/amd-has-a-gpu-to-rival-nvidias-h100/