home.social
  1. NVFP4 and MXFP4 multiply the same sixteen 4-bit values; the difference is one byte of scale per block.

    We measured what that byte costs (41.5% of MXFP4 blocks clip their own maximum; about 9% of perplexity under plain rounding) and asked ptxas what eight NVIDIA targets do with each format.

    No GPU needed, every number tiered by confidence.

    open.substack.com/pub/software

    #GPU #CUDA #MachineLearning #Quantization

Share on Mastodon

Enter the server where you have an account.