home.social

#nvlm β€” Public Fediverse posts

Live and recent posts from across the Fediverse tagged #nvlm, aggregated by home.social.

fetched live
  1. #NVIDIA introduces #NVLM 1.0, a family of open-source #multimodal #LLMs:

    πŸ† Achieves state-of-the-art results on vision-language tasks, competing with #GPT4 and #Llama3V

    πŸ“Š 72B model outperforms on #OCRBench and #VQAv2 benchmarks

    πŸ“ˆ Shows improved accuracy on text-only tasks after multimodal training

    πŸ’» Excels in #math, #coding, and #reasoning across modalities

    🧠 Novel architecture enhances training efficiency and multimodal reasoning

    πŸ–ΌοΈ Introduces 1-D tile-tagging for improved performance on high-resolution images

    πŸ”¬ Emphasizes dataset quality and task diversity over scale in training

    πŸ”— Open-sourcing model weights and training code in Megatron-Core

    Learn more: research.nvidia.com/labs/adlr/

  2. #NVIDIA introduces #NVLM 1.0, a family of open-source #multimodal #LLMs:

    πŸ† Achieves state-of-the-art results on vision-language tasks, competing with #GPT4 and #Llama3V

    πŸ“Š 72B model outperforms on #OCRBench and #VQAv2 benchmarks

    πŸ“ˆ Shows improved accuracy on text-only tasks after multimodal training

    πŸ’» Excels in #math, #coding, and #reasoning across modalities

    🧠 Novel architecture enhances training efficiency and multimodal reasoning

    πŸ–ΌοΈ Introduces 1-D tile-tagging for improved performance on high-resolution images

    πŸ”¬ Emphasizes dataset quality and task diversity over scale in training

    πŸ”— Open-sourcing model weights and training code in Megatron-Core

    Learn more: research.nvidia.com/labs/adlr/

  3. #NVIDIA introduces #NVLM 1.0, a family of open-source #multimodal #LLMs:

    πŸ† Achieves state-of-the-art results on vision-language tasks, competing with #GPT4 and #Llama3V

    πŸ“Š 72B model outperforms on #OCRBench and #VQAv2 benchmarks

    πŸ“ˆ Shows improved accuracy on text-only tasks after multimodal training

    πŸ’» Excels in #math, #coding, and #reasoning across modalities

    🧠 Novel architecture enhances training efficiency and multimodal reasoning

    πŸ–ΌοΈ Introduces 1-D tile-tagging for improved performance on high-resolution images

    πŸ”¬ Emphasizes dataset quality and task diversity over scale in training

    πŸ”— Open-sourcing model weights and training code in Megatron-Core

    Learn more: research.nvidia.com/labs/adlr/

  4. #NVIDIA introduces #NVLM 1.0, a family of open-source #multimodal #LLMs:

    πŸ† Achieves state-of-the-art results on vision-language tasks, competing with #GPT4 and #Llama3V

    πŸ“Š 72B model outperforms on #OCRBench and #VQAv2 benchmarks

    πŸ“ˆ Shows improved accuracy on text-only tasks after multimodal training

    πŸ’» Excels in #math, #coding, and #reasoning across modalities

    🧠 Novel architecture enhances training efficiency and multimodal reasoning

    πŸ–ΌοΈ Introduces 1-D tile-tagging for improved performance on high-resolution images

    πŸ”¬ Emphasizes dataset quality and task diversity over scale in training

    πŸ”— Open-sourcing model weights and training code in Megatron-Core

    Learn more: research.nvidia.com/labs/adlr/

  5. #NVIDIA introduces #NVLM 1.0, a family of open-source #multimodal #LLMs:

    πŸ† Achieves state-of-the-art results on vision-language tasks, competing with #GPT4 and #Llama3V

    πŸ“Š 72B model outperforms on #OCRBench and #VQAv2 benchmarks

    πŸ“ˆ Shows improved accuracy on text-only tasks after multimodal training

    πŸ’» Excels in #math, #coding, and #reasoning across modalities

    🧠 Novel architecture enhances training efficiency and multimodal reasoning

    πŸ–ΌοΈ Introduces 1-D tile-tagging for improved performance on high-resolution images

    πŸ”¬ Emphasizes dataset quality and task diversity over scale in training

    πŸ”— Open-sourcing model weights and training code in Megatron-Core

    Learn more: research.nvidia.com/labs/adlr/