home.social

#nanochat — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #nanochat, aggregated by home.social.

fetched live
  1. Т9 за сотню долларов, ассистент за миллиард: сколько на самом деле стоит обучить LLM с нуля

    «Лучший ChatGPT за $100», «китайцы обучили GPT-4 за шесть миллионов», «модель с нуля на одной игровой карте за вечер» — за последний год такие заголовки идут потоком. Почти все цифры в них честные. Беда в том, что они про разные вещи. За словами «обучить LLM» прячется лестница в семь порядков: от десятков долларов на карте под столом до миллиарда на фронтире. На каждой ступени за деньги покупают разное, и перепутать ступени дорого. Я работаю CTO ML-команды. Расчет «взять открытые веса / дообучить / обучить с нуля» мы проделываем регулярно, и почти всегда он заканчивается не там, где ждет заказчик. Ниже — вся лестница с числами и первоисточниками: что реально обучить на RTX 5090, где начинается аренда кластера, из чего складывается смета фронтир-модели и почему «модель за $6 млн» и «компания с инфраструктурой на миллиард» — правда одновременно. Посчитать свою LLM

    habr.com/ru/articles/1060916/

    #LLM #обучение_LLM #pretraining #машинное_обучение #нейросети #GPU #RTX_5090 #H100 #стоимость_обучения_моделей #nanochat

  2. Т9 за сотню долларов, ассистент за миллиард: сколько на самом деле стоит обучить LLM с нуля

    «Лучший ChatGPT за $100», «китайцы обучили GPT-4 за шесть миллионов», «модель с нуля на одной игровой карте за вечер» — за последний год такие заголовки идут потоком. Почти все цифры в них честные. Беда в том, что они про разные вещи. За словами «обучить LLM» прячется лестница в семь порядков: от десятков долларов на карте под столом до миллиарда на фронтире. На каждой ступени за деньги покупают разное, и перепутать ступени дорого. Я работаю CTO ML-команды. Расчет «взять открытые веса / дообучить / обучить с нуля» мы проделываем регулярно, и почти всегда он заканчивается не там, где ждет заказчик. Ниже — вся лестница с числами и первоисточниками: что реально обучить на RTX 5090, где начинается аренда кластера, из чего складывается смета фронтир-модели и почему «модель за $6 млн» и «компания с инфраструктурой на миллиард» — правда одновременно. Посчитать свою LLM

    habr.com/ru/articles/1060916/

    #LLM #обучение_LLM #pretraining #машинное_обучение #нейросети #GPU #RTX_5090 #H100 #стоимость_обучения_моделей #nanochat

  3. Т9 за сотню долларов, ассистент за миллиард: сколько на самом деле стоит обучить LLM с нуля

    «Лучший ChatGPT за $100», «китайцы обучили GPT-4 за шесть миллионов», «модель с нуля на одной игровой карте за вечер» — за последний год такие заголовки идут потоком. Почти все цифры в них честные. Беда в том, что они про разные вещи. За словами «обучить LLM» прячется лестница в семь порядков: от десятков долларов на карте под столом до миллиарда на фронтире. На каждой ступени за деньги покупают разное, и перепутать ступени дорого. Я работаю CTO ML-команды. Расчет «взять открытые веса / дообучить / обучить с нуля» мы проделываем регулярно, и почти всегда он заканчивается не там, где ждет заказчик. Ниже — вся лестница с числами и первоисточниками: что реально обучить на RTX 5090, где начинается аренда кластера, из чего складывается смета фронтир-модели и почему «модель за $6 млн» и «компания с инфраструктурой на миллиард» — правда одновременно. Посчитать свою LLM

    habr.com/ru/articles/1060916/

    #LLM #обучение_LLM #pretraining #машинное_обучение #нейросети #GPU #RTX_5090 #H100 #стоимость_обучения_моделей #nanochat

  4. 🤣 Ah, the noble quest of porting #nanochat to a TPU—because apparently, #PyTorch wasn't hard enough! 🚀 Spoiler alert: everything breaks, but hey, at least you'll learn to love debugging! 😅
    github.com/tucan9389/nanochat- #TPU #debugging #techhumor #learningjourney #HackerNews #ngated

  5. 🤣 Ah, the noble quest of porting #nanochat to a TPU—because apparently, #PyTorch wasn't hard enough! 🚀 Spoiler alert: everything breaks, but hey, at least you'll learn to love debugging! 😅
    github.com/tucan9389/nanochat- #TPU #debugging #techhumor #learningjourney #HackerNews #ngated

  6. 🤣 Ah, the noble quest of porting #nanochat to a TPU—because apparently, #PyTorch wasn't hard enough! 🚀 Spoiler alert: everything breaks, but hey, at least you'll learn to love debugging! 😅
    github.com/tucan9389/nanochat- #TPU #debugging #techhumor #learningjourney #HackerNews #ngated

  7. 🤣 Ah, the noble quest of porting #nanochat to a TPU—because apparently, #PyTorch wasn't hard enough! 🚀 Spoiler alert: everything breaks, but hey, at least you'll learn to love debugging! 😅
    github.com/tucan9389/nanochat- #TPU #debugging #techhumor #learningjourney #HackerNews #ngated

  8. 🤣 Ah, the noble quest of porting #nanochat to a TPU—because apparently, #PyTorch wasn't hard enough! 🚀 Spoiler alert: everything breaks, but hey, at least you'll learn to love debugging! 😅
    github.com/tucan9389/nanochat- #TPU #debugging #techhumor #learningjourney #HackerNews #ngated

  9. habrGPT. Обучим LLM 0.5B с нуля на статьях Хабра с помощью nanochat от Карпатого. Обучение fp8 дома и сравнение с bf16

    Раскрученный проект nanochat обещает "за $100 обучите свой ChatGPT". В оригинале правда обучение на 8xH100 80Gb, но никто не мешает обучить его на домашнем более слабом железе. Попробуем обучить LLM с нуля на статьях Хабра, посмотреть хватит ли там материала, чтобы модель хотя бы смогла связать пару слов.

    habr.com/ru/articles/1054710/

    #nanochat #sft #dpo #обучение #llm #llm_с_нуля #saiga

  10. habrGPT. Обучим LLM 0.5B с нуля на статьях Хабра с помощью nanochat от Карпатого. Обучение fp8 дома и сравнение с bf16

    Раскрученный проект nanochat обещает "за $100 обучите свой ChatGPT". В оригинале правда обучение на 8xH100 80Gb, но никто не мешает обучить его на домашнем более слабом железе. Попробуем обучить LLM с нуля на статьях Хабра, посмотреть хватит ли там материала, чтобы модель хотя бы смогла связать пару слов.

    habr.com/ru/articles/1054710/

    #nanochat #sft #dpo #обучение #llm #llm_с_нуля #saiga

  11. habrGPT. Обучим LLM 0.5B с нуля на статьях Хабра с помощью nanochat от Карпатого. Обучение fp8 дома и сравнение с bf16

    Раскрученный проект nanochat обещает "за $100 обучите свой ChatGPT". В оригинале правда обучение на 8xH100 80Gb, но никто не мешает обучить его на домашнем более слабом железе. Попробуем обучить LLM с нуля на статьях Хабра, посмотреть хватит ли там материала, чтобы модель хотя бы смогла связать пару слов.

    habr.com/ru/articles/1054710/

    #nanochat #sft #dpo #обучение #llm #llm_с_нуля #saiga

  12. @[email protected]

    nanochat server in 916 bytes, implemented entirely in uxn :D

    canon impl by @d6: https://git.phial.org/d6/nanochat

    difference is that mine has a limit of 65535 messages before the count overflows and you can't read any more :P
    oh and also you can DoS by starting a connection and never sending a newline…

    but at least it can avoid buffer overflow by kicking you!

    src: https://codeberg.org/notchoc/nanochat.tal
    rom: https://codeberg.org/notchoc/nanochat.tal/raw/branch/main/bin/nanochat.rom
    run: uxn2 -N nanochat.rom

    !!! DISCLAIMER: YOU ACCEPT EVERYTHING THAT WILL HAPPEN FROM NOW ON !!! UXN2 STILL DOES NOT HAVE SANDBOXING (YET) SO USE AT YOUR OWN RISK


    #uxn #uxntal #nanochat
  13. @[email protected]

    nanochat server in 916 bytes, implemented entirely in uxn :D

    canon impl by @d6: https://git.phial.org/d6/nanochat

    difference is that mine has a limit of 65535 messages before the count overflows and you can't read any more :P
    oh and also you can DoS by starting a connection and never sending a newline…

    but at least it can avoid buffer overflow by kicking you!

    src: https://codeberg.org/notchoc/nanochat.tal
    rom: https://codeberg.org/notchoc/nanochat.tal/raw/branch/main/bin/nanochat.rom
    run: uxn2 -N nanochat.rom

    !!! DISCLAIMER: YOU ACCEPT EVERYTHING THAT WILL HAPPEN FROM NOW ON !!! UXN2 STILL DOES NOT HAVE SANDBOXING (YET) SO USE AT YOUR OWN RISK


    #uxn #uxntal #nanochat
  14. @[email protected]

    nanochat server in 916 bytes, implemented entirely in uxn :D

    canon impl by @d6: https://git.phial.org/d6/nanochat

    difference is that mine has a limit of 65535 messages before the count overflows and you can't read any more :P
    oh and also you can DoS by starting a connection and never sending a newline…

    but at least it can avoid buffer overflow by kicking you!

    src: https://codeberg.org/notchoc/nanochat.tal
    rom: https://codeberg.org/notchoc/nanochat.tal/raw/branch/main/bin/nanochat.rom
    run: uxn2 -N nanochat.rom

    !!! DISCLAIMER: YOU ACCEPT EVERYTHING THAT WILL HAPPEN FROM NOW ON !!! UXN2 STILL DOES NOT HAVE SANDBOXING (YET) SO USE AT YOUR OWN RISK


    #uxn #uxntal #nanochat
  15. @[email protected]

    nanochat server in 916 bytes, implemented entirely in uxn :D

    canon impl by @d6: https://git.phial.org/d6/nanochat

    difference is that mine has a limit of 65535 messages before the count overflows and you can't read any more :P
    oh and also you can DoS by starting a connection and never sending a newline…

    but at least it can avoid buffer overflow by kicking you!

    src: https://codeberg.org/notchoc/nanochat.tal
    rom: https://codeberg.org/notchoc/nanochat.tal/raw/branch/main/bin/nanochat.rom
    run: uxn2 -N nanochat.rom

    !!! DISCLAIMER: YOU ACCEPT EVERYTHING THAT WILL HAPPEN FROM NOW ON !!! UXN2 STILL DOES NOT HAVE SANDBOXING (YET) SO USE AT YOUR OWN RISK


    #uxn #uxntal #nanochat
  16. @[email protected]

    nanochat server in 916 bytes, implemented entirely in uxn :D

    canon impl by @d6: https://git.phial.org/d6/nanochat

    difference is that mine has a limit of 65535 messages before the count overflows and you can't read any more :P
    oh and also you can DoS by starting a connection and never sending a newline…

    but at least it can avoid buffer overflow by kicking you!

    src: https://codeberg.org/notchoc/nanochat.tal
    rom: https://codeberg.org/notchoc/nanochat.tal/raw/branch/main/bin/nanochat.rom
    run: uxn2 -N nanochat.rom

    !!! DISCLAIMER: YOU ACCEPT EVERYTHING THAT WILL HAPPEN FROM NOW ON !!! UXN2 STILL DOES NOT HAVE SANDBOXING (YET) SO USE AT YOUR OWN RISK


    #uxn #uxntal #nanochat
  17. Ah yes, we've reached the zenith of #AI #enlightenment where 'agents' conduct 'research' on a 'nanochat' with a singular GPU — the equivalent of having a hamster power your nuclear reactor 🚀🔍⚡. Is this #progress or is the tech world just throwing #buzzwords on a digital dartboard and calling it innovation? 😂🤖
    github.com/karpathy/autoresear #Nanochat #Innovation #HackerNews #ngated

  18. Ah yes, we've reached the zenith of #AI #enlightenment where 'agents' conduct 'research' on a 'nanochat' with a singular GPU — the equivalent of having a hamster power your nuclear reactor 🚀🔍⚡. Is this #progress or is the tech world just throwing #buzzwords on a digital dartboard and calling it innovation? 😂🤖
    github.com/karpathy/autoresear #Nanochat #Innovation #HackerNews #ngated

  19. Ah yes, we've reached the zenith of #AI #enlightenment where 'agents' conduct 'research' on a 'nanochat' with a singular GPU — the equivalent of having a hamster power your nuclear reactor 🚀🔍⚡. Is this #progress or is the tech world just throwing #buzzwords on a digital dartboard and calling it innovation? 😂🤖
    github.com/karpathy/autoresear #Nanochat #Innovation #HackerNews #ngated

  20. Ah yes, we've reached the zenith of #AI #enlightenment where 'agents' conduct 'research' on a 'nanochat' with a singular GPU — the equivalent of having a hamster power your nuclear reactor 🚀🔍⚡. Is this #progress or is the tech world just throwing #buzzwords on a digital dartboard and calling it innovation? 😂🤖
    github.com/karpathy/autoresear #Nanochat #Innovation #HackerNews #ngated

  21. 🚀 Đánh bại GPT-2 với chi phí dưới $100! Andrej Karpathy chia sẻ hành trình nanochat - chỉ 3 giờ huấn luyện trên 8×H100 đã vượt qua GPT-2 trong benchmark CORE. Bài viết tiết lộ chi tiết kiến trúc, tối ưu hóa và script để tái tạo kết quả.

    #AI #MachineLearning #TríTuệNhânTạo #HọcMáy #NanoChat #GPT2

    reddit.com/r/LocalLLaMA/commen

  22. #nanochat #AMD #AI #MáyHọc #CôngNghệ
    Phân tích từ đầu tới cuối về nanochat với phần cứng AMD MI300X và tín dụng phát triển. Bài viết cập nhật tiến trình xây dựng mô hình, bao gồm RMSNorm, RoPE, GQA và KVCache. Tiếp theo: Muon, DistAdamW. Mời mọi người góp ý, phản hồi để cải thiện!

    #AIimplementation #MachineLearning #AMDmi300x #Code #Math #Debug #VietnamAI #Transformer #OpenSource

    reddit.com/r/LocalLLaMA/commen

  23. NanoChat đã chính thức được tích hợp vào thư viện Hugging Face Transformers! Bài viết chuyên sâu mới nhất đi sâu vào kiến trúc NanoChat, quy trình tích hợp và hướng dẫn sử dụng các công cụ như Torch, TRL, vLLM cho suy luận và huấn luyện. Khám phá ngay!

    #NanoChat #Transformers #HuggingFace #AIVietnam #MôHìnhAI #HọcMáy #AI #MachineLearning #DeepLearning #LLM #NLP #TechNews

    reddit.com/r/LocalLLaMA/commen

  24. Practicing #TUI app development with #python #textual package.

    Femtochat is a #nanochat client for staying in the #uxn loop.

  25. Practicing #TUI app development with #python #textual package.

    Femtochat is a #nanochat client for staying in the #uxn loop.

  26. Practicing #TUI app development with #python #textual package.

    Femtochat is a #nanochat client for staying in the #uxn loop.

  27. Practicing #TUI app development with #python #textual package.

    Femtochat is a #nanochat client for staying in the #uxn loop.

  28. Practicing #TUI app development with #python #textual package.

    Femtochat is a #nanochat client for staying in the #uxn loop.

  29. 💡 Build your own ChatGPT-style AI in 4 hours for ₹10,000 💬
    Yes — that’s real. Nanochat is unlocking AI for every curious mind in India 🇮🇳. No paywalls, no secrets — just open code, fast GPUs, and imagination.
    Be the one who doesn’t just use AI — builds it. 🧠
    #Nanochat #AIforAll #OpenSource
    🔗
    medium.com/@rogt.x1997/from-10

  30. Tech update: Andrej Karpathy lancer nanochat để huấn luyện mô hình AI chỉ $100! Hãy thử nghiệm qua repo GitHub của họ. #AI #NanoChat #AndréKarpathy #MachineLearning #C regarda

    reddit.com/r/LocalLLaMA/commen

  31. Karpathy's NanoChat gần 90% code do AI viết cùng nhiều chú thích, thể hiện sự kết hợp giữa trí tuệ nhân tạo và sự sáng tạo của người. #karpathy #nanochat #AI #technology #codingagent #vi versuchtech #artificialintelligence #code

    reddit.com/r/singularity/comme

  32. Karpathy's NanoChat gần 90% code do AI viết cùng nhiều chú thích, thể hiện sự kết hợp giữa trí tuệ nhân tạo và sự sáng tạo của người. #karpathy #nanochat #AI #technology #codingagent #vi versuchtech #artificialintelligence #code

    reddit.com/r/singularity/comme

  33. Karpathy's NanoChat gần 90% code do AI viết cùng nhiều chú thích, thể hiện sự kết hợp giữa trí tuệ nhân tạo và sự sáng tạo của người. #karpathy #nanochat #AI #technology #codingagent #vi versuchtech #artificialintelligence #code

    reddit.com/r/singularity/comme

  34. Karpathy's NanoChat gần 90% code do AI viết cùng nhiều chú thích, thể hiện sự kết hợp giữa trí tuệ nhân tạo và sự sáng tạo của người. #karpathy #nanochat #AI #technology #codingagent #vi versuchtech #artificialintelligence #code

    reddit.com/r/singularity/comme