#adamw — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #adamw, aggregated by home.social.
-
Qwen4: Архитектура будущего
Тут надо сразу расставить точки, потому что вокруг названия путаница. Полноценного Qwen4 не существует . Есть Qwen3.8-Flash-Next – модель, которую сама команда Qwen называет «экспериментальная версия архитектуры, которая ляжет в основу Qwen4». 26 августа 2026 интернет на пару часов забыл про всё. «Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!» И цифры… 125 миллиардов параметров всего. Плюс ещё 51 млрд отдельным слоем. А активируется на каждый токен – лишь 6 миллиардов. Первая мысль была: так не бывает.
https://habr.com/ru/companies/gptunnel/articles/1076382/
#Qwen38FlashNext #архитектура_Qwen4 #Alibaba #MoE #Gated_DeltaNet #sparse_attention #Ngram_embedding #Muon_optimizer #AdamW #GLM53Flash
-
Qwen4: Архитектура будущего
Тут надо сразу расставить точки, потому что вокруг названия путаница. Полноценного Qwen4 не существует . Есть Qwen3.8-Flash-Next – модель, которую сама команда Qwen называет «экспериментальная версия архитектуры, которая ляжет в основу Qwen4». 26 августа 2026 интернет на пару часов забыл про всё. «Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!» И цифры… 125 миллиардов параметров всего. Плюс ещё 51 млрд отдельным слоем. А активируется на каждый токен – лишь 6 миллиардов. Первая мысль была: так не бывает.
https://habr.com/ru/companies/gptunnel/articles/1076382/
#Qwen38FlashNext #архитектура_Qwen4 #Alibaba #MoE #Gated_DeltaNet #sparse_attention #Ngram_embedding #Muon_optimizer #AdamW #GLM53Flash
-
Qwen4: Архитектура будущего
Тут надо сразу расставить точки, потому что вокруг названия путаница. Полноценного Qwen4 не существует . Есть Qwen3.8-Flash-Next – модель, которую сама команда Qwen называет «экспериментальная версия архитектуры, которая ляжет в основу Qwen4». 26 августа 2026 интернет на пару часов забыл про всё. «Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight!» И цифры… 125 миллиардов параметров всего. Плюс ещё 51 млрд отдельным слоем. А активируется на каждый токен – лишь 6 миллиардов. Первая мысль была: так не бывает.
https://habr.com/ru/companies/gptunnel/articles/1076382/
#Qwen38FlashNext #архитектура_Qwen4 #Alibaba #MoE #Gated_DeltaNet #sparse_attention #Ngram_embedding #Muon_optimizer #AdamW #GLM53Flash
-
Adam W And Anwar Jibawi To Star In Animated Series ‘Ollie The Octopus’
EXCLUSIVE: AI Studio Staircase Studios AI and Made by Us Studios are teaming on a joint venture to…
#NewsBeep #News #Movies #AdamW #AU #Australia #Entertainment #MadeByUs #OllietheOctupus
https://www.newsbeep.com/au/700267/ -
Adam W And Anwar Jibawi To Star In Animated Series ‘Ollie The Octopus’
EXCLUSIVE: AI Studio Staircase Studios AI and Made by Us Studios are teaming on a joint venture to…
#NewsBeep #News #Movies #AdamW #AU #Australia #Entertainment #MadeByUs #OllietheOctupus
https://www.newsbeep.com/au/700267/ -
Staircase Studios AI And Made By Us Partner On New Animated Series ‘Ollie The Octopus’ Featuring Digital Creators Adam W, Hannah Stocking And Anwar Jibawi
#News #AdamW #MadeByUs #OllietheOctupushttps://deadline.com/2026/05/ollie-the-octpus-adam-w-hannah-stocking-and-anwar-jibawi-1236927962/
-
Staircase Studios AI And Made By Us Partner On New Animated Series ‘Ollie The Octopus’ Featuring Digital Creators Adam W, Hannah Stocking And Anwar Jibawi
#News #AdamW #MadeByUs #OllietheOctupushttps://deadline.com/2026/05/ollie-the-octpus-adam-w-hannah-stocking-and-anwar-jibawi-1236927962/
-
Staircase Studios AI And Made By Us Partner On New Animated Series ‘Ollie The Octopus’ Featuring Digital Creators Adam W, Hannah Stocking And Anwar Jibawi
#News #AdamW #MadeByUs #OllietheOctupushttps://deadline.com/2026/05/ollie-the-octpus-adam-w-hannah-stocking-and-anwar-jibawi-1236927962/
-
Staircase Studios AI And Made By Us Partner On New Animated Series ‘Ollie The Octopus’ Featuring Digital Creators Adam W, Hannah Stocking And Anwar Jibawi
#News #AdamW #MadeByUs #OllietheOctupushttps://deadline.com/2026/05/ollie-the-octpus-adam-w-hannah-stocking-and-anwar-jibawi-1236927962/
-
So sánh Muon và AdamW trong đào tạo mô hình AI. Muon có thể underfit trong khi AdamW overfit. Cả hai mô hình đều đạt độ chính xác cao nhưng AdamW nhỉnh hơn. #Muon #AdamW #AI #MachineLearning #ĐàoTạoMôHình #TríTuệNhânTạo #Optimization #DeepLearning
https://www.reddit.com/r/LocalLLaMA/comments/1owa4ag/muon_underfits_adamw_overfits/
-
‘YouTube Does Not Operate That Way’: How YouTube Creators Are Schooling Hollywood
#IndieWire #Analysis #News #AdamW #DharMann #FutureofFilmmaking #InDevelopment #NealMohan #YouTube -
‘YouTube Does Not Operate That Way’: How YouTube Creators Are Schooling Hollywood
#IndieWire #Analysis #News #AdamW #DharMann #FutureofFilmmaking #InDevelopment #NealMohan #YouTube -
‘YouTube Does Not Operate That Way’: How YouTube Creators Are Schooling Hollywood
#IndieWire #Analysis #News #AdamW #DharMann #FutureofFilmmaking #InDevelopment #NealMohan #YouTube -
‘YouTube Does Not Operate That Way’: How YouTube Creators Are Schooling Hollywood
#IndieWire #Analysis #News #AdamW #DharMann #FutureofFilmmaking #InDevelopment #NealMohan #YouTube -
Practical Efficiency of Muon for Pretraining
O Muon alcança o mesmo loss com 10–15% menos tokens e converge mais depressa, preservando a eficiência de dados mesmo com tamanhos de lote muito grandes. Recomenda-se como sucessor “drop-in” do AdamW em grande escala.
-
Practical Efficiency of Muon for Pretraining
O Muon alcança o mesmo loss com 10–15% menos tokens e converge mais depressa, preservando a eficiência de dados mesmo com tamanhos de lote muito grandes. Recomenda-se como sucessor “drop-in” do AdamW em grande escala.
-
Practical Efficiency of Muon for Pretraining
O Muon alcança o mesmo loss com 10–15% menos tokens e converge mais depressa, preservando a eficiência de dados mesmo com tamanhos de lote muito grandes. Recomenda-se como sucessor “drop-in” do AdamW em grande escala.
-
Practical Efficiency of Muon for Pretraining
O Muon alcança o mesmo loss com 10–15% menos tokens e converge mais depressa, preservando a eficiência de dados mesmo com tamanhos de lote muito grandes. Recomenda-se como sucessor “drop-in” do AdamW em grande escala.
-
Understanding AdamW through Proximal Methods and Scale-Freeness
Zhenxun Zhuang, Mingrui Liu, Ashok Cutkosky, Francesco Orabona
-
Understanding AdamW through Proximal Methods and Scale-Freeness
Zhenxun Zhuang, Mingrui Liu, Ashok Cutkosky, Francesco Orabona
-
Understanding AdamW through Proximal Methods and Scale-Freeness
Zhenxun Zhuang, Mingrui Liu, Ashok Cutkosky, Francesco Orabona
-
Understanding AdamW through Proximal Methods and Scale-Freeness
Zhenxun Zhuang, Mingrui Liu, Ashok Cutkosky, Francesco Orabona
-
Understanding AdamW through Proximal Methods and Scale-Freeness
Zhenxun Zhuang, Mingrui Liu, Ashok Cutkosky, Francesco Orabona