home.social

#deepseekv3 — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #deepseekv3, aggregated by home.social.

fetched live
  1. RT @dunik_7: TRANSLASATION: Ein Labor der Tsinghua-Universität hat ein Projekt auf GitHub veröffentlicht, das einen H100-Rack im Wert von 400.000 US-Dollar durch eine einzelne 24-GB-Grafikkarte ersetzt. Das Projekt heißt ktransformers, und der Trick ist fast schon lächerlich einfach: Die Experten, die Sie tatsächlich nutzen, bleiben auf der GPU, während die anderen auf der CPU warten, bis sie benötigt werden. / DeepSeek-V3 und R1 mit 139K Kontext in 24GB VRAM / bis zu 28-fache Geschwindigkeitssteigerung gegenüber dem Standard-Setup / Fine-Tuning von DeepSeek-V3 über vier RTX 4090 statt eines Rechenzentrums / entwickelt vom MADSys-Labor der Tsinghua-Universität, nicht von einem Startup mit einer Landing Page. Apache 2.0, bereits über 17.000 Sterne. - github.com/kvcache-ai/ktransfo merken.

    mehr auf Arint.info

    #AIResearch #DeepSeekV3 #ktransformers #MachineLearning #OpenSource #TsinghuaUniversity #arint_info

    https://x.com/dunik_7/status/2078065378563887290#m

  2. DeepSeek-V3 from Scratch: Mixture of Experts (MoE) Table of Contents DeepSeek-V3 from Scratch: Mixture of Experts (MoE) The Scaling Challenge in Neural Networks Mixture of Experts (MoE): Mathematic...

    #Deep #Learning #DeepSeek #Machine #Learning #Neural #Networks #Tutorial #deepseek-v3 #expert #routing

    Origin | Interest | Match
  3. DeepSeek-V3 from Scratch: Mixture of Experts (MoE) Table of Contents DeepSeek-V3 from Scratch: Mixture of Experts (MoE) The Scaling Challenge in Neural Networks Mixture of Experts (MoE): Mathematic...

    #Deep #Learning #DeepSeek #Machine #Learning #Neural #Networks #Tutorial #deepseek-v3 #expert #routing

    Origin | Interest | Match
  4. 日本樂天推自家「AI 3.0」模型 源碼竟顯示使用 DeepSeek 基礎模型
    樂天集團 (Rakuten) 3 月 17 日公開旗下最新日語大型語言模型「Rakuten AI 3.0」,惟技術人員隨即發現 Hugging Face 上的設定檔案顯示其架構與中國 AI 公司 DeepSeek 的 DeepSeek-V3 模型高度吻合,兼且發布時被指悄然移除 DeepSeek-V3 原有開源授權聲明,觸發開源社群強烈批評,樂天面對查詢時拒絕披露基礎模型來源,僅稱「非公開」。
    #人工智能 #AI #DeepSeek #DeepSeek-V3
    unwire.hk/2026/03/22/rakuten-a

  5. Build DeepSeek-V3: Multi-Head Latent Attention (MLA) Architecture Table of Contents Build DeepSeek-V3: Multi-Head Latent Attention (MLA) Architecture The KV Cache Memory Problem in DeepSeek-V3 Mult...

    #Deep #Learning #Large #Language #Models #PyTorch #Transformers #Tutorial #attention #mechanisms #deepseek-v3

    Origin | Interest | Match
  6. DeepSeek-V3 Model: Theory, Config, and Rotary Positional Embeddings Table of Contents DeepSeek-V3 Model: Theory, Config, and Rotary Positional Embeddings Introduction to the DeepSeek-V3 Model The F...

    #DeepSeek-V3 #KV #Cache #MultiHead #Latent #Attention #RoPE #Tutorial #deepseekv3 #kv #cache

    Origin | Interest | Match
  7. Beating GPT-5: DeepSeekMath-V2 Self-Corrects Logic Errors Presentational View Introduction Mathematics with the aid of artificial intelligence, is advancing rapidly. Innovations such as informal th...

    #ai-in-mathematics #deepseekmath-v2 #deepseek-v3 #open-source-ai-model #theorem-proving

    Origin | Interest | Match
  8. New benchmark shows Gemini 3 Pro outpaces Gemini 2.5 in trust, ethics and safety—69% vs 16%. The study, led by Phelim Bradley and Prolific, also pits DeepSeek V3 against the models, highlighting gaps in performance and reasoning. Dive into the full analysis for the numbers and implications. #Gemini3Pro #Gemini2_5 #DeepSeekV3 #TrustAndSafety

    🔗 aidailypost.com/news/gemini-3-

  9. New benchmarks show Mixture‑of‑Experts models on NVIDIA’s Blackwell NVL72 run up to 10× faster than on Hopper GPUs. The GB200 architecture and DeepSeek‑V3 optimizations push open‑source AI research forward. Dive into the details and see how this leap could reshape training pipelines. #MixtureOfExperts #NVIDIA #Blackwell #DeepSeekV3

    🔗 aidailypost.com/news/mixtureof

  10. 🚀 Welcome GLM-4.6 the Latest flagship #opensource #AI #llm with advanced agentic, reasoning & coding capabilities

    ⚡ Performance improvements over #GLM45 with competitive advantages against #DeepSeekV3 and #ClaudeSonnet4 across 8 public benchmarks covering agents, reasoning & coding

    🧵 👇

  11. 🚀 Welcome GLM-4.6 the Latest flagship #opensource #AI #llm with advanced agentic, reasoning & coding capabilities

    ⚡ Performance improvements over #GLM45 with competitive advantages against #DeepSeekV3 and #ClaudeSonnet4 across 8 public benchmarks covering agents, reasoning & coding

    🧵 👇

  12. Насколько зацензурен и опасен DeepSeek?

    Насколько предвзят искусственный интеллект? Принято ругать нейросети за трансляцию стереотипов человеческого мышления, которые были подсмотрены в датасетах предобучения. На деле ИИ куда более аккуратен, чем можно ожидать. Хороший пример — генерация фотографий бабочек. Как правило, дизайнеры-люди очень любят изображать бабочек в мёртвом виде. Дело в том, что энтомологи руководствуются строгими визуальными стандартами: вид сверху, расправленные на 180° крылья, чистый фон, симметрия.

    habr.com/ru/articles/949540/

    #DeepSeek #DeepSeekR1 #DeepSeekV3 #КНР #Китай #большие_языковые_модели #БЯМ #искусственный_интеллект #предвзятость #цензура

  13. 🧩 #Llama4Maverick nutzt 128 Experten für deutlich mehr Rechenleistung und schlägt sogar #GPT4o und #Gemini20 in Benchmarks – bei nur der Hälfte der aktiven Parameter von #DeepSeekv3.

    🎓 Beide #KIModelle wurden mithilfe des riesigen Lehrmodells #Llama4 Behemoth trainiert, das mit 288 Milliarden aktiven Parametern zu den leistungsstärksten weltweit zählt.

    👉 eicker.TV #Technik #Medien #Politik #Wirtschaft (2/2)

  14. 🧩 #Llama4Maverick nutzt 128 Experten für deutlich mehr Rechenleistung und schlägt sogar #GPT4o und #Gemini20 in Benchmarks – bei nur der Hälfte der aktiven Parameter von #DeepSeekv3.

    🎓 Beide #KIModelle wurden mithilfe des riesigen Lehrmodells #Llama4 Behemoth trainiert, das mit 288 Milliarden aktiven Parametern zu den leistungsstärksten weltweit zählt.

    👉 eicker.TV #Technik #Medien #Politik #Wirtschaft (2/2)

  15. Studie: #KI #Chatbots sind beim Zitieren von #News unbrauchbar
    derstandard.at/story/300000026

    "Untersucht wurden #ChatGPT Search (#OpenAI), #Perplexity, Perplexity Pro (Perplexity AI), #Gemini 2.0 Flash (#Google), #DeepseekV3 Search (#Deepseek), #Grok-2 Search, Grok-3 Search Beta (#xAI) sowie #Copilot (#Microsoft und OpenAI)."

    "#Grok3 [...] lieferte gleich in 96 Prozent aller Fälle falsche Antworten." 🤣

    #Nachrichten #Algorithmen #Automatisierung

  16. Studie: #KI #Chatbots sind beim Zitieren von #News unbrauchbar
    derstandard.at/story/300000026

    "Untersucht wurden #ChatGPT Search (#OpenAI), #Perplexity, Perplexity Pro (Perplexity AI), #Gemini 2.0 Flash (#Google), #DeepseekV3 Search (#Deepseek), #Grok-2 Search, Grok-3 Search Beta (#xAI) sowie #Copilot (#Microsoft und OpenAI)."

    "#Grok3 [...] lieferte gleich in 96 Prozent aller Fälle falsche Antworten." 🤣

    #Nachrichten #Algorithmen #Automatisierung

  17. »Chinese #AIlab #DeepSeek just released the latest version of their enormous #DeepSeekv3 model: The license is #MIT (that's new - previous DeepSeek v3 had a custom license).« simonwillison.net/2025/Mar/24/ #tech #media

  18. »Chinese #AIlab #DeepSeek just released the latest version of their enormous #DeepSeekv3 model: The license is #MIT (that's new - previous DeepSeek v3 had a custom license).« simonwillison.net/2025/Mar/24/ #tech #media

  19. 🚀 DeepSeek V3 vs ChatGPT-4o: Which One Reigns Supreme?🤖

    AI is evolving fast! 🏎️ DeepSeek V3 and ChatGPT-4o are two of the most powerful LLMs in 2025. But which one is better?

    🔍 We compare:
    ✅ Accuracy & performance
    ✅ Multimodal capabilities
    ✅ Speed & efficiency
    ✅ Real-world applications

    📖 Read the full breakdown here:

    radargit.com/2025/02/03/deepse

    Which AI model do you prefer? Comment below! 👇

    #AI #DeepSeekV3 #ChatGPT4o #ArtificialIntelligence #Tech #MachineLearning #AICompari

  20. "A key component of the success is that it is #opensource. #DeepSeek-V3 is on GitHub with detailed docs on how it can be replicated. This has fueled a rush of people to try to make their own models." https://baixacultura.org/2025/01/29/a-corrida-da-ia-ganha-um-novo-capitulo-chines-e-open-source/

    A corrida da IA ganha um novo ...

  21. The Chinese firm said training the model cost just $5.6 million. Alibaba Cloud followed with a new generative AI model, while Microsoft alleges DeepSeek ‘distilled’ OpenAI’s work.#artificialintelligence #chatgpt #deepseek #deepseekr1 #deepseek-v3 #generativeai #Microsoft #nvidia #openai #reasoningmodels
    DeepSeek Chatbot Beats OpenAI on App Store Leaderboard
  22. Anyone out there who tried to run #deepseek V3 locally on a #linux machine? I'm curious if it can run with a consumer #nvidia or #amd card?

    #deepseekv3

  23. Hey everyone, and welcome to Lety Does Unironically Posting About Girls Going to Glory Holes on Main Cause She's Really Upset About All the AI Misinformation Going Around Right Now

    peertube.doesstuff.social/w/e7

    #DeepSeek #DeepSeekR1 #DeepSeekV3 #OpenAI #ChatGPT #AI #LLM

  24. DeepSeek V3 vs. Gemini 2.0: Wer dominiert?

    - DeepSeek V3: 671 Mrd. Parameter!
    - Gemini 2.0: Multimodal & 1 Mio. Token-Kontext!
    - Revolution in Effizienz & Vielseitigkeit.

    #ai #ki #deepseekv3 #gemini2 #artificialintelligence

    kinews24.de/deepseek-v3-vs-gem

  25. DeepSeek-V3: A New Era in Open-Source AI with a Comedic Twist

    In the ever-evolving landscape of artificial intelligence, public and open models are once again catching up with their proprietary counterparts. The recent launch of DeepSeek-V3 has stirred excitement by outperforming even Sonnet and ChatGPT in certain benchmarks. This open-source model, developed by Hangzhou DeepSeek Artificial Intelligence and Beijing DeepSeek Artificial Intelligence, boasts an impressive 671 billion parameters, making it the largest model in the open-source community. With its advanced Mixture-of-Experts (MoE) architecture and innovative technologies, DeepSeek-V3 not only outperforms its open-source counterparts like LLaMA but also rivals closed models such as Sonnet 3.5 and ChatGPT-4o.
    The Power of Open Source

    One of the most remarkable aspects of DeepSeek-V3 is its commitment to accessibility. By being open-source, it allows researchers and developers from around the globe to experiment, innovate, and contribute to the AI community. While it does require around 400GB of RAM to run locally, that likely won’t deter anyone from deploying it on their server. The model's documentation and training frameworks are readily available on platforms like Hugging Face, fostering collaboration and knowledge sharing. This democratization of AI technology is a significant step forward, enabling a diverse range of applications from education to programming.
    Technical Innovations

    DeepSeek-V3 is not just about size; it’s about performance. With a processing speed of 60 tokens per second—three times faster than its predecessor—this model is designed for efficiency. The incorporation of FP8 mixed precision training reduces GPU memory consumption without sacrificing accuracy, while the DualPipe algorithm enhances processing efficiency. These advancements not only improve performance but also keep training costs competitive, making DeepSeek-V3 a viable option for various applications.
    A Comedic Twist: The Identity Crisis

    However, the launch of DeepSeek-V3 has not been without its quirks. In a rather amusing turn of events, the model seems to have developed an identity crisis, often identifying itself as ChatGPT during interactions. This phenomenon has sparked laughter and intrigue within the tech community. As reported by TechCrunch, DeepSeek-V3 claimed to be a version of OpenAI’s GPT-4 model in five out of eight generations during tests. This raises questions about the model's training data and the potential for "hallucinations"—a term used to describe AI's tendency to generate inaccurate or misleading information.

    While some may find this amusing, it highlights a critical issue in the AI field: the challenge of ensuring the integrity and accuracy of training data. As AI models increasingly draw from a web saturated with AI-generated content, the risk of misidentification and misinformation grows. This situation serves as a reminder of the importance of transparency and ethical practices in AI development.
    Looking Ahead

    Despite its comedic missteps, DeepSeek-V3 stands as a testament to the potential of open-source AI. Its impressive performance metrics and commitment to accessibility position it as a benchmark in the industry. As the DeepSeek team continues to innovate—introducing features like “Deep Roles” for customizable AI interactions—the future looks bright for this model.

    In conclusion, DeepSeek-V3 not only represents a significant leap forward in open-source AI technology but also provides a lighthearted reminder of the complexities and quirks inherent in artificial intelligence. As we navigate this exciting frontier, let’s embrace both the advancements and the occasional hilarity that comes with it. Whether you’re a developer looking to explore new possibilities or simply someone who enjoys a good laugh at AI’s expense, DeepSeek-V3 is worth keeping an eye on.

    Try it out here: DeepSeek-V3 chat.deepseek.com/sign_in

    Explore more on Hugging Face: Hugging Face - DeepSeek-V3 huggingface.co/deepseek-ai/Dee

    #DeepSeekV3 #OpenSourceAI #ArtificialIntelligence #AIInnovation #TechNews #MachineLearning #AICommunity #DeepLearning #AIModels #HuggingFace
    #DeepSeek #DeepSeekV3

  26. analyticsindiamag.com/ai-news-

    Chinese AI research lab backed by High-Flyer Capital Management has released #DeepSeekV3 open-source Mixture-of-Experts model features a total of 671B total parameters, with 37B activated for each token. The model has been trained on 14.8T tokens. DeepSeek has released the model on GitHub and a detailed technical paper outlining its capabilities.

    DeepSeek AI also released the benchmark scores, and it outperformed Meta’s flagship llama3.1 405B parameter model…