home.social

#qwen2 — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #qwen2, aggregated by home.social.

fetched live
  1. The #OpenCode is working. Check when the Nvidia is used for it
    ⎿  [Image #1]

    ● I'll check how #Ollama uses the #Nvidia GPU.

    ● Bash(nvidia-smi --query-gpu=name,memory.used,memory.total,utilization.gpu --format=csv,noheader 2>/dev/null)
    ⎿  NVIDIA GeForce RTX 2060, 952 MiB, 6144 MiB, 9 %

    ● Bash(ollama ps 2>/dev/null)
    ⎿  NAME ID SIZE PROCESSOR UNTIL
    qwen2.5-coder:7b-16k 0533dd01f237 7.2 GB 30%/70% CPU/GPU 4 minutes from now

    ● The model is currently loaded and actively using the GPU. Here's the breakdown:

    - Model: #qwen2.5-coder:7b-16k — 7.2 GB total
    - Split: 30% CPU / 70% GPU — the model is split across RAM and VRAM because it doesn't fit entirely in the RTX 2060's 6 GB VRAM
    - Current #GPU #VRAM usage: 952 MiB used of 6144 MiB (the rest of the model layers sit in system RAM)
    - GPU compute utilization: 9% right now (idle between requests)

    The GPU is used whenever OpenCode sends a prompt to Ollama — during inference, 70% of the model layers run on the #RTX2060 and 30% on the #CPU. The model stays loaded in VRAM for 4 more minutes of idle time before Ollama unloads it.

    #LocalLLM

  2. The #OpenCode is working. Check when the Nvidia is used for it
    ⎿  [Image #1]

    ● I'll check how #Ollama uses the #Nvidia GPU.

    ● Bash(nvidia-smi --query-gpu=name,memory.used,memory.total,utilization.gpu --format=csv,noheader 2>/dev/null)
    ⎿  NVIDIA GeForce RTX 2060, 952 MiB, 6144 MiB, 9 %

    ● Bash(ollama ps 2>/dev/null)
    ⎿  NAME ID SIZE PROCESSOR UNTIL
    qwen2.5-coder:7b-16k 0533dd01f237 7.2 GB 30%/70% CPU/GPU 4 minutes from now

    ● The model is currently loaded and actively using the GPU. Here's the breakdown:

    - Model: #qwen2.5-coder:7b-16k — 7.2 GB total
    - Split: 30% CPU / 70% GPU — the model is split across RAM and VRAM because it doesn't fit entirely in the RTX 2060's 6 GB VRAM
    - Current #GPU #VRAM usage: 952 MiB used of 6144 MiB (the rest of the model layers sit in system RAM)
    - GPU compute utilization: 9% right now (idle between requests)

    The GPU is used whenever OpenCode sends a prompt to Ollama — during inference, 70% of the model layers run on the #RTX2060 and 30% on the #CPU. The model stays loaded in VRAM for 4 more minutes of idle time before Ollama unloads it.

    #LocalLLM

  3. The #OpenCode is working. Check when the Nvidia is used for it
    ⎿  [Image #1]

    ● I'll check how #Ollama uses the #Nvidia GPU.

    ● Bash(nvidia-smi --query-gpu=name,memory.used,memory.total,utilization.gpu --format=csv,noheader 2>/dev/null)
    ⎿  NVIDIA GeForce RTX 2060, 952 MiB, 6144 MiB, 9 %

    ● Bash(ollama ps 2>/dev/null)
    ⎿  NAME ID SIZE PROCESSOR UNTIL
    qwen2.5-coder:7b-16k 0533dd01f237 7.2 GB 30%/70% CPU/GPU 4 minutes from now

    ● The model is currently loaded and actively using the GPU. Here's the breakdown:

    - Model: #qwen2.5-coder:7b-16k — 7.2 GB total
    - Split: 30% CPU / 70% GPU — the model is split across RAM and VRAM because it doesn't fit entirely in the RTX 2060's 6 GB VRAM
    - Current #GPU #VRAM usage: 952 MiB used of 6144 MiB (the rest of the model layers sit in system RAM)
    - GPU compute utilization: 9% right now (idle between requests)

    The GPU is used whenever OpenCode sends a prompt to Ollama — during inference, 70% of the model layers run on the #RTX2060 and 30% on the #CPU. The model stays loaded in VRAM for 4 more minutes of idle time before Ollama unloads it.

    #LocalLLM

  4. The #OpenCode is working. Check when the Nvidia is used for it
    ⎿  [Image #1]

    ● I'll check how #Ollama uses the #Nvidia GPU.

    ● Bash(nvidia-smi --query-gpu=name,memory.used,memory.total,utilization.gpu --format=csv,noheader 2>/dev/null)
    ⎿  NVIDIA GeForce RTX 2060, 952 MiB, 6144 MiB, 9 %

    ● Bash(ollama ps 2>/dev/null)
    ⎿  NAME ID SIZE PROCESSOR UNTIL
    qwen2.5-coder:7b-16k 0533dd01f237 7.2 GB 30%/70% CPU/GPU 4 minutes from now

    ● The model is currently loaded and actively using the GPU. Here's the breakdown:

    - Model: #qwen2.5-coder:7b-16k — 7.2 GB total
    - Split: 30% CPU / 70% GPU — the model is split across RAM and VRAM because it doesn't fit entirely in the RTX 2060's 6 GB VRAM
    - Current #GPU #VRAM usage: 952 MiB used of 6144 MiB (the rest of the model layers sit in system RAM)
    - GPU compute utilization: 9% right now (idle between requests)

    The GPU is used whenever OpenCode sends a prompt to Ollama — during inference, 70% of the model layers run on the #RTX2060 and 30% on the #CPU. The model stays loaded in VRAM for 4 more minutes of idle time before Ollama unloads it.

    #LocalLLM

  5. RT @HuggingModels: Lernen Sie Qwen2-32B-N64-Decomp kennen, eine leistungsstarke konversationelle KI, die ab sofort im GGUF-Format verfügbar ist. Dieses Modell bringt Dialogfunktionen auf Enterprise-Niveau auf lokale Maschinen und ermöglicht es Ihnen, anspruchsvolle KI-Chats ohne Cloud-Abhängigkeiten zu führen. Perfekt für Entwickler, die volle Kontrolle wünschen.

    mehr auf Arint.info

    #AI #GGUF #LLM #LocalAI #MachineLearning #Qwen2 #arint_info

    https://x.com/HuggingModels/status/2043963227521069367#m

  6. BTW, these are the #AI #LLM models I settled on using with #JanAI:

    #Qwen2.5 at 0.5B (Qwen2_5-0_5B-Instruct-uncensored_Q8_0), for fastest performance on low-end hardware

    #Qwen2 at 1.5B (Qwen2-1_5B-Instruct-Abliterated-Q5_K_M), for balanced performance and good enough output quality

    #Llama3.2 at 3B (Llama-3_2-3B-Instruct-heretic-ablitered-uncensored_Q5_K_M), for higher quality output

    #Llama3 actually doesn’t run too poorly on my machine, although it can take some time to load up responses sometimes.

  7. BTW, these are the #AI #LLM models I settled on using with #JanAI:

    #Qwen2.5 at 0.5B (Qwen2_5-0_5B-Instruct-uncensored_Q8_0), for fastest performance on low-end hardware

    #Qwen2 at 1.5B (Qwen2-1_5B-Instruct-Abliterated-Q5_K_M), for balanced performance and good enough output quality

    #Llama3.2 at 3B (Llama-3_2-3B-Instruct-heretic-ablitered-uncensored_Q5_K_M), for higher quality output

    #Llama3 actually doesn’t run too poorly on my machine, although it can take some time to load up responses sometimes.

  8. Vấn đề với hệ thống chat RAG: Qwen2.5 bỏ qua ngữ cảnh cuộc trò chuyện trước và trả lời không liên quan cho các câu hỏi tiếp theo. Người dùng gặp khó khăn khi mô hình chỉ dựa vào truy vấn mới nhất thay vì sử dụng lịch sử chat.

    #RAG #AI #Qwen2.5 #Chatbot #LỗiKỹThuật

    reddit.com/r/ollama/comments/1

  9. Hướng dẫn tinh chỉnh mô hình Qwen2.5-Coder-1.5B cho phân tích cảm xúc tiếng Trung. Có thể chạy trên Google Colab miễn phí trong 20-30 phút. Độ chính xác tăng từ 91,6% lên 97,8%. #AI #MachineLearning #Qwen2.5 #PhânTíchCảmXúc #GoogleColab #TinhChỉnhMôHình #TríTuệNhânTạo #HọcMáy

    i.redd.it/7xx856mftfzf1.png

  10. Ch peque nhá! Tôi vừa chuyển sang dùng Qwen2.5 Code Instruct bản tự-host thành công! M المقابل với Claude đầu tiên (lần nào 1h phải chờ), Qwen2.5 có thể xử lý comuni code, debug, và nhiếp ý nhanh lùi ởстром đường công việc. Ưbrochen ở máy MBook Pro 48GB và PC 2x RTX 5060TI 16GB (không cần quantize). Cài đặt đơn giản, chất lượng tốt cho công việc lẻ lậu.
    Tham khảo GitHub: @reliableJARED/qwen_coder
    Tags: #AI #Qwen2.5 #CodeAssistant #LocalTech #MáyTínhLâu
    #TechTips #OfflineAI #DevelopersCommu

  11. 🧠 #ByteDance ha rilasciato UI-TARS-1.5, un agente multimodale basato su #Qwen2.5-VL-7B che unisce visione e linguaggio con "reasoning". 

    👉 I dettagli: linkedin.com/posts/alessiopoma

    ___ 

    ✉️ 𝗦𝗲 𝘃𝘂𝗼𝗶 𝗿𝗶𝗺𝗮𝗻𝗲𝗿𝗲 𝗮𝗴𝗴𝗶𝗼𝗿𝗻𝗮𝘁𝗼/𝗮 𝘀𝘂 𝗾𝘂𝗲𝘀𝘁𝗲 𝘁𝗲𝗺𝗮𝘁𝗶𝗰𝗵𝗲, 𝗶𝘀𝗰𝗿𝗶𝘃𝗶𝘁𝗶 𝗮𝗹𝗹𝗮 𝗺𝗶𝗮 𝗻𝗲𝘄𝘀𝗹𝗲𝘁𝘁𝗲𝗿: bit.ly/newsletter-alessiopomar 

    #AI #GenAI #GenerativeAI #IntelligenzaArtificiale #LLM 

  12. 🧠 #ByteDance ha rilasciato UI-TARS-1.5, un agente multimodale basato su #Qwen2.5-VL-7B che unisce visione e linguaggio con "reasoning". 

    👉 I dettagli: linkedin.com/posts/alessiopoma

    ___ 

    ✉️ 𝗦𝗲 𝘃𝘂𝗼𝗶 𝗿𝗶𝗺𝗮𝗻𝗲𝗿𝗲 𝗮𝗴𝗴𝗶𝗼𝗿𝗻𝗮𝘁𝗼/𝗮 𝘀𝘂 𝗾𝘂𝗲𝘀𝘁𝗲 𝘁𝗲𝗺𝗮𝘁𝗶𝗰𝗵𝗲, 𝗶𝘀𝗰𝗿𝗶𝘃𝗶𝘁𝗶 𝗮𝗹𝗹𝗮 𝗺𝗶𝗮 𝗻𝗲𝘄𝘀𝗹𝗲𝘁𝘁𝗲𝗿: bit.ly/newsletter-alessiopomar 

    #AI #GenAI #GenerativeAI #IntelligenzaArtificiale #LLM 

  13. 🧠 #ByteDance ha rilasciato UI-TARS-1.5, un agente multimodale basato su #Qwen2.5-VL-7B che unisce visione e linguaggio con "reasoning". 

    👉 I dettagli: linkedin.com/posts/alessiopoma

    ___ 

    ✉️ 𝗦𝗲 𝘃𝘂𝗼𝗶 𝗿𝗶𝗺𝗮𝗻𝗲𝗿𝗲 𝗮𝗴𝗴𝗶𝗼𝗿𝗻𝗮𝘁𝗼/𝗮 𝘀𝘂 𝗾𝘂𝗲𝘀𝘁𝗲 𝘁𝗲𝗺𝗮𝘁𝗶𝗰𝗵𝗲, 𝗶𝘀𝗰𝗿𝗶𝘃𝗶𝘁𝗶 𝗮𝗹𝗹𝗮 𝗺𝗶𝗮 𝗻𝗲𝘄𝘀𝗹𝗲𝘁𝘁𝗲𝗿: bit.ly/newsletter-alessiopomar 

    #AI #GenAI #GenerativeAI #IntelligenzaArtificiale #LLM 

  14. 🧠 #ByteDance ha rilasciato UI-TARS-1.5, un agente multimodale basato su #Qwen2.5-VL-7B che unisce visione e linguaggio con "reasoning". 

    👉 I dettagli: linkedin.com/posts/alessiopoma

    ___ 

    ✉️ 𝗦𝗲 𝘃𝘂𝗼𝗶 𝗿𝗶𝗺𝗮𝗻𝗲𝗿𝗲 𝗮𝗴𝗴𝗶𝗼𝗿𝗻𝗮𝘁𝗼/𝗮 𝘀𝘂 𝗾𝘂𝗲𝘀𝘁𝗲 𝘁𝗲𝗺𝗮𝘁𝗶𝗰𝗵𝗲, 𝗶𝘀𝗰𝗿𝗶𝘃𝗶𝘁𝗶 𝗮𝗹𝗹𝗮 𝗺𝗶𝗮 𝗻𝗲𝘄𝘀𝗹𝗲𝘁𝘁𝗲𝗿: bit.ly/newsletter-alessiopomar 

    #AI #GenAI #GenerativeAI #IntelligenzaArtificiale #LLM 

  15. ​Qwen2.5-VL & QVQ-Max: Neue Maßstäbe in der visuellen KI

    Fortschrittliche Bild- und Videoanalyse
    Präzise Objekterkennung
    Verbesserte Dokumentenverarbeitung

    #ai #ki #artificialintelligence #kuenstlicheintelligenz #Qwen2.5-VL #QVQ-Max

    Jetzt lesen und folgen!

    kinews24.de/qwen2-5-vl-qvq-max/

  16. ​Qwen2.5-VL & QVQ-Max: Neue Maßstäbe in der visuellen KI

    Fortschrittliche Bild- und Videoanalyse
    Präzise Objekterkennung
    Verbesserte Dokumentenverarbeitung

    #ai #ki #artificialintelligence #kuenstlicheintelligenz #Qwen2.5-VL #QVQ-Max

    Jetzt lesen und folgen!

    kinews24.de/qwen2-5-vl-qvq-max/

  17. ​Qwen2.5-VL & QVQ-Max: Neue Maßstäbe in der visuellen KI

    Fortschrittliche Bild- und Videoanalyse
    Präzise Objekterkennung
    Verbesserte Dokumentenverarbeitung

    #ai #ki #artificialintelligence #kuenstlicheintelligenz #Qwen2.5-VL #QVQ-Max

    Jetzt lesen und folgen!

    kinews24.de/qwen2-5-vl-qvq-max/

  18. ​Qwen2.5-VL & QVQ-Max: Neue Maßstäbe in der visuellen KI

    Fortschrittliche Bild- und Videoanalyse
    Präzise Objekterkennung
    Verbesserte Dokumentenverarbeitung

    #ai #ki #artificialintelligence #kuenstlicheintelligenz #Qwen2.5-VL #QVQ-Max

    Jetzt lesen und folgen!

    kinews24.de/qwen2-5-vl-qvq-max/

  19. Alibaba Cloud shakes up the AI scene with **Qwen2.5-Omni-7B!** This cutting-edge multimodal model processes text, images, audio, and video, making it perfect for mobile devices. It's designed for cost-effective AI agents, especially in voice applications for the visually impaired. With a hefty **$53 billion** investment in AI and cloud infrastructure, Alibaba is positioning itself for success in the booming AI market—don’t miss the full story. [Read more](cnbc.com/2025/03/27/alibaba-la) #ArtificialIntelligence #AlibabaCloud #Qwen2 #TechInnovation

  20. Alibaba Cloud shakes up the AI scene with **Qwen2.5-Omni-7B!** This cutting-edge multimodal model processes text, images, audio, and video, making it perfect for mobile devices. It's designed for cost-effective AI agents, especially in voice applications for the visually impaired. With a hefty **$53 billion** investment in AI and cloud infrastructure, Alibaba is positioning itself for success in the booming AI market—don’t miss the full story. [Read more](cnbc.com/2025/03/27/alibaba-la) #ArtificialIntelligence #AlibabaCloud #Qwen2 #TechInnovation

  21. Alibaba Cloud shakes up the AI scene with **Qwen2.5-Omni-7B!** This cutting-edge multimodal model processes text, images, audio, and video, making it perfect for mobile devices. It's designed for cost-effective AI agents, especially in voice applications for the visually impaired. With a hefty **$53 billion** investment in AI and cloud infrastructure, Alibaba is positioning itself for success in the booming AI market—don’t miss the full story. [Read more](cnbc.com/2025/03/27/alibaba-la) #ArtificialIntelligence #AlibabaCloud #Qwen2 #TechInnovation

  22. Qwen2.5-VL-32B: because nothing says "cutting-edge" like moaning about parameter scales and reinforcement learning 🙄. Apparently, this 32B thing is "smarter" and "lighter" – sounds like a diet ad for AI models. 😂🍩 #Innovation!
    qwenlm.github.io/blog/qwen2.5- #Qwen2.5VL32B #AIModels #ReinforcementLearning #CuttingEdge #TechHumor #HackerNews #ngated

  23. Qwen2.5-VL-32B: because nothing says "cutting-edge" like moaning about parameter scales and reinforcement learning 🙄. Apparently, this 32B thing is "smarter" and "lighter" – sounds like a diet ad for AI models. 😂🍩 #Innovation!
    qwenlm.github.io/blog/qwen2.5- #Qwen2.5VL32B #AIModels #ReinforcementLearning #CuttingEdge #TechHumor #HackerNews #ngated

  24. Qwen2.5-VL-32B: because nothing says "cutting-edge" like moaning about parameter scales and reinforcement learning 🙄. Apparently, this 32B thing is "smarter" and "lighter" – sounds like a diet ad for AI models. 😂🍩 #Innovation!
    qwenlm.github.io/blog/qwen2.5- #Qwen2.5VL32B #AIModels #ReinforcementLearning #CuttingEdge #TechHumor #HackerNews #ngated

  25. Qwen2.5-VL-32B: because nothing says "cutting-edge" like moaning about parameter scales and reinforcement learning 🙄. Apparently, this 32B thing is "smarter" and "lighter" – sounds like a diet ad for AI models. 😂🍩 #Innovation!
    qwenlm.github.io/blog/qwen2.5- #Qwen2.5VL32B #AIModels #ReinforcementLearning #CuttingEdge #TechHumor #HackerNews #ngated

  26. OLMo 2 32B offers unprecedented transparency in #LLM development:

    • 🚀 State-of-the-art results: Outperforms GPT3.5, GPT4o-mini, matches top open-weight models like #Qwen2.5 and approaches #Llama3

  27. OLMo 2 32B offers unprecedented transparency in #LLM development:

    • 🚀 State-of-the-art results: Outperforms GPT3.5, GPT4o-mini, matches top open-weight models like #Qwen2.5 and approaches #Llama3

  28. OLMo 2 32B offers unprecedented transparency in #LLM development:

    • 🚀 State-of-the-art results: Outperforms GPT3.5, GPT4o-mini, matches top open-weight models like #Qwen2.5 and approaches #Llama3

  29. OLMo 2 32B offers unprecedented transparency in #LLM development:

    • 🚀 State-of-the-art results: Outperforms GPT3.5, GPT4o-mini, matches top open-weight models like #Qwen2.5 and approaches #Llama3

  30. #AI2 releases OLMo 2 32B, trained on 6T tokens with #Tulu3.1 post-training. Matches or exceeds GPT3.5 Turbo while using just 1/3 the compute of #Qwen2.5 32B. Complete open recipe includes data, code, weights and training methodology.

  31. #AI2 releases OLMo 2 32B, trained on 6T tokens with #Tulu3.1 post-training. Matches or exceeds GPT3.5 Turbo while using just 1/3 the compute of #Qwen2.5 32B. Complete open recipe includes data, code, weights and training methodology.

  32. #AI2 releases OLMo 2 32B, trained on 6T tokens with #Tulu3.1 post-training. Matches or exceeds GPT3.5 Turbo while using just 1/3 the compute of #Qwen2.5 32B. Complete open recipe includes data, code, weights and training methodology.

  33. #AI2 releases OLMo 2 32B, trained on 6T tokens with #Tulu3.1 post-training. Matches or exceeds GPT3.5 Turbo while using just 1/3 the compute of #Qwen2.5 32B. Complete open recipe includes data, code, weights and training methodology.

  34. 🎯 #OpenSource Language Model Platform Launch

    🔧 Leverages #vLLM technology with custom #GPU scheduler for running various #LLM models
    🤖 Supports major models: #Llama3 (405B/70B/8B), #Qwen2 72B, #Mixtral, #Gemma2, #Jamba15, #Phi3

    glhf.chat/

  35. 🎯 #OpenSource Language Model Platform Launch

    🔧 Leverages #vLLM technology with custom #GPU scheduler for running various #LLM models
    🤖 Supports major models: #Llama3 (405B/70B/8B), #Qwen2 72B, #Mixtral, #Gemma2, #Jamba15, #Phi3

    glhf.chat/

  36. 🎯 #OpenSource Language Model Platform Launch

    🔧 Leverages #vLLM technology with custom #GPU scheduler for running various #LLM models
    🤖 Supports major models: #Llama3 (405B/70B/8B), #Qwen2 72B, #Mixtral, #Gemma2, #Jamba15, #Phi3

    glhf.chat/

  37. 🎯 #OpenSource Language Model Platform Launch

    🔧 Leverages #vLLM technology with custom #GPU scheduler for running various #LLM models
    🤖 Supports major models: #Llama3 (405B/70B/8B), #Qwen2 72B, #Mixtral, #Gemma2, #Jamba15, #Phi3

    glhf.chat/

  38. Advanced Reasoning Model: #ai #llm Marco-o1 Pushes Boundaries in Problem-Solving 🧠

    🔬 Built on #Qwen2, focusing on open-ended reasoning beyond traditional tasks

    💡 Key Innovations:
    🤔 #ChainOfThought fine-tuning for structured reasoning
    🌳 Monte Carlo Tree Search (#MCTS) for solution space exploration
    🔄 Novel reflection mechanisms for self-improvement
    🎯 Multiple action granularities for complex problem-solving

    📊 Performance Highlights:
    📈 +6.17% accuracy on MGSM English dataset
    📈 +5.60% accuracy on MGSM Chinese dataset
    🌐 Excels in translation tasks, especially with colloquial expressions

    🛠️ Technical Features:
    • Fine-tuned on 60,266 training samples
    • Implements step & mini-step MCTS strategies
    • Utilizes confidence scoring for path selection
    • Incorporates self-reflection mechanisms

    ⚡️ Project Status: Research work in progress with continuous optimization
    github.com/AIDC-AI/Marco-o1

  39. Advanced Reasoning Model: #ai #llm Marco-o1 Pushes Boundaries in Problem-Solving 🧠

    🔬 Built on #Qwen2, focusing on open-ended reasoning beyond traditional tasks

    💡 Key Innovations:
    🤔 #ChainOfThought fine-tuning for structured reasoning
    🌳 Monte Carlo Tree Search (#MCTS) for solution space exploration
    🔄 Novel reflection mechanisms for self-improvement
    🎯 Multiple action granularities for complex problem-solving

    📊 Performance Highlights:
    📈 +6.17% accuracy on MGSM English dataset
    📈 +5.60% accuracy on MGSM Chinese dataset
    🌐 Excels in translation tasks, especially with colloquial expressions

    🛠️ Technical Features:
    • Fine-tuned on 60,266 training samples
    • Implements step & mini-step MCTS strategies
    • Utilizes confidence scoring for path selection
    • Incorporates self-reflection mechanisms

    ⚡️ Project Status: Research work in progress with continuous optimization
    github.com/AIDC-AI/Marco-o1

  40. Advanced Reasoning Model: #ai #llm Marco-o1 Pushes Boundaries in Problem-Solving 🧠

    🔬 Built on #Qwen2, focusing on open-ended reasoning beyond traditional tasks

    💡 Key Innovations:
    🤔 #ChainOfThought fine-tuning for structured reasoning
    🌳 Monte Carlo Tree Search (#MCTS) for solution space exploration
    🔄 Novel reflection mechanisms for self-improvement
    🎯 Multiple action granularities for complex problem-solving

    📊 Performance Highlights:
    📈 +6.17% accuracy on MGSM English dataset
    📈 +5.60% accuracy on MGSM Chinese dataset
    🌐 Excels in translation tasks, especially with colloquial expressions

    🛠️ Technical Features:
    • Fine-tuned on 60,266 training samples
    • Implements step & mini-step MCTS strategies
    • Utilizes confidence scoring for path selection
    • Incorporates self-reflection mechanisms

    ⚡️ Project Status: Research work in progress with continuous optimization
    github.com/AIDC-AI/Marco-o1

  41. Advanced Reasoning Model: #ai #llm Marco-o1 Pushes Boundaries in Problem-Solving 🧠

    🔬 Built on #Qwen2, focusing on open-ended reasoning beyond traditional tasks

    💡 Key Innovations:
    🤔 #ChainOfThought fine-tuning for structured reasoning
    🌳 Monte Carlo Tree Search (#MCTS) for solution space exploration
    🔄 Novel reflection mechanisms for self-improvement
    🎯 Multiple action granularities for complex problem-solving

    📊 Performance Highlights:
    📈 +6.17% accuracy on MGSM English dataset
    📈 +5.60% accuracy on MGSM Chinese dataset
    🌐 Excels in translation tasks, especially with colloquial expressions

    🛠️ Technical Features:
    • Fine-tuned on 60,266 training samples
    • Implements step & mini-step MCTS strategies
    • Utilizes confidence scoring for path selection
    • Incorporates self-reflection mechanisms

    ⚡️ Project Status: Research work in progress with continuous optimization
    github.com/AIDC-AI/Marco-o1

  42. Edge-Ready #Vision Language Model Advances Visual #AI Processing 🌟

    🧠 #OmniVision (968M params) sets new benchmark as world's smallest #VisionLanguageModel

    🔄 Architecture combines #Qwen2 (0.5B) for text & #SigLIP (400M) for vision processing

    💡 Key Innovations:
    • 9x token reduction (729 → 81) for faster processing
    • Enhanced accuracy through #DPO training
    • Only 988MB RAM & 948MB storage required
    • Outperforms #nanoLLAVA across multiple benchmarks

    🎯 Use Cases:
    • Image analysis & description
    • Visual memory assistance
    • Recipe generation from food images
    • Technical documentation support

    Try it now: huggingface.co/spaces/NexaAIDe
    Source: nexa.ai/blogs/omni-vision

  43. Edge-Ready #Vision Language Model Advances Visual #AI Processing 🌟

    🧠 #OmniVision (968M params) sets new benchmark as world's smallest #VisionLanguageModel

    🔄 Architecture combines #Qwen2 (0.5B) for text & #SigLIP (400M) for vision processing

    💡 Key Innovations:
    • 9x token reduction (729 → 81) for faster processing
    • Enhanced accuracy through #DPO training
    • Only 988MB RAM & 948MB storage required
    • Outperforms #nanoLLAVA across multiple benchmarks

    🎯 Use Cases:
    • Image analysis & description
    • Visual memory assistance
    • Recipe generation from food images
    • Technical documentation support

    Try it now: huggingface.co/spaces/NexaAIDe
    Source: nexa.ai/blogs/omni-vision

  44. Edge-Ready #Vision Language Model Advances Visual #AI Processing 🌟

    🧠 #OmniVision (968M params) sets new benchmark as world's smallest #VisionLanguageModel

    🔄 Architecture combines #Qwen2 (0.5B) for text & #SigLIP (400M) for vision processing

    💡 Key Innovations:
    • 9x token reduction (729 → 81) for faster processing
    • Enhanced accuracy through #DPO training
    • Only 988MB RAM & 948MB storage required
    • Outperforms #nanoLLAVA across multiple benchmarks

    🎯 Use Cases:
    • Image analysis & description
    • Visual memory assistance
    • Recipe generation from food images
    • Technical documentation support

    Try it now: huggingface.co/spaces/NexaAIDe
    Source: nexa.ai/blogs/omni-vision

  45. Edge-Ready #Vision Language Model Advances Visual #AI Processing 🌟

    🧠 #OmniVision (968M params) sets new benchmark as world's smallest #VisionLanguageModel

    🔄 Architecture combines #Qwen2 (0.5B) for text & #SigLIP (400M) for vision processing

    💡 Key Innovations:
    • 9x token reduction (729 → 81) for faster processing
    • Enhanced accuracy through #DPO training
    • Only 988MB RAM & 948MB storage required
    • Outperforms #nanoLLAVA across multiple benchmarks

    🎯 Use Cases:
    • Image analysis & description
    • Visual memory assistance
    • Recipe generation from food images
    • Technical documentation support

    Try it now: huggingface.co/spaces/NexaAIDe
    Source: nexa.ai/blogs/omni-vision