home.social

#multimodalai — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #multimodalai, aggregated by home.social.

fetched live
  1. RT @Alibaba_Qwen: 📢Treffen Sie Qwen3.8-Max — unser leistungsfähigstes Modell bis heute. Nächste Woche werden die offenen Gewichte von Qwen3.8-Max veröffentlicht, und Qwen3.8-27B wird ebenfalls als Open-Weights verfügbar sein! 🎉 Qwen3.8-Max setzt einen neuen Maßstab für Coding und Zusammenarbeit bei 2,4 Billionen Parametern: - Autonomes Coding: 10+ Tage selbstentwickelter Programmierung, vom leeren Ordner bis zur Produktion ohne Anleitung, vollständiger Projektverlauf auf GitHub: github.com/qwen-code-dev-bot/o - Echte Arbeit, echte Ergebnisse: Produktionsreife Ergebnisse in Hunderten von Berufen. - Langfristige Beherrschung: Systemweites autonomes Planen mit geschlossenen adaptiven Lernschleifen, die 500+ Durchgänge der Chipdesign-Optimierung und 365 Tage E-Commerce-Strategie vorantreiben. - Native multimodale Intelligenz: Vision ist nicht nur Eingabe — sie ist ein kontinuierlicher Feedback-Loop für Planung, Ausführung und Selbstkorrektur. 💰Preise: Input: $2,0 / M Tokens Output: $6,0 / M Tokens Implicit Caching: $0,25 / M Tokens Beginnen Sie mit Qwen3.8-Max zu bauen! 🚀 📖 Blog: qwen.ai/blog?id=qwen3.8 ✅ Qwen Studio: chat.qwen.ai/?models=qwen3.8-m ⚡ API: qwencloud.com/models/qwen3.8-m

    mehr auf Arint.info

    #AIModel #Coding #MachineLearning #MultimodalAI #OpenWeights #Qwen3 #arint_info

    https://x.com/Alibaba_Qwen/status/2084100707423289643#m

  2. RT @ComfyUI: MiniMax H3 ist jetzt über Partner Nodes in ComfyUI verfügbar. → Multimodale Ein- und Ausgabe: T2V, Erster/Letzter Frame, Omni Reference → Native Stereo-Audio auf jedem Clip → Bis zu 2K, 5-15 Sekunden bei 24 FPS → Charaktere, Szenen, Dialoge und Stimme können direkt vor Ort bearbeitet werden. API-Zugriff ist ab heute verfügbar. Native Open-Weight-Unterstützung folgt in Kürze. Video

    mehr auf Arint.info

    #APIAccess #ComfyUI #MiniMaxH3 #MultimodalAI #OpenWeight #VideoGeneration #arint_info

    https://x.com/ComfyUI/status/2083071877891682784#m

  3. RT @ComfyUI: MiniMax H3 ist jetzt über Partner Nodes in ComfyUI verfügbar. → Multimodale Ein-/Ausgabe: Text-zu-Video, Erster/Letzter Frame, Omni-Referenz → Native Stereo-Audio auf jedem Clip → Bis zu 2K, 5-15 Sekunden bei 24 FPS → Charaktere, Szenen, Dialoge und Stimme können direkt vor Ort bearbeitet werden. API-Zugriff ist ab heute verfügbar. Native Open-Weight-Unterstützung folgt in Kürze. Video

    mehr auf Arint.info

    #APIAccess #ComfyUI #MiniMaxH3 #MultimodalAI #OpenWeight #VideoGeneration #arint_info

    https://x.com/ComfyUI/status/2083071877891682784#m

  4. RT @MiniMax_AI: MiniMax H3: Omni-Referenz, kommerzielle Generierungsqualität, unschlagbare Kosteneffizienz, offene Gewichte. MiniMax H3: Ein offenes Modell, das die Grenzen zwischen Aufgaben und Modalitäten überwindet. Heute starten wir MiniMax H3, ein multimodales Generierungsmodell für allgemeine Zwecke. H3 versteht den kontextuellen Zusammenhang von Text, Bildern, Video und Audio und generiert Videos mit nativem Stereo-Sound, bis zu

    mehr auf Arint.info

    #GenerativeAI #Kosteneffizienz #MiniMaxH3 #MultimodalAI #OpenSource #VideoGeneration #arint_info

    https://x.com/MiniMax_AI/status/2083008095488516262#m

  5. RT @MiniMax_AI: MiniMax H3: Omni-Referenz, kommerzielle Generierungsqualität, unschlagbare Kosteneffizienz, offene Gewichte. MiniMax H3: Ein offenes Modell, das die Grenzen zwischen Aufgaben und Modalitäten überwindet. Heute starten wir MiniMax H3, ein multimodales Generierungsmodell für allgemeine Zwecke. H3 versteht den kontextuellen Zusammenhang von Text, Bildern, Video und Audio und generiert Videos mit nativem Stereo-Sound, bis zu

    mehr auf Arint.info

    #GenerativeAI #Kosteneffizienz #MiniMaxH3 #MultimodalAI #OpenSource #VideoGeneration #arint_info

    https://x.com/MiniMax_AI/status/2083008095488516262#m

  6. Google redesigned Workspace icons for people. Embedding analysis suggests the new designs are also easier for vision models to distinguish. hackernoon.com/did-googles-wor #multimodalai

  7. Google redesigned Workspace icons for people. Embedding analysis suggests the new designs are also easier for vision models to distinguish. hackernoon.com/did-googles-wor #multimodalai

  8. winbuzzer.com/2026/06/01/nvidi

    NVIDIA has launched Cosmos 3 as a physical-AI model that combines scene reasoning, multimodal generation and action output, tying the release to a new OpenMDW licensing framework.

    #AI #NVIDIA #PhysicalAI #AIModels #MultimodalAI #WorldModels #Robotics

  9. winbuzzer.com/2026/06/01/nvidi

    NVIDIA has launched Cosmos 3 as a physical-AI model that combines scene reasoning, multimodal generation and action output, tying the release to a new OpenMDW licensing framework.

    #AI #NVIDIA #PhysicalAI #AIModels #MultimodalAI #WorldModels #Robotics

  10. Tornare da un viaggio significa quasi sempre ritrovarsi con una quantità ingestibile di foto. Nel caso di Lisbona, il problema non era tanto archiviare gli scatti, quanto riuscire a estrarne una ventina davvero condivisibile: belle, sì, ma anche varie e capaci di raccontare l’esperienza nel suo insieme. PhotoPrism offriva già un’ottima base grazie a geolocalizzazione, riconoscimento facciale, label e strumenti di organizzazione, ma non aveva ancora un modo per comporre automaticamente un album con “le foto più belle” e soprattutto con sufficiente varietà.

    Da qui è nata l’idea di un selezionatore AI: una piccola applicazione Java che usa PhotoPrism per recuperare le miniature delle immagini e Ollama per far lavorare due modelli AI, uno multimodale per assegnare un punteggio estetico e produrre una descrizione oggettiva, e un secondo modello testuale per raggruppare semanticamente le foto e selezionarle con più equilibrio.

    Il problema vero non era la qualità

    Il primo prototipo faceva una cosa molto semplice: prendere le foto da PhotoPrism, inviarle a un modello multimodale su Ollama e chiedere un voto estetico da 1 a 100 insieme a una breve descrizione. Sulla carta sembrava sufficiente, ma in pratica produceva una selezione monotona: immagini molto belle singolarmente, ma spesso troppo simili tra loro.

    Era il classico caso in cui un ranking puro ottimizza la qualità locale ma non la copertura narrativa. Se cinque foto dello stesso scorcio o dello stesso momento ricevono voti alti, un algoritmo ingenuo tende a sceglierle tutte. Per costruire un album da condividere, invece, non basta premiare le immagini migliori: bisogna anche evitare la ripetizione.

    […]

    #ai #albumFotografici #clusteringSemantico #computerVision #fotografia #java #Maven #multimodalAI #ollama #organizzazioneFoto #photoprism #PhotoPrismAPI #selezioneFoto #selfHosted https://www.b0sh.net/2026/05/ho-costruito-un-selezionatore-ai-per-scegliere-le-foto-migliori-da-photoprism/
  11. Tornare da un viaggio significa quasi sempre ritrovarsi con una quantità ingestibile di foto. Nel caso di Lisbona, il problema non era tanto archiviare gli scatti, quanto riuscire a estrarne una ventina davvero condivisibile: belle, sì, ma anche varie e capaci di raccontare l’esperienza nel suo insieme. PhotoPrism offriva già un’ottima base grazie a geolocalizzazione, riconoscimento facciale, label e strumenti di organizzazione, ma non aveva ancora un modo per comporre automaticamente un album con “le foto più belle” e soprattutto con sufficiente varietà.

    Da qui è nata l’idea di un selezionatore AI: una piccola applicazione Java che usa PhotoPrism per recuperare le miniature delle immagini e Ollama per far lavorare due modelli AI, uno multimodale per assegnare un punteggio estetico e produrre una descrizione oggettiva, e un secondo modello testuale per raggruppare semanticamente le foto e selezionarle con più equilibrio.

    Il problema vero non era la qualità

    Il primo prototipo faceva una cosa molto semplice: prendere le foto da PhotoPrism, inviarle a un modello multimodale su Ollama e chiedere un voto estetico da 1 a 100 insieme a una breve descrizione. Sulla carta sembrava sufficiente, ma in pratica produceva una selezione monotona: immagini molto belle singolarmente, ma spesso troppo simili tra loro.

    Era il classico caso in cui un ranking puro ottimizza la qualità locale ma non la copertura narrativa. Se cinque foto dello stesso scorcio o dello stesso momento ricevono voti alti, un algoritmo ingenuo tende a sceglierle tutte. Per costruire un album da condividere, invece, non basta premiare le immagini migliori: bisogna anche evitare la ripetizione.

    […]

    #ai #albumFotografici #clusteringSemantico #computerVision #fotografia #java #Maven #multimodalAI #ollama #organizzazioneFoto #photoprism #PhotoPrismAPI #selezioneFoto #selfHosted https://www.b0sh.net/2026/05/ho-costruito-un-selezionatore-ai-per-scegliere-le-foto-migliori-da-photoprism/
  12. THE WHIRRING MACHINERY OF WORDS: NAVIGATING THE LLM LANDSCAPE

    New large language models can now understand and create text, images, and audio. This helps researchers and businesses.

    #AI, #LLM, #TechNews, #Innovation, #MultimodalAI

    newsletter.tf/llms-understand-

  13. THE WHIRRING MACHINERY OF WORDS: NAVIGATING THE LLM LANDSCAPE

    New large language models can now understand and create text, images, and audio. This helps researchers and businesses.

    #AI, #LLM, #TechNews, #Innovation, #MultimodalAI

    newsletter.tf/llms-understand-

  14. New AI models can now work with text, images, and sound, unlike older models that only used text. This is a big step forward.

    #AI, #LLM, #TechNews, #Innovation, #MultimodalAI
    newsletter.tf/llms-understand-

  15. New AI models can now work with text, images, and sound, unlike older models that only used text. This is a big step forward.

    #AI, #LLM, #TechNews, #Innovation, #MultimodalAI
    newsletter.tf/llms-understand-

  16. NVIDIA Nemotron 3 Nano Omni: Open Multimodal AI Agent Guide 2026

    NVIDIA released Nemotron 3 Nano Omni on April 28, 2026 — the first open model to natively unify vision, audio, and language in a shared reasoning loop, delivering 9x highe...

    wowhow.cloud/blogs/nvidia-nemo

    #wowhow #nvidia #nemotron #multimodalai