#gpt4v — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #gpt4v, aggregated by home.social.
-
🔍 Major breakthrough in multimodal AI research:
#InfinityMM dataset launches with 43.4M entries across 4 categories: 10M image descriptions, 24.4M visual instructions, 6M high-quality instructions & 3M #AI generated data
🧠 Technical highlights:
New #AquilaVL2B model uses #LLaVA architecture with #Qwen25 language model & #SigLIP for image processing
Despite only 2B parameters, achieves state-of-the-art results in multiple benchmarks
Exceptional performance: #MMStar (54.9%), #MathVista (59%), #MMBench (75.2%)🚀 Training innovation:
4-stage training process with increasing complexity
Combines image recognition, instruction classification & response generation
Uses #opensource models like RAM++ for data generation💡 Industry impact:
Model trained on both #Nvidia A100 GPUs & Chinese chips
Complete dataset & model available to research community
Shows promising results compared to commercial systems like #GPT4V -
🔍 Major breakthrough in multimodal AI research:
#InfinityMM dataset launches with 43.4M entries across 4 categories: 10M image descriptions, 24.4M visual instructions, 6M high-quality instructions & 3M #AI generated data
🧠 Technical highlights:
New #AquilaVL2B model uses #LLaVA architecture with #Qwen25 language model & #SigLIP for image processing
Despite only 2B parameters, achieves state-of-the-art results in multiple benchmarks
Exceptional performance: #MMStar (54.9%), #MathVista (59%), #MMBench (75.2%)🚀 Training innovation:
4-stage training process with increasing complexity
Combines image recognition, instruction classification & response generation
Uses #opensource models like RAM++ for data generation💡 Industry impact:
Model trained on both #Nvidia A100 GPUs & Chinese chips
Complete dataset & model available to research community
Shows promising results compared to commercial systems like #GPT4V -
🔍 #Microsoft introduces #OmniParser, a new screen parsing module for #GUI interactions:
• Converts UI screenshots into structured elements for improved #AI agent navigation
• Works with #GPT4V to generate precise actions for interface regions
• Achieves top performance on #WindowsAgentArena benchmark🛠️ Key Components:
• Specialized datasets for icon detection and description
• Fine-tuned detection model for identifying actionable regions
• Captioning model for extracting functional semantics📊 Performance Highlights:
• Outperforms standard #GPT4V on #ScreenSpot benchmarks
• Compatible with #Phi35V and #Llama32V models
• Functions across PC and mobile platforms without HTML dependencies🔗 Learn more: https://www.microsoft.com/en-us/research/articles/omniparser-for-pure-vision-based-gui-agent/
-
[Перевод] Картинка стоит 170 токенов: как GPT-4o кодирует изображения?
Интересный факт : GPT-4o взимает по 170 токенов за обработку каждого тайла 512x512 , используемого в режиме высокого разрешения. При соотношении примерно 0,75 токенов на слово можно предположить, что картинка стоит примерно 227 слов, что всего в четыре раза меньше, чем в поговорке «картинка стоит тысячи слов». (Кроме того, взимается 85 токенов за master thumbnail низкого разрешения каждого изображения, а изображения более высокого разрешения разбиваются на множество таких тайлов 512x512 , но давайте ограничимся одним тайлом высокого разрешения.) Но почему же 170? Необычное число, неправда ли? В своих ценах OpenAI указывает округлённые числа, например, $20 или $0,50, а в своих внутренних размерностях — степени двойки и тройки. Почему же в этом случае выбрано число 170? Числа, которые без объяснений вставляют в кодовую базу, называют в программировании « магическими числами », и 170 кажется очевидным магическим числом. И почему затраты на изображения вообще преобразуются в стоимость в токенах? Если бы это нужно было только для определения цены, то разве не удобнее было бы просто указать цену за тайл? Что, если OpenAI выбрала 170 не в рамках своей запутанной стратегии ценообразования, а потому что это в буквальном смысле так? Что, если тайлы изображений действительно представлены в виде 170 последовательных векторов эмбеддингов? А если это так, то как реализовано?
-
Fun and interesting experiment with #DallE and #GPT4V: Prompt Dall-E for an image, then let GPT-4 Vision describe that image and feed the result back into Dall-E. Example: https://dalle.party/?party=42riPROf
-
Pretty wild what someone can put together with tools like #GPT4V, #Whisper and others. This would have been a big company's keynote demo accompanied by gasps and applause not that long ago: https://old.reddit.com/r/ChatGPT/comments/17ywjsv/have_a_live_conversation_about_a_basketball_game/
-
“David Attenborough” AI clone narrates developer’s life → https://arstechnica.com/information-technology/2023/11/unauthorized-david-attenborough-ai-clone-narrates-developers-life-goes-viral/
> "We observe the sophisticated Homo sapiens engaging in the ritual of hydration."
Developer Charlie Holtz hat mit seinem Narrator-Code eine ziemlich unterhaltsame Parodie des Naturdokumentarfilmers David Attenborough gebaut: In einem auf Twitter/X veröffentlichten Video – das auch oben in dem verlinkten Artikel als Kopie enthalten ist, falls ihr verständlicherweise… → https://eay.li/3ox #blog #AI #GPT4V
-
“David Attenborough” AI clone narrates developer’s life → https://arstechnica.com/information-technology/2023/11/unauthorized-david-attenborough-ai-clone-narrates-developers-life-goes-viral/
> "We observe the sophisticated Homo sapiens engaging in the ritual of hydration."
Developer Charlie Holtz hat mit seinem Narrator-Code eine ziemlich unterhaltsame Parodie des Naturdokumentarfilmers David Attenborough gebaut: In einem auf Twitter/X veröffentlichten Video – das auch oben in dem verlinkten Artikel als Kopie enthalten ist, falls ihr verständlicherweise… → https://eay.li/3ox #blog #AI #GPT4V
-
Oma elämäsi - mutta taustalla sitä selostaa Sir David Attenborough
Hakkeri yhdisteli nipun tekoälyjä ja nyt hänen elämäänsä selostaa legendaarinen luontodokumenttien ääni, kuvaillen hyvin yksityiskohtaisesti vaikkapa urospuolisen homo sapiens -yksilön vesilasin juontia.
https://dawn.fi/uutiset/2023/11/17/david-attenborough-selostaa-elama-tekoaly
#DavidAttenborough #Attenborough #tekoäly #gpt4v #uutiset #teknologia
-
Oma elämäsi - mutta taustalla sitä selostaa Sir David Attenborough
Hakkeri yhdisteli nipun tekoälyjä ja nyt hänen elämäänsä selostaa legendaarinen luontodokumenttien ääni, kuvaillen hyvin yksityiskohtaisesti vaikkapa urospuolisen homo sapiens -yksilön vesilasin juontia.
https://dawn.fi/uutiset/2023/11/17/david-attenborough-selostaa-elama-tekoaly
#DavidAttenborough #Attenborough #tekoäly #gpt4v #uutiset #teknologia
-
Intrigued by multi-modal (text + vision) models like LLaVa, I tried an experiment to create a browser extension that walks the DOM, finds images without good alt, converts the image to Base64, sends it to LLaVa 7B 1.5 (running in LlamaCPP's server) and injects the rich description back into the image tag's alt. Needs much more work and testing, but amazing what a ~5GB (quantised at 5bit) model can do!
Edit: now on Github: https://github.com/daaain/image-alt-text-generator-extension
-
Heute ab 18 Uhr reden @nSonic und @chrismarquardt über Kamera-Umwandler, Lightroom-KI, Stasi-Fotografie und GPT-4V. Auf Happy Shooting. #hslive #gpt4v #lightroom #stasi #fotografie
Live hier: https://www.youtube.com/live/gq4-ulmld6Q
-
GPT-4V(ision) – Features und Möglichkeiten (während Europa jetzt darauf wartet…)
#multimodalAI, #GPT4V, #ChatGPT, #KI, #OpenAI, #Bildverarbeitung, #Sprachmodelle, #NLP, #Zukunft, #Innovation
-
[#GPT4V] GPT-4 With Vision: Examples, Limitations, And Potential Risks
https://www.searchenginejournal.com/gpt-4-with-vision-examples-limitations-and-potential-risks/497250/ -
[#GPT4V] GPT-4 With Vision: Examples, Limitations, And Potential Risks
https://www.searchenginejournal.com/gpt-4-with-vision-examples-limitations-and-potential-risks/497250/