home.social

#gpt4v — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #gpt4v, aggregated by home.social.

fetched live
  1. 🔍 Major breakthrough in multimodal AI research:

    #InfinityMM dataset launches with 43.4M entries across 4 categories: 10M image descriptions, 24.4M visual instructions, 6M high-quality instructions & 3M #AI generated data

    🧠 Technical highlights:

    New #AquilaVL2B model uses #LLaVA architecture with #Qwen25 language model & #SigLIP for image processing
    Despite only 2B parameters, achieves state-of-the-art results in multiple benchmarks
    Exceptional performance: #MMStar (54.9%), #MathVista (59%), #MMBench (75.2%)

    🚀 Training innovation:

    4-stage training process with increasing complexity
    Combines image recognition, instruction classification & response generation
    Uses #opensource models like RAM++ for data generation

    💡 Industry impact:

    Model trained on both #Nvidia A100 GPUs & Chinese chips
    Complete dataset & model available to research community
    Shows promising results compared to commercial systems like #GPT4V

    arxiv.org/abs/2410.18558

  2. 🔍 Major breakthrough in multimodal AI research:

    #InfinityMM dataset launches with 43.4M entries across 4 categories: 10M image descriptions, 24.4M visual instructions, 6M high-quality instructions & 3M #AI generated data

    🧠 Technical highlights:

    New #AquilaVL2B model uses #LLaVA architecture with #Qwen25 language model & #SigLIP for image processing
    Despite only 2B parameters, achieves state-of-the-art results in multiple benchmarks
    Exceptional performance: #MMStar (54.9%), #MathVista (59%), #MMBench (75.2%)

    🚀 Training innovation:

    4-stage training process with increasing complexity
    Combines image recognition, instruction classification & response generation
    Uses #opensource models like RAM++ for data generation

    💡 Industry impact:

    Model trained on both #Nvidia A100 GPUs & Chinese chips
    Complete dataset & model available to research community
    Shows promising results compared to commercial systems like #GPT4V

    arxiv.org/abs/2410.18558

  3. 🔍 #Microsoft introduces #OmniParser, a new screen parsing module for #GUI interactions:
    • Converts UI screenshots into structured elements for improved #AI agent navigation
    • Works with #GPT4V to generate precise actions for interface regions
    • Achieves top performance on #WindowsAgentArena benchmark

    🛠️ Key Components:
    • Specialized datasets for icon detection and description
    • Fine-tuned detection model for identifying actionable regions
    • Captioning model for extracting functional semantics

    📊 Performance Highlights:
    • Outperforms standard #GPT4V on #ScreenSpot benchmarks
    • Compatible with #Phi35V and #Llama32V models
    • Functions across PC and mobile platforms without HTML dependencies

    🔗 Learn more: microsoft.com/en-us/research/a

  4. [Перевод] Картинка стоит 170 токенов: как GPT-4o кодирует изображения?

    Интересный факт : GPT-4o взимает по 170 токенов за обработку каждого тайла 512x512 , используемого в режиме высокого разрешения. При соотношении примерно 0,75 токенов на слово можно предположить, что картинка стоит примерно 227 слов, что всего в четыре раза меньше, чем в поговорке «картинка стоит тысячи слов». (Кроме того, взимается 85 токенов за master thumbnail низкого разрешения каждого изображения, а изображения более высокого разрешения разбиваются на множество таких тайлов 512x512 , но давайте ограничимся одним тайлом высокого разрешения.) Но почему же 170? Необычное число, неправда ли? В своих ценах OpenAI указывает округлённые числа, например, $20 или $0,50, а в своих внутренних размерностях — степени двойки и тройки. Почему же в этом случае выбрано число 170? Числа, которые без объяснений вставляют в кодовую базу, называют в программировании « магическими числами », и 170 кажется очевидным магическим числом. И почему затраты на изображения вообще преобразуются в стоимость в токенах? Если бы это нужно было только для определения цены, то разве не удобнее было бы просто указать цену за тайл? Что, если OpenAI выбрала 170 не в рамках своей запутанной стратегии ценообразования, а потому что это в буквальном смысле так? Что, если тайлы изображений действительно представлены в виде 170 последовательных векторов эмбеддингов? А если это так, то как реализовано?

    habr.com/ru/articles/834548/

    #openai #gpt4 #gpt4o #gpt4v #эмбеддинги

  5. Fun and interesting experiment with #DallE and #GPT4V: Prompt Dall-E for an image, then let GPT-4 Vision describe that image and feed the result back into Dall-E. Example: dalle.party/?party=42riPROf

  6. Pretty wild what someone can put together with tools like #GPT4V, #Whisper and others. This would have been a big company's keynote demo accompanied by gasps and applause not that long ago: old.reddit.com/r/ChatGPT/comme

  7. “David Attenborough” AI clone narrates developer’s life → arstechnica.com/information-te

    > "We observe the sophisticated Homo sapiens engaging in the ritual of hydration."

    Developer Charlie Holtz hat mit seinem Narrator-Code eine ziemlich unterhaltsame Parodie des Naturdokumentarfilmers David Attenborough gebaut: In einem auf Twitter/X veröffentlichten Video – das auch oben in dem verlinkten Artikel als Kopie enthalten ist, falls ihr verständlicherweise… → eay.li/3ox #blog #AI #GPT4V

  8. “David Attenborough” AI clone narrates developer’s life → arstechnica.com/information-te

    > "We observe the sophisticated Homo sapiens engaging in the ritual of hydration."

    Developer Charlie Holtz hat mit seinem Narrator-Code eine ziemlich unterhaltsame Parodie des Naturdokumentarfilmers David Attenborough gebaut: In einem auf Twitter/X veröffentlichten Video – das auch oben in dem verlinkten Artikel als Kopie enthalten ist, falls ihr verständlicherweise… → eay.li/3ox #blog #AI #GPT4V

  9. Oma elämäsi - mutta taustalla sitä selostaa Sir David Attenborough

    Hakkeri yhdisteli nipun tekoälyjä ja nyt hänen elämäänsä selostaa legendaarinen luontodokumenttien ääni, kuvaillen hyvin yksityiskohtaisesti vaikkapa urospuolisen homo sapiens -yksilön vesilasin juontia.

    dawn.fi/uutiset/2023/11/17/dav

    #DavidAttenborough #Attenborough #tekoäly #gpt4v #uutiset #teknologia

  10. Oma elämäsi - mutta taustalla sitä selostaa Sir David Attenborough

    Hakkeri yhdisteli nipun tekoälyjä ja nyt hänen elämäänsä selostaa legendaarinen luontodokumenttien ääni, kuvaillen hyvin yksityiskohtaisesti vaikkapa urospuolisen homo sapiens -yksilön vesilasin juontia.

    dawn.fi/uutiset/2023/11/17/dav

    #DavidAttenborough #Attenborough #tekoäly #gpt4v #uutiset #teknologia

  11. Intrigued by multi-modal (text + vision) models like LLaVa, I tried an experiment to create a browser extension that walks the DOM, finds images without good alt, converts the image to Base64, sends it to LLaVa 7B 1.5 (running in LlamaCPP's server) and injects the rich description back into the image tag's alt. Needs much more work and testing, but amazing what a ~5GB (quantised at 5bit) model can do!

    Edit: now on Github: github.com/daaain/image-alt-te

  12. Heute ab 18 Uhr reden @nSonic und @chrismarquardt über Kamera-Umwandler, Lightroom-KI, Stasi-Fotografie und GPT-4V. Auf Happy Shooting. #hslive #gpt4v #lightroom #stasi #fotografie

    Live hier: youtube.com/live/gq4-ulmld6Q

  13. #OpenAI has unveiled new voice & image features for #ChatGPT!

    Say hello to #GPT4V, the mastermind behind image inputs, and the updated #DALLE model for generating images.

    More details on #InfoQ: bit.ly/3F6jDAl

    #GenerativeAI

  14. has unveiled new voice & image features for !

    Say hello to , the mastermind behind image inputs, and the updated model for generating images.

    More details on : bit.ly/3F6jDAl