home.social

#groq — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #groq, aggregated by home.social.

  1. Google Plans New ‘Frozen’ Chip to Run Its AI Models Much More Efficiently – The Information

    Google Plans New ‘Frozen’ Chip to Run Its AI Models Much More Efficiently  The InformationAlphabet perks up on report…
    #NewsBeep #News #US #USA #UnitedStates #UnitedStatesOfAmerica #Business #AI #AIchips #aiinference #frozenv2 #geminiaimodel #Google #Groq #JeffDean #NVIDIA #OpenAI #SambaNova #Semiconductors #TPU
    newsbeep.com/us/774001/

  2. Google Plans New ‘Frozen’ Chip to Run Its AI Models Much More Efficiently – The Information

    Google Plans New ‘Frozen’ Chip to Run Its AI Models Much More Efficiently  The InformationAlphabet perks up on report…
    #NewsBeep #News #US #USA #UnitedStates #UnitedStatesOfAmerica #Business #AI #AIchips #aiinference #frozenv2 #geminiaimodel #Google #Groq #JeffDean #NVIDIA #OpenAI #SambaNova #Semiconductors #TPU
    newsbeep.com/us/774001/

  3. Samsung Foundry’s 4nm Capacity Sold Out, AI Order Backlog Exceeds 50 Trillion Won — BigGo Finance

    Artificial intelligence (AI) semiconductor demand is reshaping the global foundry landscape at an unprecedented pace. Samsung Electronics’ foundry…
    #EuropeSays #Korea #KR #SamsungElectronics #2nmprocess #4nmprocess #Anthropic #Groq #HBM #Meta #NVIDIA #OpenAI #Samsung #Tesla #TSMC
    europesays.com/korea/74169/

  4. Finally, the #Harness / Orchestrator is taking the form I desire.
    See that list at the bottom? Thats some of the LLM AI engines that are connected to work on a prompt.

    There are two types ORCH (Orchestrator) and WORKER, there is actual 4 of different classes of router asset.

    And for the first time when #Claude hit the weekly limit, the harness did not shut down

    Opencode-big-pickle #Opencode is picking up the Orchestrator role when Claude dies. The bar at the top shows which engines are used most, and surprisingly #Groq (Not Grok) is doing the heaviest lifting for the worker units (WUs)

    Only Claude and Opencode are agentic capable.

    TLDR: Big prompt get crank crank AFK. Ape happy.

    #Agentic #Ai #LLM #Vibecode

  5. Finally, the #Harness / Orchestrator is taking the form I desire.
    See that list at the bottom? Thats some of the LLM AI engines that are connected to work on a prompt.

    There are two types ORCH (Orchestrator) and WORKER, there is actual 4 of different classes of router asset.

    And for the first time when #Claude hit the weekly limit, the harness did not shut down

    Opencode-big-pickle #Opencode is picking up the Orchestrator role when Claude dies. The bar at the top shows which engines are used most, and surprisingly #Groq (Not Grok) is doing the heaviest lifting for the worker units (WUs)

    Only Claude and Opencode are agentic capable.

    TLDR: Big prompt get crank crank AFK. Ape happy.

    #Agentic #Ai #LLM #Vibecode

  6. Finally, the #Harness / Orchestrator is taking the form I desire.
    See that list at the bottom? Thats some of the LLM AI engines that are connected to work on a prompt.

    There are two types ORCH (Orchestrator) and WORKER, there is actual 4 of different classes of router asset.

    And for the first time when #Claude hit the weekly limit, the harness did not shut down

    Opencode-big-pickle #Opencode is picking up the Orchestrator role when Claude dies. The bar at the top shows which engines are used most, and surprisingly #Groq (Not Grok) is doing the heaviest lifting for the worker units (WUs)

    Only Claude and Opencode are agentic capable.

    TLDR: Big prompt get crank crank AFK. Ape happy.

    #Agentic #Ai #LLM #Vibecode

  7. Day 2 of rearchitecting the #Ai harness.
    Multiple provider #LLM hooked up
    I have also installed 3 local LLMs which will run like absolute dogshit on a 4GB VPS... but we will see if its worthwhile... because we can always pump up the server. Some #Hosting providers even offer GPUs

    Ive also added a bar graph (I love graphs) up the top that shows frequency of models used. Not surprisingly #Claude is up the top... but happy to see #groq (Not the fascist grok) up there too.

    Now that we can spare the compute for the organ monkey let it run for a few days and see.

    #vibecoding #Harnessengineering

  8. Day 2 of rearchitecting the #Ai harness.
    Multiple provider #LLM hooked up
    I have also installed 3 local LLMs which will run like absolute dogshit on a 4GB VPS... but we will see if its worthwhile... because we can always pump up the server. Some #Hosting providers even offer GPUs

    Ive also added a bar graph (I love graphs) up the top that shows frequency of models used. Not surprisingly #Claude is up the top... but happy to see #groq (Not the fascist grok) up there too.

    Now that we can spare the compute for the organ monkey let it run for a few days and see.

    #vibecoding #Harnessengineering

  9. Day 2 of rearchitecting the #Ai harness.
    Multiple provider #LLM hooked up
    I have also installed 3 local LLMs which will run like absolute dogshit on a 4GB VPS... but we will see if its worthwhile... because we can always pump up the server. Some #Hosting providers even offer GPUs

    Ive also added a bar graph (I love graphs) up the top that shows frequency of models used. Not surprisingly #Claude is up the top... but happy to see #groq (Not the fascist grok) up there too.

    Now that we can spare the compute for the organ monkey let it run for a few days and see.

    #vibecoding #Harnessengineering

  10. AI chipmaker Groq confirms $650M raise, re-staffs after Nvidia’s $20B not-acqui-hire deal

    What does an AI company do after one of those not-acqui-hire deals, where a rival pays investors a…
    #NewsBeep #News #US #USA #UnitedStates #UnitedStatesOfAmerica #Artificialintelligence #AI #ArtificialIntelligence #Groq #Technology
    newsbeep.com/us/720212/

  11. AI chipmaker Groq confirms $650M raise, re-staffs after Nvidia’s $20B not-acqui-hire deal

    What does an AI company do after one of those not-acqui-hire deals, where a rival pays investors a…
    #NewsBeep #News #US #USA #UnitedStates #UnitedStatesOfAmerica #Artificialintelligence #AI #ArtificialIntelligence #Groq #Technology
    newsbeep.com/us/720212/

  12. Telegram-бот с RAG на Cloudflare Workers: база знаний без векторов и без базы данных

    Строим Telegram-бота с RAG-поиском по базе знаний — без векторных БД, без эмбеддингов, без платной инфраструктуры. Поиск по ключевым словам через Jaccard, LLM через Groq, история сессий в Cloudflare KV, деплой одной командой. Стек: TypeScript + Telegraf + Cloudflare Workers.

    habr.com/ru/articles/1046495/

    #telegrambot #cloudflare_workers #typescript #llm #jaccard #groq #telegraf #serverless #knowledgebase #rag

  13. Telegram-бот с RAG на Cloudflare Workers: база знаний без векторов и без базы данных

    Строим Telegram-бота с RAG-поиском по базе знаний — без векторных БД, без эмбеддингов, без платной инфраструктуры. Поиск по ключевым словам через Jaccard, LLM через Groq, история сессий в Cloudflare KV, деплой одной командой. Стек: TypeScript + Telegraf + Cloudflare Workers.

    habr.com/ru/articles/1046495/

    #telegrambot #cloudflare_workers #typescript #llm #jaccard #groq #telegraf #serverless #knowledgebase #rag

  14. Telegram-бот с RAG на Cloudflare Workers: база знаний без векторов и без базы данных

    Строим Telegram-бота с RAG-поиском по базе знаний — без векторных БД, без эмбеддингов, без платной инфраструктуры. Поиск по ключевым словам через Jaccard, LLM через Groq, история сессий в Cloudflare KV, деплой одной командой. Стек: TypeScript + Telegraf + Cloudflare Workers.

    habr.com/ru/articles/1046495/

    #telegrambot #cloudflare_workers #typescript #llm #jaccard #groq #telegraf #serverless #knowledgebase #rag

  15. Как я довёл расходы на LLM до нуля: почему на бесплатных тарифах параллелизм — враг

    Это продолжение первой статьи про Briefka — там я описывал самого бота и базовую архитектуру каскада LLM-провайдеров. За прошедшие 4 месяца бот органически вырос с 59 до 84 пользователей, и именно на этом масштабе бесплатный каскад начал срываться на платного провайдера. Расскажу, почему так вышло и как я вернул расходы к нулю — с цифрами и кодом. Код ниже — реальные фрагменты из боевого Briefka, слегка сокращённые для читаемости: убраны логирование и сбор статистики.

    habr.com/ru/articles/1044546/

    #llm #ratelimit #asyncio #telegrambot #groq #deepseek #fallback #circuit_breaker

  16. Как я довёл расходы на LLM до нуля: почему на бесплатных тарифах параллелизм — враг

    Это продолжение первой статьи про Briefka — там я описывал самого бота и базовую архитектуру каскада LLM-провайдеров. За прошедшие 4 месяца бот органически вырос с 59 до 84 пользователей, и именно на этом масштабе бесплатный каскад начал срываться на платного провайдера. Расскажу, почему так вышло и как я вернул расходы к нулю — с цифрами и кодом. Код ниже — реальные фрагменты из боевого Briefka, слегка сокращённые для читаемости: убраны логирование и сбор статистики.

    habr.com/ru/articles/1044546/

    #llm #ratelimit #asyncio #telegrambot #groq #deepseek #fallback #circuit_breaker

  17. Как я довёл расходы на LLM до нуля: почему на бесплатных тарифах параллелизм — враг

    Это продолжение первой статьи про Briefka — там я описывал самого бота и базовую архитектуру каскада LLM-провайдеров. За прошедшие 4 месяца бот органически вырос с 59 до 84 пользователей, и именно на этом масштабе бесплатный каскад начал срываться на платного провайдера. Расскажу, почему так вышло и как я вернул расходы к нулю — с цифрами и кодом. Код ниже — реальные фрагменты из боевого Briefka, слегка сокращённые для читаемости: убраны логирование и сбор статистики.

    habr.com/ru/articles/1044546/

    #llm #ratelimit #asyncio #telegrambot #groq #deepseek #fallback #circuit_breaker

  18. Как я довёл расходы на LLM до нуля: почему на бесплатных тарифах параллелизм — враг Это продолжение первой ст...

    #llm #rate-limit #asyncio #telegram-bot #groq #deepseek #fallback #circuit #breaker

    Origin | Interest | Match
  19. RT @TechCrunch: Nach Nvidias 20-Mrd.-Dollar-„Not-Aqui-Hire“ soll das KI-Chip-Startup Groq reportedly 650 Millionen Dollar einsammeln.

    mehr auf Arint.info

    #Groq #Halbleiter #Investment #KI #Nvidia #Startup #arint_info

    https://x.com/TechCrunch/status/2060413760268087583#m

  20. А есть ли бесплатные API нейросетей?

    Третьего дня я решил сделать лид-магнит для своего Telegram-канала. Схема такая - бот собирает у пользователя текст, обрабатывает его нейросетью, выдает что-то полезное, и в конце просит подписаться на канал в обмен на результат. Aiogram 3, Python, VPS за 150 рублей - ничего необычного. Встал первый вопрос - за что платить? Бот прототипный, аудитория на входе пока еще, собственно, не особо и понятно сколько человек. Платить $20 в месяц ради теста гипотезы - нет. Мы не ищем легких путей. Пошел разбираться, что вообще бесплатного есть.

    habr.com/ru/articles/1041398/

    #бесплатные_api #groq #groq_api #openrouter #gemini_api #telegram_бот

  21. А есть ли бесплатные API нейросетей?

    Третьего дня я решил сделать лид-магнит для своего Telegram-канала. Схема такая - бот собирает у пользователя текст, обрабатывает его нейросетью, выдает что-то полезное, и в конце просит подписаться на канал в обмен на результат. Aiogram 3, Python, VPS за 150 рублей - ничего необычного. Встал первый вопрос - за что платить? Бот прототипный, аудитория на входе пока еще, собственно, не особо и понятно сколько человек. Платить $20 в месяц ради теста гипотезы - нет. Мы не ищем легких путей. Пошел разбираться, что вообще бесплатного есть.

    habr.com/ru/articles/1041398/

    #бесплатные_api #groq #groq_api #openrouter #gemini_api #telegram_бот

  22. А есть ли бесплатные API нейросетей?

    Третьего дня я решил сделать лид-магнит для своего Telegram-канала. Схема такая - бот собирает у пользователя текст, обрабатывает его нейросетью, выдает что-то полезное, и в конце просит подписаться на канал в обмен на результат. Aiogram 3, Python, VPS за 150 рублей - ничего необычного. Встал первый вопрос - за что платить? Бот прототипный, аудитория на входе пока еще, собственно, не особо и понятно сколько человек. Платить $20 в месяц ради теста гипотезы - нет. Мы не ищем легких путей. Пошел разбираться, что вообще бесплатного есть.

    habr.com/ru/articles/1041398/

    #бесплатные_api #groq #groq_api #openrouter #gemini_api #telegram_бот

  23. Times of India | Jensen Huang may claim Nvidia chips can do "a whole bunch of applications" vs Google TPUs, but as Google AI Demis Hassabis says: A lot of people would like to run ...

    AI generated summary, Read the full article for complete information.

    The rivalry between Google and Nvidia is intensifying as the AI boom shifts from training models to running them efficiently for rapid answers. Nvidia’s CEO Jensen Huang argues that its GPUs are more versatile, handling a wide range of applications, while Google is betting on its custom Tensor Processing Units (TPUs) specially tuned for inference workloads, which it expects to dominate as demand for fast, cost‑effective AI services grows. Google’s chief scientist Jeff Dean emphasizes the need to specialize chips for training or inference, and DeepMind CEO Demis Hassabis notes that many AI labs now prefer running on both Nvidia GPUs and Google TPUs, with interest in TPUs at an all‑time high. Analysts see inference as the next battleground, citing Google’s decade‑long experience in chip design and its Gemini model’s strong reasoning performance, while Nvidia has invested heavily in inference technology through acquisitions such as Groq. The competition reflects a broader industry shift toward specialized hardware that can deliver scalable, low‑latency AI services.

    Read more: timesofindia.indiatimes.com/te

    #JensenHuang #Nvidia #Google #DemisHassabis #JeffDean #DeepMind #Groq #ChatGPT #Gemini #

    AI generated summary, Read the full article for complete information.

  24. SwiftSlate ist so eine App, die sofort hängen bleibt.

    Ein systemweiter AI-Textassistent für Android: Du tippst z. B. ?fix oder ?formal direkt im Eingabefeld – und dein Text wird sofort ersetzt. Kein Copy-Paste. Kein App-Wechsel. Unterstützt Gemini, Groq und OpenAI-kompatible Endpunkte.

    Für alle, die AI auf Android wirklich im Alltag nutzen wollen, ist das richtig stark.

    #Android #AI #SwiftSlate #OpenSource #Gemini #Groq #Productivity #FOSS #RawInstinctAI #OpenAI

  25. SwiftSlate ist so eine App, die sofort hängen bleibt.

    Ein systemweiter AI-Textassistent für Android: Du tippst z. B. ?fix oder ?formal direkt im Eingabefeld – und dein Text wird sofort ersetzt. Kein Copy-Paste. Kein App-Wechsel. Unterstützt Gemini, Groq und OpenAI-kompatible Endpunkte.

    Für alle, die AI auf Android wirklich im Alltag nutzen wollen, ist das richtig stark.

    #Android #AI #SwiftSlate #OpenSource #Gemini #Groq #Productivity #FOSS #RawInstinctAI #OpenAI

  26. SwiftSlate ist so eine App, die sofort hängen bleibt.

    Ein systemweiter AI-Textassistent für Android: Du tippst z. B. ?fix oder ?formal direkt im Eingabefeld – und dein Text wird sofort ersetzt. Kein Copy-Paste. Kein App-Wechsel. Unterstützt Gemini, Groq und OpenAI-kompatible Endpunkte.

    Für alle, die AI auf Android wirklich im Alltag nutzen wollen, ist das richtig stark.

    #Android #AI #SwiftSlate #OpenSource #Gemini #Groq #Productivity #FOSS #RawInstinctAI #OpenAI

  27. Все переводчики речи в реальном времени — херня. Я написал свой. Тоже херня, но бесплатная

    Перепробовал всё что есть на рынке, потратил на подписки больше чем на кофе, и в итоге сел писать с нуля. Вот что вышло AI Open Source Voice AI Real-time перевод Deepgram Groq Piper TTS STT TTS LLM Google Meet Zoom Личный опыт Elixir Rust macOS Apple Silicon Speech-to-Text Text-to-Speech Сижу на рабочем созвоне. Обсуждаем архитектуру нового сервиса. Технически я всё понимаю - документацию на английском читаю без словаря, код ревьюю, в Slack переписываюсь нормально. А вот когда надо открыть рот и сказать что-то сложнее "I agree" - начинается цирк. Пауза. Подбираю слова. Коллега уже ответил за меня. Знакомо? Мне - до зубного скрежета. Я CTO, последние годы плотно работаю с AI-интеграциями. Могу собрать систему автоматического обзвона клиентов с клонированием голосов, поднять флот ботов для скана Телеги, собрать архитектуру которая выдержит тысячи пользователей за копейки. А сам на созвоне звучу как иностранец с разговорником. Ирония уровня бог. И вот в голове простая картинка: я говорю по-русски, собеседник слышит английский. Он отвечает по-английски, я слышу русский. В реальном времени. Без пауз на 10 секунд. Без субтитров - именно голосом. С любым приложением: Meet, Zoom, Slack, Discord. Пошёл искать. И тут началось.

    habr.com/ru/articles/1019458/

    #realtime_communications #translations #speechtotext #texttospeech #deepgram #groq #elixir #rust #open_source #voice_ai

  28. Все переводчики речи в реальном времени — херня. Я написал свой. Тоже херня, но бесплатная

    Перепробовал всё что есть на рынке, потратил на подписки больше чем на кофе, и в итоге сел писать с нуля. Вот что вышло AI Open Source Voice AI Real-time перевод Deepgram Groq Piper TTS STT TTS LLM Google Meet Zoom Личный опыт Elixir Rust macOS Apple Silicon Speech-to-Text Text-to-Speech Сижу на рабочем созвоне. Обсуждаем архитектуру нового сервиса. Технически я всё понимаю - документацию на английском читаю без словаря, код ревьюю, в Slack переписываюсь нормально. А вот когда надо открыть рот и сказать что-то сложнее "I agree" - начинается цирк. Пауза. Подбираю слова. Коллега уже ответил за меня. Знакомо? Мне - до зубного скрежета. Я CTO, последние годы плотно работаю с AI-интеграциями. Могу собрать систему автоматического обзвона клиентов с клонированием голосов, поднять флот ботов для скана Телеги, собрать архитектуру которая выдержит тысячи пользователей за копейки. А сам на созвоне звучу как иностранец с разговорником. Ирония уровня бог. И вот в голове простая картинка: я говорю по-русски, собеседник слышит английский. Он отвечает по-английски, я слышу русский. В реальном времени. Без пауз на 10 секунд. Без субтитров - именно голосом. С любым приложением: Meet, Zoom, Slack, Discord. Пошёл искать. И тут началось.

    habr.com/ru/articles/1019458/

    #realtime_communications #translations #speechtotext #texttospeech #deepgram #groq #elixir #rust #open_source #voice_ai

  29. NVIDIA’s new Vera Rubin platform brings together specialized chips (Vera CPUs, Rubin GPUs, Groq LPUs, and BlueField-4 DPUs) into coordinated, rack-scale systems designed for real-time AI.

    The big shift: AI isn’t just about training models anymore — it’s about orchestrating entire systems to power intelligent, autonomous agents in real time.
    buysellram.com/blog/the-agenti
    #NVIDIAGTC #AgenticAI #VeraRubin #DataCenter #GPU #InferenceFactory #AIInfrastructure #Groq #NVIDIA #NVLink #AIHardware #technology

  30. NVIDIA’s new Vera Rubin platform brings together specialized chips (Vera CPUs, Rubin GPUs, Groq LPUs, and BlueField-4 DPUs) into coordinated, rack-scale systems designed for real-time AI.

    The big shift: AI isn’t just about training models anymore — it’s about orchestrating entire systems to power intelligent, autonomous agents in real time.
    buysellram.com/blog/the-agenti
    #NVIDIAGTC #AgenticAI #VeraRubin #DataCenter #GPU #InferenceFactory #AIInfrastructure #Groq #NVIDIA #NVLink #AIHardware #technology

  31. US stock markets added on Monday, closing higher in New York. The Dow finished the day at 46,946 points, up 0.8 % from the previous session. A few minutes earli... news.osna.fm/?p=38427 | #news #amid #fuels #groq #hormuz

  32. SRAM. Static RAM. The stuff used for CPU caches, including AMD's 3D chips.

    There have been mumbles of CPU prices spiking like DRAM....this may be part of why.

    "Companies like Cerebras, Groq, and d-Matrix are designing AI inference chips that use massive amounts of on-chip SRAM instead of relying on external DRAM (HBM), which significantly reduces latency and power consumption."

    nVidia bought Groq. Amazon and Cerebras just signed a deal. Cerebras’ WSE-3 chip includes 900,000 cores and 44 gigabytes of on-chip SRAM.

    Wait for it...............

    #ai #dram #memory #sram #datacenters #gpu #cerebras #groq #amazon #nvidia

  33. 🚀 Ra mắt Oddvision – tiện ích Chrome cho phép trả lời ngay trên mọi trang web bằng phím tắt Alt+1 (capture), Alt+2 (analyze), Alt+3 (overlay). Chuyển từ API OpenAI (2.5s) sang Groq Llama‑3‑70b (<400ms) nên trải nghiệm “instant”. Có gói miễn phí 3 truy vấn/tuần, thích chia sẻ kinh nghiệm giới hạn Manifest V3. #CôngCụ #Extension #Chrome #AI #Oddvision #Llama3 #Groq #Developer #SinhViên

    reddit.com/r/SaaS/comments/1qh

  34. So sánh hiệu suất và chi phí giữa các mô hình AI:

    - **Ollama (CPU cục bộ)**: Miễn phí nhưng chậm (45 phút).
    - **OpenAI (GPT-4o)**: $5, nhanh (5 phút).
    - **Groq (Llama-3-70b)**: Chỉ $0.10, siêu nhanh (30 giây) - "Chén Thánh" của AI!

    #AI #TríTuệNhânTạo #CôngNghệ #SoSánh #Ollama #OpenAI #Groq #Llama3

    i.redd.it/zoa4sb80jbcg1.png