home.social

#googlebot — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #googlebot, aggregated by home.social.

fetched live
  1. SEO трафик на сайте есть, а заявок нет. Я полез в серверные логи и посмотрел, кто на самом деле «ходит» на сайт

    У сайта рос органический трафик, позиции были в порядке, а заявок всё равно не хватало. Клиент задал простой вопрос: «Если людей стало больше, где заявки?» Я полез не только в Метрику, но и в access.log — и там уже начался отдельный зоопарк: Googlebot, YandexBot, GPTBot, Ahrefs, Semrush, MJ12bot, DataForSEO и Chrome, который вёл себя совсем не как человек. В статье покажу, как отличать визиты от серверных запросов, проверять настоящего Googlebot, разбирать crawler'ов и решать, кого вообще стоит пускать на коммерческий сайт через robots.txt .

    habr.com/ru/articles/1070350/

    #SEO #техническое_SEO #серверные_логи #accesslog #robotstxt #Googlebot #YandexBot #вебкраулеры #роботный_трафик #UserAgent

  2. Momenteel lijkt het wel alsof het halve internet los gaat. De logfiles staan vol met troep... Ik denk een overblijfsel van een recente zwakheid in #WordPress. Hopelijk waait de boel snel weer over.

    Tussen alle troep, ben ik dan wel weer blij om deze bot te zien:

    "Mozilla/5.0 (compatible; Googlebot/2.1; +google.com/bot.html)"

    Ga maar eens lekker de inhoud van mijn blogs indexeren ja 👍

    #GoogleBot

  3. Momenteel lijkt het wel alsof het halve internet los gaat. De logfiles staan vol met troep... Ik denk een overblijfsel van een recente zwakheid in #WordPress. Hopelijk waait de boel snel weer over.

    Tussen alle troep, ben ik dan wel weer blij om deze bot te zien:

    "Mozilla/5.0 (compatible; Googlebot/2.1; +google.com/bot.html)"

    Ga maar eens lekker de inhoud van mijn blogs indexeren ja 👍

    #GoogleBot

  4. ICYMI: Google's expiry tag forces a re-crawl, and it won't fix 50% of 404 hits: Classifieds site cycles 10,000 listings monthly on 24-72 hour lifespans against 5,000 stable URLs. Half of Googlebot hits land on 404s. What the tag misses. ppc.land/googles-expiry-tag-fo #SEO #Googlebot #404Errors #Webmasters #DigitalMarketing

  5. ICYMI: Google's expiry tag forces a re-crawl, and it won't fix 50% of 404 hits: Classifieds site cycles 10,000 listings monthly on 24-72 hour lifespans against 5,000 stable URLs. Half of Googlebot hits land on 404s. What the tag misses. ppc.land/googles-expiry-tag-fo #SEO #Googlebot #404Errors #Webmasters #DigitalMarketing

  6. Google's expiry tag forces a re-crawl, and it won't fix 50% of 404 hits: Classifieds site cycles 10,000 listings monthly on 24-72 hour lifespans against 5,000 stable URLs. Half of Googlebot hits land on 404s. What the tag misses. ppc.land/googles-expiry-tag-fo #SEO #Googlebot #404errors #DigitalMarketing #WebDevelopment

  7. Google's expiry tag forces a re-crawl, and it won't fix 50% of 404 hits: Classifieds site cycles 10,000 listings monthly on 24-72 hour lifespans against 5,000 stable URLs. Half of Googlebot hits land on 404s. What the tag misses. ppc.land/googles-expiry-tag-fo #SEO #Googlebot #404errors #DigitalMarketing #WebDevelopment

  8. Cloudflare: из «трубы» — в вахтёра и кассира

    (про трубу в заголовке я немного зря - CF всегда были средством доставки, но любили их и за защиту; другими словами, надо бы писать про "трубу с ситечком", но в заголовок такое не ложится) Сама новость: Cloudflare, всем нам хорошо знакомый провайдер CDN и WAF, теперь будет не просто выделять в статистике посещений ИИ-ботов, но и разделять их по категориям, разрешая одним то, что запрещено другим. Ну и да, прямо замаячила тема оплаты за доступ. Brave new world, как по мне! Cloudflare разделила ИИ-ботов на три категории: Search , Agent и Training . Владельцы сайтов теперь могут независимо разрешать или блокировать поисковую индексацию, обращения агентов от имени пользователей и сбор данных для обучения моделей. Для каждой категории доступны три режима: разрешить , заблокировать везде или заблокировать только на страницах с рекламой . В результате Cloudflare уже не только пропускает трафик и отфильтровывает мусор. Она начинает решать, кто и зачем может читать сайт, а следом собирается брать оплату за проход.

    habr.com/ru/articles/1058108/

    #Cloudflare #ИИботы #вебкраулеры #WAF #CDN #BotBase #robotstxt #Googlebot #Pay_Per_Crawl #монетизация_контента

  9. Почему Google не индексирует страницы, хотя технически всё в порядке

    У меня есть сайт на Next.js. Часть страниц индексируется почти сразу. Часть застряла в статусе «Обнаружено, не проиндексировано» уже две недели. Самое неприятное в том что все страницы технически одинаковые. Тот же фреймворк, тот же сервер, тот же sitemap. Расскажу, как я перебирал гипотезы одну за другой, и что в итоге осталось. Читать разбор

    habr.com/ru/articles/1052898/

    #google #seo #googlebot #sitemap #sitemapxml #robotstxt #searchconsoler #nextjs #caching

  10. Is it now okay to consider GoogleBot hostile and just 403 them away?

    #WebAdmin #FreeWeb #GoogleBot #Google

  11. Is it now okay to consider GoogleBot hostile and just 403 them away?

    #WebAdmin #FreeWeb #GoogleBot #Google

  12. After Google’s announcement that they will start showing AI results rather than links in search results a lot of people showed interest in blocking them from scanning their sites.
    If you want to do this rather block their IP ranges than use robots.txt. I recommend blocking all Google Cloud IP ranges as well. I only see malicious bot traffic from there.
    Here is a good resource.

    searchengineworld.com/googles-

    #google #googlebot #googleai #hosting #AI #seo #googlesearch

  13. After Google’s announcement that they will start showing AI results rather than links in search results a lot of people showed interest in blocking them from scanning their sites.
    If you want to do this rather block their IP ranges than use robots.txt. I recommend blocking all Google Cloud IP ranges as well. I only see malicious bot traffic from there.
    Here is a good resource.

    searchengineworld.com/googles-

    #google #googlebot #googleai #hosting #AI #seo #googlesearch

  14. Пять неочевидных вещей, которые я узнал, запуская кино-соцсеть: от robots.txt-ловушки до 24-мерной математики вкуса

    Последние полгода я работаю над VibeMuvik — кино-соцсетью с рецензиями, дебатами и синхронным просмотром фильмов. Одна из тех штук, которые «ну вроде несложно», пока не начинаешь копать. Эта статья — про неожиданные находки . Не про «как я выбрал стек» (скучно) и не про «туториал по WebRTC» (и без меня есть). Это пять ситуаций, в которых я споткнулся, обнаружил что-то интересное, и подумал «об этом стоит рассказать — другим пригодится». Поехали.

    habr.com/ru/articles/1027876/

    #robotstxt #SEO #WebRTC #Nextjs #IndexNow #sitemap #Googlebot #Cinema_DNA #синхронный_просмотр #рекомендательные_системы

  15. FYI: Google rewrites Googlebot's rulebook: 2MB limits, IP moves, and what crawlers really are: Google today published two blog posts revealing Googlebot's true architecture as a shared SaaS platform, a 2MB fetch limit, and a new IP ranges directory path. ppc.land/google-rewrites-googl #Googlebot #SEO #WebCrawlers #DigitalMarketing #SaaS

  16. FYI: Google rewrites Googlebot's rulebook: 2MB limits, IP moves, and what crawlers really are: Google today published two blog posts revealing Googlebot's true architecture as a shared SaaS platform, a 2MB fetch limit, and a new IP ranges directory path. ppc.land/google-rewrites-googl #Googlebot #SEO #WebCrawlers #DigitalMarketing #SaaS

  17. ICYMI: Google rewrites Googlebot's rulebook: 2MB limits, IP moves, and what crawlers really are: Google today published two blog posts revealing Googlebot's true architecture as a shared SaaS platform, a 2MB fetch limit, and a new IP ranges directory path. ppc.land/google-rewrites-googl #Google #Googlebot #SEO #Crawlers #WebMaster

  18. Inside Googlebot: demystifying crawling, fetching, and the bytes we process: developers.google.com/search/b. A great post that clarifies how #Googlebot does its business for people in the #SEO community.

  19. Inside Googlebot: demystifying crawling, fetching, and the bytes we process: developers.google.com/search/b. A great post that clarifies how #Googlebot does its business for people in the #SEO community.

  20. Google rewrites Googlebot's rulebook: 2MB limits, IP moves, and what crawlers really are: Google today published two blog posts revealing Googlebot's true architecture as a shared SaaS platform, a 2MB fetch limit, and a new IP ranges directory path. ppc.land/google-rewrites-googl #Google #Googlebot #SEO #WebCrawlers #DigitalMarketing

  21. Google rewrites Googlebot's rulebook: 2MB limits, IP moves, and what crawlers really are: Google today published two blog posts revealing Googlebot's true architecture as a shared SaaS platform, a 2MB fetch limit, and a new IP ranges directory path. ppc.land/google-rewrites-googl #Google #Googlebot #SEO #WebCrawlers #DigitalMarketing

  22. FYI: Googlebot is not a program - Google engineers finally explain what it really is: Google engineers reveal Googlebot is a misnomer for a central SaaS crawling platform serving dozens of products, with a 15 MB default file size limit and geo-crawling constraints. ppc.land/googlebot-is-not-a-pr #Googlebot #SEO #WebCrawling #DigitalMarketing #SaaS

  23. FYI: Googlebot is not a program - Google engineers finally explain what it really is: Google engineers reveal Googlebot is a misnomer for a central SaaS crawling platform serving dozens of products, with a 15 MB default file size limit and geo-crawling constraints. ppc.land/googlebot-is-not-a-pr #Googlebot #SEO #WebCrawling #DigitalMarketing #SaaS

  24. ICYMI: Googlebot is not a program - Google engineers finally explain what it really is: Google engineers reveal Googlebot is a misnomer for a central SaaS crawling platform serving dozens of products, with a 15 MB default file size limit and geo-crawling constraints. ppc.land/googlebot-is-not-a-pr #Googlebot #SEO #Crawling #SaaS #DigitalMarketing

  25. ICYMI: Googlebot is not a program - Google engineers finally explain what it really is: Google engineers reveal Googlebot is a misnomer for a central SaaS crawling platform serving dozens of products, with a 15 MB default file size limit and geo-crawling constraints. ppc.land/googlebot-is-not-a-pr #Googlebot #SEO #Crawling #SaaS #DigitalMarketing

  26. Googlebot is not a program - Google engineers finally explain what it really is: Google engineers reveal Googlebot is a misnomer for a central SaaS crawling platform serving dozens of products, with a 15 MB default file size limit and geo-crawling constraints. ppc.land/googlebot-is-not-a-pr #Googlebot #SEO #WebCrawling #SaaS #DigitalMarketing

  27. Googlebot is not a program - Google engineers finally explain what it really is: Google engineers reveal Googlebot is a misnomer for a central SaaS crawling platform serving dozens of products, with a 15 MB default file size limit and geo-crawling constraints. ppc.land/googlebot-is-not-a-pr #Googlebot #SEO #WebCrawling #SaaS #DigitalMarketing

  28. FYI: Google's secret crawl logic, finally explained in one page: Google published a new web crawling overview on March 3, 2026, detailing how Googlebot discovers, renders, and manages site access across 30+ years of web indexing. ppc.land/googles-secret-crawl- #Google #SEO #WebCrawling #Googlebot #DigitalMarketing

  29. FYI: Google's secret crawl logic, finally explained in one page: Google published a new web crawling overview on March 3, 2026, detailing how Googlebot discovers, renders, and manages site access across 30+ years of web indexing. ppc.land/googles-secret-crawl- #Google #SEO #WebCrawling #Googlebot #DigitalMarketing

  30. ICYMI: Google's secret crawl logic, finally explained in one page: Google published a new web crawling overview on March 3, 2026, detailing how Googlebot discovers, renders, and manages site access across 30+ years of web indexing. ppc.land/googles-secret-crawl- #Google #SEO #WebCrawling #Googlebot #DigitalMarketing

  31. ICYMI: Google's secret crawl logic, finally explained in one page: Google published a new web crawling overview on March 3, 2026, detailing how Googlebot discovers, renders, and manages site access across 30+ years of web indexing. ppc.land/googles-secret-crawl- #Google #SEO #WebCrawling #Googlebot #DigitalMarketing

  32. Google's secret crawl logic, finally explained in one page: Google published a new web crawling overview on March 3, 2026, detailing how Googlebot discovers, renders, and manages site access across 30+ years of web indexing. ppc.land/googles-secret-crawl- #Google #SEO #WebCrawling #Googlebot #DigitalMarketing

  33. Google's secret crawl logic, finally explained in one page: Google published a new web crawling overview on March 3, 2026, detailing how Googlebot discovers, renders, and manages site access across 30+ years of web indexing. ppc.land/googles-secret-crawl- #Google #SEO #WebCrawling #Googlebot #DigitalMarketing

  34. Testing tool simulates Google's 2MB HTML limit as SEO professionals assess crawling impact: Dave Smart added 2MB truncation feature to Tame the Bots fetch tool on February 6, enabling technical SEO professionals to simulate Googlebot's reduced file size limits. ppc.land/testing-tool-simulate #SEO #GoogleBot #HTML #Crawling #DigitalMarketing

  35. Testing tool simulates Google's 2MB HTML limit as SEO professionals assess crawling impact: Dave Smart added 2MB truncation feature to Tame the Bots fetch tool on February 6, enabling technical SEO professionals to simulate Googlebot's reduced file size limits. ppc.land/testing-tool-simulate #SEO #GoogleBot #HTML #Crawling #DigitalMarketing