home.social

#robots-txt — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #robots-txt, aggregated by home.social.

fetched live
  1. Карго-культы технического SEO: что robots.txt и sitemap.xml делают на самом деле

    Priority и changefreq, которые Google игнорирует. Disallow как «запрет индексации». robots.txt в роли замка для чувствительных данных. Разбираем семь живучих ритуалов вокруг двух старейших файлов технического SEO — и что эти файлы делают на самом деле.

    habr.com/ru/articles/1067780/

    #robotstxt #sitemapxml #техническое_seo #индексация #краулинговый_бюджет #noindex #поисковые_роботы #seo #search_console #яндекс_вебмастер

  2. AI-краулеры в 2026: кого пускать в robots.txt, чтобы попадать в ответы ИИ

    Полгода назад ко мне постучался один из будущих клиентов с вопросом, почему ChatGPT не знает про его компанию, хотя сайт в топ-3 Google по нужному запросу. Индексация была нормальной, ссылки — тоже. Проблема нашлась за пять минут в логах: в robots.txt висел Disallow: / для всех user-agent, кроме Googlebot и YandexBot. GPTBot, ClaudeBot, PerplexityBot получали 403 ещё на первом запросе. Полгода сайт был невидим для трёх из четырёх крупных AI-поисковиков, и никто об этом не знал — в Search Console и Вебмастере всё выглядело зелёным. Случай типовой. Проверка индексации показывает, видит ли сайт поисковик. Попадёт ли страница в ответ модели, которая формирует его вместо поисковой выдачи, — вопрос отдельный. Это две разные аудитории ботов, и вторая почти никогда не попадает в стандартный SEO-чеклист.

    habr.com/ru/articles/1063216/

    #AIкраулеры #robotstxt #GPTBot #AEO #GEO #AIвидимость #нейросети #AIO #YandexBot #яндексвебмастер

  3. GEO: техника нейровыдачи: чанкование с оверлапом, Schema.org, robots.txt и пр. Часть 3 из 3

    В этой части — самое приземлённое про GEO-продвижение. Никакой философии, только то, что можно взять и в понедельник начать делать руками: как резать страницу под чанки, какой JSON-LD писать, какие боты пускать на сайт, как честно замерить цитируемость и как поймать переходы в Метрике и GA. То есть форма, в которую упаковывается уже сформулированная ценность. Если ценности нет, форма не спасёт — об этом была вторая часть. Часть 1 — Нейросайт глазами бизнеса: позиции, трафик, конверсия и скорость изменений Часть 2 — Как принудить нейросеть рассказать про ваш продукт

    habr.com/ru/articles/1060370/

    #GEO #Schemaorg #JSONLD #robotstxt #llmstxt #чанкование #SEO #нейровыдача #AI_Overview #Share_of_Voice

  4. Cloudflare: из «трубы» — в вахтёра и кассира

    (про трубу в заголовке я немного зря - CF всегда были средством доставки, но любили их и за защиту; другими словами, надо бы писать про "трубу с ситечком", но в заголовок такое не ложится) Сама новость: Cloudflare, всем нам хорошо знакомый провайдер CDN и WAF, теперь будет не просто выделять в статистике посещений ИИ-ботов, но и разделять их по категориям, разрешая одним то, что запрещено другим. Ну и да, прямо замаячила тема оплаты за доступ. Brave new world, как по мне! Cloudflare разделила ИИ-ботов на три категории: Search , Agent и Training . Владельцы сайтов теперь могут независимо разрешать или блокировать поисковую индексацию, обращения агентов от имени пользователей и сбор данных для обучения моделей. Для каждой категории доступны три режима: разрешить , заблокировать везде или заблокировать только на страницах с рекламой . В результате Cloudflare уже не только пропускает трафик и отфильтровывает мусор. Она начинает решать, кто и зачем может читать сайт, а следом собирается брать оплату за проход.

    habr.com/ru/articles/1058108/

    #Cloudflare #ИИботы #вебкраулеры #WAF #CDN #BotBase #robotstxt #Googlebot #Pay_Per_Crawl #монетизация_контента

  5. The rise of consumer generative AI has created an insatiable demand for content. Though I've tried to block the bots, I've failed. Here's why.

    plagiarismtoday.com/2026/06/30

    #Copyright #AI #RobotsTXT #SEO

  6. The rise of consumer generative AI has created an insatiable demand for content. Though I've tried to block the bots, I've failed. Here's why.

    plagiarismtoday.com/2026/06/30

    #Copyright #AI #RobotsTXT #SEO

  7. Почему Google не индексирует страницы, хотя технически всё в порядке

    У меня есть сайт на Next.js. Часть страниц индексируется почти сразу. Часть застряла в статусе «Обнаружено, не проиндексировано» уже две недели. Самое неприятное в том что все страницы технически одинаковые. Тот же фреймворк, тот же сервер, тот же sitemap. Расскажу, как я перебирал гипотезы одну за другой, и что в итоге осталось. Читать разбор

    habr.com/ru/articles/1052898/

    #google #seo #googlebot #sitemap #sitemapxml #robotstxt #searchconsoler #nextjs #caching

  8. Маленький файл robots.txt и большие последствия одной строки

    Разбираемся, как работает robots.txt, почему его часто путают с инструментами индексации и какую роль он играет в эпоху ИИ-сканеров.

    habr.com/ru/companies/hostkey/

    #robotstxt #SEO #поисковая_выдача #индексация #Яндекс #Google #ИИ #сканеры #тестирование #hostkey

  9. Спор про llms.txt не сходится: и критики, и хайп меряют не тот слой

    Один лагерь показывает 0,1% обращений в логах и хоронит файл. Другой обещает прирост цитируемости на 30–60%. Обе цифры реальны. Они измеряют разные вещи, и пока спорщики этого не видят, спор идёт по кругу. Я полгода вожусь с llms.txt на клиентских проектах и на собственном сайте. В мае прогнал восемь AI-систем через контролируемый тест, чтобы перестать гадать и увидеть, кто реально читает файл. Результат не подтвердил ни одну из двух громких позиций целиком. Он показал третью картину, которую обе стороны пропускают: llms.txt живёт не в логах фоновых краулеров и не в магии ранжирования. Он живёт в агентном слое реального времени и в IDE-агентах. Это узкое место, но там он работает.

    habr.com/ru/articles/1043736/

    #llmstxt #AI_SEO #GEO #агентный_веб #LLM #robotstxt #AEO

  10. SEO-админка для большого каталога: sitemap, robots, мета-превью и тревоги поисковиков в одном месте

    Рассказываю, как мы собрали SEO-панель для динамического каталога: sitemap, robots.txt, мета-превью, RSS, диагностика и переобход в одном интерфейсе. Без секретов и полного кода, но с архитектурой и граблями продакшена.

    habr.com/ru/articles/1043402/

    #SEO #sitemap #robotstxt #IndexNow #Nextjs #FastAPI #PostgreSQL #админка #поисковые_системы #техническое_SEO

  11. CW: google AI vs. your personal webpage

    Now that Google have announced their intention to discontinue Web Search, and replace it with LLM Summaries Only, I advise everyone that cares to update your robots.txt to disallow Googlebot (their original search index robot).

    Because after summer, there will be no web search to index your website, the data gathered by the "good" index robot will at best be discarded, or at worst, be fed to LLMs, leading to plagiarism of your content. In either scenario, no human visitors will be guided to your website.

    (If you don't want to be indexed by any search engine at all, you can disallow all robots. You can also add the noindex metatag to all your pages.)

    #robotstxt #webSearch #noAI

  12. CW: google AI vs. your personal webpage

    Now that Google have announced their intention to discontinue Web Search, and replace it with LLM Summaries Only, I advise everyone that cares to update your robots.txt to disallow Googlebot (their original search index robot).

    Because after summer, there will be no web search to index your website, the data gathered by the "good" index robot will at best be discarded, or at worst, be fed to LLMs, leading to plagiarism of your content. In either scenario, no human visitors will be guided to your website.

    (If you don't want to be indexed by any search engine at all, you can disallow all robots. You can also add the noindex metatag to all your pages.)

    #robotstxt #webSearch #noAI

  13. Robots.txt zůstává základní signál pro slušné crawlery, ale už neumí popsat hlavní problém: stejný veřejný obsah může sloužit klasickému vyhledávání, AI odpovědím, tréninku modelů i načtení na pokyn uživatele. Provozovatel webu proto musí oddělit účel přístupu, ověřovat identitu botů, měřit dopad na infrastrukturu a u hodnotného obsahu řešit i vynucení pravidel mimo samotný robots.txt.

    https://zdrojak.cz/clanky/robots-txt-nestaci-ai-crawleri-meni-jak-weby-chrani-obsah/
  14. Robots.txt zůstává základní signál pro slušné crawlery, ale už neumí popsat hlavní problém: stejný veřejný obsah může sloužit klasickému vyhledávání, AI odpovědím, tréninku modelů i načtení na pokyn uživatele. Provozovatel webu proto musí oddělit účel přístupu, ověřovat identitu botů, měřit dopad na infrastrukturu a u hodnotného obsahu řešit i vynucení pravidel mimo samotný robots.txt.

    https://zdrojak.cz/clanky/robots-txt-nestaci-ai-crawleri-meni-jak-weby-chrani-obsah/
  15. So, with Google announcing "Search is going full-AI, we won't be sending traffic to the original sites any more", someone else pointed out that this eradication of the traditional search-engine compact - we let you crawl our sites to create your index, and you send visitors to our sites when relevant - means that we can, and should, block all of Google's crawlers now. If they're going to just take, take, take and give nothing back, why let them access your content at all?

    But this is cute. Besides the fact that Google documents that some of their crawlers ignore robots.txt, there's this bit of fun. On this page (developers.google.com/crawling), they link to "the Google list of user agents" (developers.google.com/crawling).

    However, that links to 3 separate pages of them, and *each of those pages explicitly states that is not comprehensive, but only the ones they commonly get questions about*. And of course, none of the "User-triggered fetchers" obey robots.txt, along with some others.

    So Google isn't even reporting the full list of user-agents that can be used to stop their crawling.

    That is some bullshit.

    #Google #crawler #RobotsTxt #UserAgent #bullshit #antisocial #web #search #WebSearch #LLM #AI

  16. So, with Google announcing "Search is going full-AI, we won't be sending traffic to the original sites any more", someone else pointed out that this eradication of the traditional search-engine compact - we let you crawl our sites to create your index, and you send visitors to our sites when relevant - means that we can, and should, block all of Google's crawlers now. If they're going to just take, take, take and give nothing back, why let them access your content at all?

    But this is cute. Besides the fact that Google documents that some of their crawlers ignore robots.txt, there's this bit of fun. On this page (developers.google.com/crawling), they link to "the Google list of user agents" (developers.google.com/crawling).

    However, that links to 3 separate pages of them, and *each of those pages explicitly states that is not comprehensive, but only the ones they commonly get questions about*. And of course, none of the "User-triggered fetchers" obey robots.txt, along with some others.

    So Google isn't even reporting the full list of user-agents that can be used to stop their crawling.

    That is some bullshit.

    #Google #crawler #RobotsTxt #UserAgent #bullshit #antisocial #web #search #WebSearch #LLM #AI

  17. Scrapers vs Wikis: Person who runs a bunch of custom Wiki websites writes about abuse from scrapers
    weirdgloop.org/blog/clankers
    #via:lobsters #robotstxt #scraping #scaling #wiki #web #ai #+

  18. Пять неочевидных вещей, которые я узнал, запуская кино-соцсеть: от robots.txt-ловушки до 24-мерной математики вкуса

    Последние полгода я работаю над VibeMuvik — кино-соцсетью с рецензиями, дебатами и синхронным просмотром фильмов. Одна из тех штук, которые «ну вроде несложно», пока не начинаешь копать. Эта статья — про неожиданные находки . Не про «как я выбрал стек» (скучно) и не про «туториал по WebRTC» (и без меня есть). Это пять ситуаций, в которых я споткнулся, обнаружил что-то интересное, и подумал «об этом стоит рассказать — другим пригодится». Поехали.

    habr.com/ru/articles/1027876/

    #robotstxt #SEO #WebRTC #Nextjs #IndexNow #sitemap #Googlebot #Cinema_DNA #синхронный_просмотр #рекомендательные_системы

  19. FYI: Only 7.4% of Fortune 500 have an llms.txt file, study finds: ProGEO.ai research reveals just 7.4% of Fortune 500 companies have implemented llms.txt, while 92.8% use robots.txt and 53.8% use JSON-LD for AI visibility. ppc.land/only-7-4-of-fortune-5 #LLMSTXT #Fortune500 #AIVisibility #RobotsTxt #JSONLD

  20. Only 7.4% of Fortune 500 have an llms.txt file, study finds: ProGEO.ai research reveals just 7.4% of Fortune 500 companies have implemented llms.txt, while 92.8% use robots.txt and 53.8% use JSON-LD for AI visibility. ppc.land/only-7-4-of-fortune-5 #Fortune500 #AI #llms #robotsTxt #JSONLD

  21. Only 7.4% of Fortune 500 have an llms.txt file, study finds: ProGEO.ai research reveals just 7.4% of Fortune 500 companies have implemented llms.txt, while 92.8% use robots.txt and 53.8% use JSON-LD for AI visibility. ppc.land/only-7-4-of-fortune-5 #Fortune500 #AI #llms #robotsTxt #JSONLD

  22. Oh, this is #fun.

    #Applebot - Apple's web crawler, used for various things - is ignoring robots.txt rules governing crawling of websites.

    I have Applebot (and Applebot-Extended, which isn't really a crawler) in my robots.txt files, set to disallow all access. Has been that way for #yonks.

    And Applebot is consistently the highest-traffic crawler to my sites - at least of ones that actually bother to fetch robots.txt. Yesterday, for example, Applebot fetched robots.txt from one of my websites almost 800 times.

    Yes, it's really Apple, not someone faking the user-agent identifier. It's coming from the networks that Apple says can be used to identify Applebot access. DNS matches, everything.
    e.g. support.apple.com/en-ca/119829

    So: legendary Apple software quality. Documented to do the right thing, but actually doing the wrong thing. And completely failing to cache content, fetching the same file 800 times a day when it hasn't changed in years.

    Hey, Apple! Need a software engineer who's actually, you know, good at it? I'm available.

    #Apple #AppleInc #TimApple #WebCrawler #RobotsTxt #quality #WeveHeardOfIt #qwality #AppleQwality #legendary #TwoHardThings #caching #fail #engineer #software #SoftwareEngineer

  23. Oh, this is #fun.

    #Applebot - Apple's web crawler, used for various things - is ignoring robots.txt rules governing crawling of websites.

    I have Applebot (and Applebot-Extended, which isn't really a crawler) in my robots.txt files, set to disallow all access. Has been that way for #yonks.

    And Applebot is consistently the highest-traffic crawler to my sites - at least of ones that actually bother to fetch robots.txt. Yesterday, for example, Applebot fetched robots.txt from one of my websites almost 800 times.

    Yes, it's really Apple, not someone faking the user-agent identifier. It's coming from the networks that Apple says can be used to identify Applebot access. DNS matches, everything.
    e.g. support.apple.com/en-ca/119829

    So: legendary Apple software quality. Documented to do the right thing, but actually doing the wrong thing. And completely failing to cache content, fetching the same file 800 times a day when it hasn't changed in years.

    Hey, Apple! Need a software engineer who's actually, you know, good at it? I'm available.

    #Apple #AppleInc #TimApple #WebCrawler #RobotsTxt #quality #WeveHeardOfIt #qwality #AppleQwality #legendary #TwoHardThings #caching #fail #engineer #software #SoftwareEngineer

  24. FYI: Czech publishers get new robots.txt shield against AI scrapers: SPIR on March 19 updated its standard for Czech online publishers to opt out of AI text and data mining, adding real-time response crawlers to the scope of the robots.txt framework. ppc.land/czech-publishers-get- #AI #robotstxt #datautajení #česképublikace #ochranadat

  25. FYI: Czech publishers get new robots.txt shield against AI scrapers: SPIR on March 19 updated its standard for Czech online publishers to opt out of AI text and data mining, adding real-time response crawlers to the scope of the robots.txt framework. ppc.land/czech-publishers-get- #AI #robotstxt #datautajení #česképublikace #ochranadat

  26. ICYMI: Czech publishers get new robots.txt shield against AI scrapers: SPIR on March 19 updated its standard for Czech online publishers to opt out of AI text and data mining, adding real-time response crawlers to the scope of the robots.txt framework. ppc.land/czech-publishers-get- #technologie #publikace #AI #robotstxt #czechpublishing

  27. ICYMI: Czech publishers get new robots.txt shield against AI scrapers: SPIR on March 19 updated its standard for Czech online publishers to opt out of AI text and data mining, adding real-time response crawlers to the scope of the robots.txt framework. ppc.land/czech-publishers-get- #technologie #publikace #AI #robotstxt #czechpublishing

  28. Czech publishers get new robots.txt shield against AI scrapers: SPIR on March 19 updated its standard for Czech online publishers to opt out of AI text and data mining, adding real-time response crawlers to the scope of the robots.txt framework. ppc.land/czech-publishers-get- #CzechPublishing #AIScrapers #RobotsTxt #DataMining #OnlinePrivacy

  29. 📝 New article: Why We Reject Google: Our Anti-Surveillance SEO Policy

    An in-depth look at why Virebent.art deliberately blocks Google and other surveillance-based crawlers, and our strategy for visibility in a privacy-first web.

    🔗 virebent.art/blog/seo-policy.h

    #antiseo #robotstxt #surveillancecapitalism

  30. Robots.txt Generator - Retro Terminal Edition - Mehr als 200 Bots in der kostenfreien Version. Pures HTML, Javascript und ein bisschen CSS. Keine Third Parties, kein Framework, kein CDN, keine Cookies, kein Tracking, keine Werbung, kein BigTech-Gedönse, keine KI, sehr datenschutzfreundlich. Simple und effektiv im Retro-Style. Demnächst online.

    #teufelswerk #HTML #javascript #app #entwicklung #code #retro #css #robotstxt #generator #stopbots #bots #crawler #scraper #keineKI #cookieless #datenschutz

  31. Panduan memahami tiga opsi Cloudflare untuk konfigurasi robots.txt: Content Signals Policy, Instruct AI bots to not scrape, dan Disable configuration. Pelajari cara memberi instruksi pada AI crawler.

    #fediverse #Repost #WartaTekno #Mengelola #Robotstxt #Cloudflare

    dalam.web.id/artikel/cloudflar