#crawler — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #crawler, aggregated by home.social.
-
Как создать архитектуру B2B-пайплайна, в которой не проверяется лишнее и не теряется нужное
На входе у нас вакансия, на выходе — проверенный контакт в одной из трёх кампаний. Между ними: дубли, таймауты, catch-all, неоднозначные SMTP-ответы и окно, в котором процесс может упасть после передачи контакта, но до сохранения статуса. Показываю, как мы расставили дешёвые и дорогие проверки и почему выбрали at-least-once.
https://habr.com/ru/articles/1070562/
#B2Bпайплайн #atleastonce #дедупликация #DNS #SMTP #crawler #email_validation #emailаутрич
-
Go-juggler и протокол Juggler
ИИ-агентам нужно ходить по настоящему вебу. Но настоящий веб враждебен к автоматизации: Playwright блокируют, headless-версия Chrome палится по отпечаткам, а стелс-плагины просто становятся частью отпечатка. Если вы когда-нибудь писали скраперы или инструменты для агентов, вы знаете, как это бывает - всё работает локально, а потом продакшен разваливается за Cloudflare-челленджем. Несколько проектов пытались решить эту проблему со стороны браузера. Проект https://camoufox.com патчит Firefox на уровне C++, так что navigator.hardwareConcurrency, WebGL-рендереры, AudioContext, геометрия экрана и WebRTC подменяются ещё до того, как JavaScript их увидит. Браузер https://github.com/jo-inc/camofox-browser оборачивает этот движок в REST API, заточенный под агентов: снимки доступности вместо раздутого HTML, стабильные ссылки на элементы для кликов и изоляция сессий. Остаётся только одна дыра: инструментарий вокруг этой экосистемы завязан на JavaScript/Python. Если вы живёте в Go - а весь стек Go-агентов, взорвавшийся за последние пару лет, весомый аргумент в его пользу - вам оставалось писать сырые вызовы curl. Пакет https://github.com/yvv4git/go-juggler исправляет это. Это Go-клиент для протокола автоматизации Juggler (того самого, который патчит и расширяет Camoufox) под лицензией MIT. Он управляет Firefox/Camoufox из Go с единственной зависимостью и чистым, слоистым API.
-
ICYMI: PatronView blocks Amazon's AI crawler after 117,000 daily page reads: Anthropic's crawler hit a 35,000 to 1 crawl ratio, and CAPTCHA solve rates measured just 0.24%. The findings show why small operators are locking down servers. https://ppc.land/patronview-blocks-amazons-ai-crawler-after-117-000-daily-page-reads/ #AI #Crawler #WebSecurity #CAPTCHA #DataPrivacy
-
ICYMI: PatronView blocks Amazon's AI crawler after 117,000 daily page reads: Anthropic's crawler hit a 35,000 to 1 crawl ratio, and CAPTCHA solve rates measured just 0.24%. The findings show why small operators are locking down servers. https://ppc.land/patronview-blocks-amazons-ai-crawler-after-117-000-daily-page-reads/ #AI #Crawler #WebSecurity #CAPTCHA #DataPrivacy
-
Der Online-Handel steht unter automatisiertem Dauerbeschuss. Cyberkriminelle haben den Handelssektor fest im Visier. Einem Akamai-Report zufolge ging fast jeder zweite erfasste Zugriff auf einen KI-Bot zurück. Der Anteil lag bei 47,9 Prozent. Mehr als 70 Prozent der erkannten KI-Bot-Aktivitäten entfielen auf Crawler, die Inhalte für das Training großer Sprachmodelle sammeln.
https://www.speicherguide.de/management/cybersicherheit/akamai-ki-bots-praegen-fast-die-haelfte-des-handels-traffics-27031.html
#KI #Cyber #Crawler #Cybersecurity -
Der Online-Handel steht unter automatisiertem Dauerbeschuss. Cyberkriminelle haben den Handelssektor fest im Visier. Einem Akamai-Report zufolge ging fast jeder zweite erfasste Zugriff auf einen KI-Bot zurück. Der Anteil lag bei 47,9 Prozent. Mehr als 70 Prozent der erkannten KI-Bot-Aktivitäten entfielen auf Crawler, die Inhalte für das Training großer Sprachmodelle sammeln.
https://www.speicherguide.de/management/cybersicherheit/akamai-ki-bots-praegen-fast-die-haelfte-des-handels-traffics-27031.html
#KI #Cyber #Crawler #Cybersecurity -
How to use Hister's built-in crawler: https://hister.org/docs/crawler
-
How to use Hister's built-in crawler: https://hister.org/docs/crawler
-
Just finished Book 1 of the Dungeon Crawler Carl series by Matt Dinniman. My god, it's like six or seven so far and it was so well written. We have many cats and one in particular has been dubbed Princess Donut.
#read #dungeon #crawler #carl #donut #princess #books
mattdinniman.com
Matt DinnimanMatt Dinniman -
Just finished Book 1 of the Dungeon Crawler Carl series by Matt Dinniman. My god, it's like six or seven so far and it was so well written. We have many cats and one in particular has been dubbed Princess Donut.
#read #dungeon #crawler #carl #donut #princess #books
mattdinniman.com
Matt DinnimanMatt Dinniman -
And finally, #btracker instance for #I2P
http://btrackrqkjp6kgelov5a3uxisis77ofxqt5nvy5hvvtoybjpmq4q.b32.i2p* at this moment, crawler does not support B32 address family by the #librqbit dependency, but the catalog already returns the actual I2P peers for existing torrents from Yggdrasil and Mycelium nodes
-
And finally, #btracker instance for #I2P
http://btrackrqkjp6kgelov5a3uxisis77ofxqt5nvy5hvvtoybjpmq4q.b32.i2p* at this moment, crawler does not support B32 address family by the #librqbit dependency, but the catalog already returns the actual I2P peers for existing torrents from Yggdrasil and Mycelium nodes
-
CGE: визуализация кравлера и скрытых связей между поддоменами
Привет, Хабр! Хотелось бы поделиться с вами моим open-source проектом для поиска директорий, поддоменов, ака crawler. Я не говорю, что он перевернёт мир краулеров или превзойдёт Katana, но, думаю, утилита будет крайне полезна для red team-команды. https://github.com/a11mut3d/CGE Проблемы, которые решает CGE Современное веб-приложение — это не монолит, где всё в одном HTML, а куча микросервисов, API и в целом эндпоинтов. Составить карту всех запросов достаточно сложно, поэтому вы не видите картину целиком. CGE помогает в этой задаче. Он: — собирает все поддомены из SSL-сертификата (как в crt.sh ); — краулит каждый эндпоинт, парсит HTML, JS, формы, аплоады, файлы; — отслеживает, куда идут запросы в реальном времени через взаимодействие с формами; — строит граф взаимодействия эндпоинта в реальном времени. Как это выглядит На данный момент у CGE есть 2 варианта использования: web UI и CLI. Если про CLI особо и нечего расписывать (он просто выдаёт все найденные эндпоинты в консоль или по желанию сохраняет в файл), то на web UI давайте остановимся подробнее. Веб-интерфейс я постарался сделать в стиле Obsidian (спойлер: получилось не очень). — Каждая нода — хост (поддомен). — Ребро между нодами — факт HTTP-обмена информацией. — При клике на ноду мы получаем список всех эндпоинтов (даже тех, которые были замечены в запросах от других хостов). — При клике на ребро мы получаем все реальные запросы между хостами. Технические детали Реализовать я решил на Python с использованием BS, requests, DNS. В качестве базы данных я решил использовать Neo4j.
-
So, with Google announcing "Search is going full-AI, we won't be sending traffic to the original sites any more", someone else pointed out that this eradication of the traditional search-engine compact - we let you crawl our sites to create your index, and you send visitors to our sites when relevant - means that we can, and should, block all of Google's crawlers now. If they're going to just take, take, take and give nothing back, why let them access your content at all?
But this is cute. Besides the fact that Google documents that some of their crawlers ignore robots.txt, there's this bit of fun. On this page (https://developers.google.com/crawling/docs/robots-txt/create-robots-txt), they link to "the Google list of user agents" (https://developers.google.com/crawling/docs/crawlers-fetchers/overview-google-crawlers).
However, that links to 3 separate pages of them, and *each of those pages explicitly states that is not comprehensive, but only the ones they commonly get questions about*. And of course, none of the "User-triggered fetchers" obey robots.txt, along with some others.
So Google isn't even reporting the full list of user-agents that can be used to stop their crawling.
That is some bullshit.
#Google #crawler #RobotsTxt #UserAgent #bullshit #antisocial #web #search #WebSearch #LLM #AI
-
So, with Google announcing "Search is going full-AI, we won't be sending traffic to the original sites any more", someone else pointed out that this eradication of the traditional search-engine compact - we let you crawl our sites to create your index, and you send visitors to our sites when relevant - means that we can, and should, block all of Google's crawlers now. If they're going to just take, take, take and give nothing back, why let them access your content at all?
But this is cute. Besides the fact that Google documents that some of their crawlers ignore robots.txt, there's this bit of fun. On this page (https://developers.google.com/crawling/docs/robots-txt/create-robots-txt), they link to "the Google list of user agents" (https://developers.google.com/crawling/docs/crawlers-fetchers/overview-google-crawlers).
However, that links to 3 separate pages of them, and *each of those pages explicitly states that is not comprehensive, but only the ones they commonly get questions about*. And of course, none of the "User-triggered fetchers" obey robots.txt, along with some others.
So Google isn't even reporting the full list of user-agents that can be used to stop their crawling.
That is some bullshit.
#Google #crawler #RobotsTxt #UserAgent #bullshit #antisocial #web #search #WebSearch #LLM #AI
-
Ooh OpenAi ist gerade auf einer meiner Seiten unterwegs und ich wundere mich, warum gerade so viel Traffic auf dem Server ist
#Crawler -
Welcome to the future, where AI agents hunt down alleged online copyright infringementAs readers of this blog have doubtless noticed, the latest hot tech – and investment – area involves “agentic AI”, where AI systems are allowed to operative autonomously on allocated tasks. There’s no doubt there are some exciting possibilities here, as well as some troubling issues concerning lack of control. It’s a rapidly-evolving area of research and experimentation, which makes […]
#agenticAi #agents #ai #ceaseAndDesist #crawler #digitalWatermarks #infringement #licensing #llms #patents #pricing #takedowns #universalMusicGroup https://walledculture.org/welcome-to-the-future-where-ai-agents-hunt-down-alleged-online-copyright-infringement/ -
Welcome to the future, where AI agents hunt down alleged online copyright infringementAs readers of this blog have doubtless noticed, the latest hot tech – and investment – area involves “agentic AI”, where AI systems are allowed to operative autonomously on allocated tasks. There’s no doubt there are some exciting possibilities here, as well as some troubling issues concerning lack of control. It’s a rapidly-evolving area of research and experimentation, which makes […]
#agenticAi #agents #ai #ceaseAndDesist #crawler #digitalWatermarks #infringement #licensing #llms #patents #pricing #takedowns #universalMusicGroup https://walledculture.org/welcome-to-the-future-where-ai-agents-hunt-down-alleged-online-copyright-infringement/ -
To all #webmasters who use their service: This may be a quick fix for your #crawler woes. But you're not going to like the future they usher in. Your own #descendents will ask you one day why you ceded the control of this wonderful public resource to the likes of CF.
[3/4]
-
If I'm visiting a site from a country that you don't expect me to be from, does that mean that I'm not a human being interested in the content? Your solution to the AI vacuum cleaner is to arbitrarily blanket ban the IP blocks we're in? Why are we denied the full benefits of the internet because of your incompetence and/or unwillingness to solve the #LLM #crawler issue technically?
[2/4]
-
📬 Google-Ranking verstehen: Was hinter den Suchergebnissen steckt
#Empfehlungen #Internet #Absprungrate #Crawler #GoogleRanking #Keywords #SEO #Suchmaschinen #URLStrukturen https://sc.tarnkappe.info/d8af41 -
📬 Google-Ranking verstehen: Was hinter den Suchergebnissen steckt
#Empfehlungen #Internet #Absprungrate #Crawler #GoogleRanking #Keywords #SEO #Suchmaschinen #URLStrukturen https://sc.tarnkappe.info/d8af41 -
Turned people into crawlers today :V
-
Turned people into crawlers today :V
-
Found out about a project with millions of randomly generated links. The author explained how #Facebook's scraping bot hit it's page 38 million times. All while the company itself claims that their bot only crawls pages that are shared on their platforms.
Why is there so much dishonesty in some hyperscaling tech companies?
Other crawlers are also listed in a short write-up by the author.
-
Found out about a project with millions of randomly generated links. The author explained how #Facebook's scraping bot hit it's page 38 million times. All while the company itself claims that their bot only crawls pages that are shared on their platforms.
Why is there so much dishonesty in some hyperscaling tech companies?
Other crawlers are also listed in a short write-up by the author.
-
Who do you think you are?
47.128.32.0 - - [18/Mar/2026:00:48:01 +0100] "GET /robots.txt HTTP/1.1" 403 239 "-" "-" 1650 4269
Good on you that #CrowdSec won't immediately block on a missing user-agent, but my httpd-ACL does.
#DarkVisitors #AI #Crawler #GenAI #SocialPermissionToBurnEnergy
-
:ablobcatheartsqueeze: I have been running iocaine on my server for a week now. During this time, 7,076,701 requests have passed through iocaine, 3,312,318 of which were identified as AI crawlers/bots. 3,741,577 requests came from crawlers/bots that got stuck in iocaine's deadly maze, consuming an infinite amount of poisoned garbage. Furthermore, 972 crawlers/bots were detected that were routed into the maze via major browsers.
All of this is managed by iocaine with just ~80 MB of memory and ~0.1% direct CPU usage. Now that’s what I call efficient! Well done, @algernon.
Let's fight back against AI crawlers and bots. Thanks to projects like iocaine, this is entirely possible, not just theory :blobcat_thisisfine:
#iocaine #ai #llm #FckAI #FckLLMs #selfhosting #crawler #bots
-
Hallo liebe Fedinauten hier auf anonsys.net. Ab sofort wird diese Instanz vor AI- bzw. KI-Crawlern geschützt. Diese werden geblockt bzw. gebannt.
Danke @rainer für den Tipp. Habe diesen jetzt auf anonsys.net aktiviert und lasse den Filter einmal täglich aktualisieren.
Verdammt interessant ist, dass nach ca. 10 Minuten der Aktivierung des Filters bereits 128 AI-Crawler gebannt wurden:
Status for the jail: apache-ai-crawler |- Filter | |- Currently failed: 0 | |- Total failed: 33 | `- File list: /var/log/apache2/useragent.log `- Actions |- Currently banned: 128 |- Total banned: 128 `- Banned IP list: 100.28.204.82 100.29.160.53 107.20.181.148 119.28.140.106 18.207.89.138 18.214.124.6 18.215.24.66 18.215.49.176 18.232.11.247 18.235.158.19 184.73.167.217 184.73.239.35 216.73.216.43 23.21.179.120 23.21.225.190 2 3.21.227.240 23.21.228.180 23.23.99.55 3.209.174.110 3.212.205.90 3.212.86.97 3.220.148.166 3.221.244.28 3.222.190.107 3.93.211.16 3.93.253.174 34.192.67.98 34.195.248.30 34.205.163.103 34.225.138.57 34.226.89.140 34.227.234.246 34.230. 124.21 34.231.45.47 35.169.102.85 35.169.119.108 35.171.117.160 43.130.101.151 43.130.116.87 43.130.26.3 43.134.186.61 43.135.115.233 43.153.192.98 43.154.140.188 43.154.250.181 43.155.157.239 43.157.20.63 43.157.46.118 43.164.195.17 43 .164.196.57 43.164.197.224 43.165.135.242 43.165.189.206 43.166.128.86 43.166.242.189 43.166.244.66 44.194.134.53 44.205.74.196 44.209.35.147 44.210.213.220 44.213.202.136 44.217.255.167 44.220.2.97 44.221.105.234 44.223.116.180 47.128. 112.235 47.128.112.241 47.128.63.217 49.51.166.228 50.19.102.70 52.0.63.151 52.2.4.213 52.201.155.215 52.203.237.170 52.4.229.9 52.5.232.250 52.54.157.23 52.6.97.88 52.70.123.241 54.145.82.217 54.147.80.137 54.157.84.74 54.159.18.27 54. 235.172.108 54.83.23.103 54.83.240.58 54.83.56.1 66.249.68.128 66.249.68.130 98.82.38.120 98.82.63.147 98.82.66.172 98.83.10.183 98.83.8.142 98.84.60.17 18.208.11.93 18.214.238.178 3.218.35.239 44.212.131.50 54.157.99.244 3.230.69.161 1 8.235.81.246 52.203.152.231 35.173.38.202 3.232.82.72 34.193.2.57 54.166.126.132 3.225.9.97 98.82.39.241 98.84.200.43 3.94.156.104 44.223.115.10 43.163.104.54 43.157.22.109 43.130.131.18 43.131.26.226 49.51.132.100 50.16.248.61 43.155.1 62.41 52.203.68.145 54.89.90.224 34.236.185.101 52.200.251.20 43.166.224.244 98.82.107.102 129.226.174.80 18.205.213.231 34.204.150.196Es werden minütlich mehr. Das ist echt Wahnsinn! 😳
Quelle: rainer.sokoll.com/?p=8353
-
I am looking for a nice tool that I could run on my home server to poison my internet useage pattern.
So far I could only find some outdated projects...
Do you have any recommendations?
-
Wie KI die Art und Weise, wie wir Inhalte finden, neu definiert
Die Art und Weise, wie Menschen Informationen online finden, ändert sich schnell. Da Künstliche Intelligenz (KI) zu einem Kernbestandteil davon wird, wie Benutzer Inhalte entdecken, müssen Ihre Inhalte härter und intelligenter arbeiten, um gesehen zu werden.
https://clearleft.com/thinking/how-ai-is-redefining-the-way-we-find-content
#Crawler #Information #Inhalt #KI #KIBots #KünstlicheIntelligenz #SEO #GEO #Suche #Suchmaschine #Optimieren
-
Robots.txt Generator - Retro Terminal Edition - Mehr als 200 Bots in der kostenfreien Version. Pures HTML, Javascript und ein bisschen CSS. Keine Third Parties, kein Framework, kein CDN, keine Cookies, kein Tracking, keine Werbung, kein BigTech-Gedönse, keine KI, sehr datenschutzfreundlich. Simple und effektiv im Retro-Style. Demnächst online.
#teufelswerk #HTML #javascript #app #entwicklung #code #retro #css #robotstxt #generator #stopbots #bots #crawler #scraper #keineKI #cookieless #datenschutz