#webdata — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #webdata, aggregated by home.social.
-
Warburg Pincus has invested $130 million in Oxylabs, valuing the Lithuanian company at $3.6 billion.
The bigger story is the growing importance, and growing controversy, of the infrastructure that gives AI access to the live web.
#AI #technology #webdata #web #investment #startups #lithuania #Europe #EU
-
Warburg Pincus has invested $130 million in Oxylabs, valuing the Lithuanian company at $3.6 billion.
The bigger story is the growing importance, and growing controversy, of the infrastructure that gives AI access to the live web.
#AI #technology #webdata #web #investment #startups #lithuania #Europe #EU
-
Some teams spend time searching for data. The most effective teams start with reliable access from day one.
At TagX, we make trusted, ethically sourced data accessible, so your teams can focus on building, analyzing, and growing.
Because every successful strategy begins with better data and a trusted partner.
#TagX #WebData #DataAnalytics #MarketResearch #ConsumerInsight #BusinessGrowth
-
Some teams spend time searching for data. The most effective teams start with reliable access from day one.
At TagX, we make trusted, ethically sourced data accessible, so your teams can focus on building, analyzing, and growing.
Because every successful strategy begins with better data and a trusted partner.
#TagX #WebData #DataAnalytics #MarketResearch #ConsumerInsight #BusinessGrowth
-
Why configure your AI Model Harness? #webscraping #podcast #ai #programming https://www.youtube.com/watch?v=xdVcK-XfxLo?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Why configure your AI Model Harness? #webscraping #podcast #ai #programming https://www.youtube.com/watch?v=xdVcK-XfxLo?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Keep your context window clean using these #webscraping #ai #podcast https://www.youtube.com/watch?v=wNbLzW1huZM?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Keep your context window clean using these #webscraping #ai #podcast https://www.youtube.com/watch?v=wNbLzW1huZM?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
No-one likes an out-of-touch AI assistant. Fortunately, rapid refreshing can keep AI models aware of the very latest public information. https://www.zyte.com/blog/enhancing-ai-model-performance-with-fresh-web-data?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
When you can scrape the web by API, a world of possibility opens up. Yes, you can extract live web data using iOS Shortcuts. https://www.zyte.com/blog/web-scraping-on-an-iphone?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
AI models are only as good as the data they learn from. Explore how web data collection supports AI training by providing large-scale, diverse, and structured datasets for machine learning and generative AI projects. Ideal for businesses, AI developers, and data scientists.
Read the full guide:
https://www.webscreenscraping.com/ai-training-data-collection-using-web-scraping/#AITrainingData #MachineLearning #WebData #DataCollection #AIInnovation
-
Developers are embracing agentic coding tools - but data engineers need tools with specialist scraping skills. https://www.zyte.com/blog/introducing-zyte-web-data-for-claude-code?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Thunderbit Rolls Out New Tools For Web Data Sifting
Thunderbit launched new tools like an API and CLI to help developers turn web page content into usable data for AI and automation. Learn how it works.
#WebData, #DeveloperTools, #AI, #DataExtraction, #Thunderbit
https://newsletter.tf/thunderbit-new-tools-for-developers-web-data/
-
Thunderbit Rolls Out New Tools For Web Data Sifting
Thunderbit launched new tools like an API and CLI to help developers turn web page content into usable data for AI and automation. Learn how it works.
#WebData, #DeveloperTools, #AI, #DataExtraction, #Thunderbit
https://newsletter.tf/thunderbit-new-tools-for-developers-web-data/
-
Thunderbit's new tools aim to make web data easier to use for AI, with a new engine achieving a high score in tests for converting web pages to Markdown.
#WebData, #DeveloperTools, #AI, #DataExtraction, #Thunderbit
https://newsletter.tf/thunderbit-new-tools-for-developers-web-data/ -
Thunderbit's new tools aim to make web data easier to use for AI, with a new engine achieving a high score in tests for converting web pages to Markdown.
#WebData, #DeveloperTools, #AI, #DataExtraction, #Thunderbit
https://newsletter.tf/thunderbit-new-tools-for-developers-web-data/ -
Your VPS is ready, but now you need to work through the same sequence you have run a dozen times before: apt update, apt install python3-pip, pip install scrapy, playwright install chromium, the Chromium dependency list that never installs cleanly on the first try, Redis, possibly Postgres, whatever else this particular project needs. https://www.zyte.com/blog/automate-deployment-of-your-web-scraper-on-vps-with-ubuntu-24-04-cloud-init?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Multi-agent orchestration is having its moment. The diagrams are everywhere now. Boxes for planners, boxes for hands, boxes for daemons, arrows to a shared brain, a human floating at the top. They keep getting prettier. The part where the web pushes back is still the part nobody draws. https://www.zyte.com/blog/multi-agent-orchestration-in-a-large-scale-web-scraping-project?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Learn how rotating proxies work, when to use them for web scraping, and why IP rotation alone is not enough for reliable data access. https://www.zyte.com/learn/how-do-rotating-proxies-work?utm_campaign=learn-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Learn what residential proxies are, how they compare to datacenter proxies, and why modern web scraping needs more than IP diversity. https://www.zyte.com/learn/what-is-a-residential-proxy?utm_campaign=learn-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Learn how much rotating proxies cost, what affects pricing, and why total web scraping costs often go beyond proxy subscriptions. https://www.zyte.com/learn/how-much-do-rotating-proxies-cost?utm_campaign=learn-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Learn whether proxies are legal for web scraping, what compliance factors matter, and why modern teams use automated unblocking beyond proxy management. https://www.zyte.com/learn/are-proxies-legal-for-web-scraping?utm_campaign=learn-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Consign bill-shock to the trashcan. New custom spending limits and usage insights put data-gatherers in control. https://www.zyte.com/blog/new-spending-controls-and-usage-insights-for-zyte-api?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Monitor your data-gathering pipelines like a boss - and act on domain issues in real-time. https://www.zyte.com/blog/zyte-domain-health-hub?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
The problem was a legacy project with 12,000 websites to crawl, and there’s no world where you write custom spiders for 12,000 websites, not with a human team and certainly not sustainably.
So Javier built a workflow: a set of AI prompts that could analyze a website, figure out its structure, and generate a crawl configuration that a generic spider could then use. https://www.zyte.com/blog/not-an-interview-i-have-superpowers-now?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Deploy your web scraper on any cloud vendor in under 2 minutes | IaaC https://www.youtube.com/watch?v=MivnbzBVFTQ?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
If you want to understand exactly how a browser scraping service works at the infrastructure level, or you have a steady workload that you want running on hardware you already own, building one yourself teaches you things that matter. Here's how I did it https://www.zyte.com/blog/building-a-self-hosted-browser-scraping-service-is-it-more-hassle-than-its-worth?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Web scraping devs discuss LLMs (good AND bad) https://www.youtube.com/watch?v=GH8cKsu4_Tk?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Data-gathering doesn’t have to be memory-intensive. You can fit the world’s weather on a 9cm-square board, when you move the work to a web scraping API. https://www.zyte.com/blog/web-scraping-on-22-kb-of-ram-fitting-the-world-on-an-esp-8266-microcontroller?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
For the last 30 days, I did one thing almost exclusively: I built scraping systems with AI agents, from the ground up, across real targets, with real deadlines. Not prototypes designed to impress in a demo, not isolated experiments running against a toy website, but production-grade pipelines that needed to ship and keep running. https://www.zyte.com/blog/i-built-scraping-agents-for-30-days-heres-what-i-learned?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
I've been running a series of conversations with developers at Zyte to understand what's actually changed in the way they work since LLMs showed up. Not the headlines. The day-to-day. What they delegate, what they don't, what they notice, what surprises them.
This one was different on two counts. https://www.zyte.com/blog/not-the-same-developer-i-was-before-llms?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
the next time you spin up a VPS to give it a persistent home, you spend the better part of an afternoon rebuilding from memory: installing Scrapy, wiring up Redis, configuring the systemd units, getting Playwright's Chromium dependencies in the right state. Here's a tool to help https://www.zyte.com/blog/flatcar-linux-for-web-scrapers-deploying-containers?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Programmers were raised on long-standing core principles of the craft. What if those tenets are no longer relevant? https://www.zyte.com/blog/are-programming-best-practices-still-relevant-ai?utm_campaign=blog-postsutm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Discover how autonomous, agent-driven data pipelines are transforming web scraping in 2026, enabling self-healing systems, API discovery, and end-to-end automation. https://www.zyte.com/blog/dawn-of-the-autonomous-data-pipeline?utm_campaign=blog-postsutm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Ayan's 4 agent team, using Claude's /goal, and the models and coding agents he uses to code effectively. https://www.zyte.com/blog/my-agentic-coding-setup-claude-code-multi-agent-orchestration-and-how-i-actually-work?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
Marketers are giving up on the idea of plain-text pages - but llms.txt and Markdown are how we’ll get our docs in the hands of LLMs and developers. https://www.zyte.com/blog/how-we-put-dev-docs-in-ai-spotlight?utm_campaign=blog-posts&utm_activity=ORS&utm_medium=social&utm_source=mastodon
-
#Reproducibility #OpenScience #ComputationalSocialScience #WebData #DigitalBehavior #MLResearch #NLPResearch #SocialMediaData
#CallForPapers: Our full-day workshop at The Web Conference 2026 #TheWebConf2026 invites submissions on reproducible and reusable computational approaches for social and web data.
👉 https://easychair.org/cfp/r2cass2026Deadline: Dec 18th, 2025
-
Discover the three best, most modern methods to access and harness web data for your projects. https://hackernoon.com/need-web-data-here-are-the-3-methods-everyones-using #webdata
-
How to Build an Efficient Data Team to Work with Public Web Data - The topic of how to assemble an efficient data team is a highly debated and freque... - https://readwrite.com/how-to-build-an-efficient-data-team-to-work-with-public-web-data/ #bigdataanalytics #dataandsecurity #dataengineering #dataanalytics #publicwebdata #datascience #datateam #bigdata #webdata #tech #work