home.social

#webdata — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #webdata, aggregated by home.social.

fetched live
  1. Warburg Pincus has invested $130 million in Oxylabs, valuing the Lithuanian company at $3.6 billion.

    The bigger story is the growing importance, and growing controversy, of the infrastructure that gives AI access to the live web.

    movetheneedle.news/brands/lith

    #AI #technology #webdata #web #investment #startups #lithuania #Europe #EU

  2. Warburg Pincus has invested $130 million in Oxylabs, valuing the Lithuanian company at $3.6 billion.

    The bigger story is the growing importance, and growing controversy, of the infrastructure that gives AI access to the live web.

    movetheneedle.news/brands/lith

    #AI #technology #webdata #web #investment #startups #lithuania #Europe #EU

  3. Some teams spend time searching for data. The most effective teams start with reliable access from day one.

    At TagX, we make trusted, ethically sourced data accessible, so your teams can focus on building, analyzing, and growing.

    Because every successful strategy begins with better data and a trusted partner.

    #TagX #WebData #DataAnalytics #MarketResearch #ConsumerInsight #BusinessGrowth

  4. Some teams spend time searching for data. The most effective teams start with reliable access from day one.

    At TagX, we make trusted, ethically sourced data accessible, so your teams can focus on building, analyzing, and growing.

    Because every successful strategy begins with better data and a trusted partner.

    #TagX #WebData #DataAnalytics #MarketResearch #ConsumerInsight #BusinessGrowth

  5. AI models are only as good as the data they learn from. Explore how web data collection supports AI training by providing large-scale, diverse, and structured datasets for machine learning and generative AI projects. Ideal for businesses, AI developers, and data scientists.

    Read the full guide:
    webscreenscraping.com/ai-train

    #AITrainingData #MachineLearning #WebData #DataCollection #AIInnovation

  6. Thunderbit Rolls Out New Tools For Web Data Sifting

    Thunderbit launched new tools like an API and CLI to help developers turn web page content into usable data for AI and automation. Learn how it works.

    #WebData, #DeveloperTools, #AI, #DataExtraction, #Thunderbit

    newsletter.tf/thunderbit-new-t

  7. Thunderbit Rolls Out New Tools For Web Data Sifting

    Thunderbit launched new tools like an API and CLI to help developers turn web page content into usable data for AI and automation. Learn how it works.

    #WebData, #DeveloperTools, #AI, #DataExtraction, #Thunderbit

    newsletter.tf/thunderbit-new-t

  8. Thunderbit's new tools aim to make web data easier to use for AI, with a new engine achieving a high score in tests for converting web pages to Markdown.

    #WebData, #DeveloperTools, #AI, #DataExtraction, #Thunderbit
    newsletter.tf/thunderbit-new-t

  9. Thunderbit's new tools aim to make web data easier to use for AI, with a new engine achieving a high score in tests for converting web pages to Markdown.

    #WebData, #DeveloperTools, #AI, #DataExtraction, #Thunderbit
    newsletter.tf/thunderbit-new-t

  10. Your VPS is ready, but now you need to work through the same sequence you have run a dozen times before: apt update, apt install python3-pip, pip install scrapy, playwright install chromium, the Chromium dependency list that never installs cleanly on the first try, Redis, possibly Postgres, whatever else this particular project needs. zyte.com/blog/automate-deploym

    #webscraping #webdata #data #web

  11. Multi-agent orchestration is having its moment. The diagrams are everywhere now. Boxes for planners, boxes for hands, boxes for daemons, arrows to a shared brain, a human floating at the top. They keep getting prettier. The part where the web pushes back is still the part nobody draws. zyte.com/blog/multi-agent-orch

    #webscraping #webdata #data #web

  12. Learn whether proxies are legal for web scraping, what compliance factors matter, and why modern teams use automated unblocking beyond proxy management. zyte.com/learn/are-proxies-leg

    #webscraping #webdata #data #web

  13. The problem was a legacy project with 12,000 websites to crawl, and there’s no world where you write custom spiders for 12,000 websites, not with a human team and certainly not sustainably.

    So Javier built a workflow: a set of AI prompts that could analyze a website, figure out its structure, and generate a crawl configuration that a generic spider could then use. zyte.com/blog/not-an-interview

    #webscraping #webdata #data #web

  14. If you want to understand exactly how a browser scraping service works at the infrastructure level, or you have a steady workload that you want running on hardware you already own, building one yourself teaches you things that matter. Here's how I did it zyte.com/blog/building-a-self-

    #webscraping #webdata #data #web

  15. For the last 30 days, I did one thing almost exclusively: I built scraping systems with AI agents, from the ground up, across real targets, with real deadlines. Not prototypes designed to impress in a demo, not isolated experiments running against a toy website, but production-grade pipelines that needed to ship and keep running. zyte.com/blog/i-built-scraping

    #webscraping #webdata #data #web

  16. I've been running a series of conversations with developers at Zyte to understand what's actually changed in the way they work since LLMs showed up. Not the headlines. The day-to-day. What they delegate, what they don't, what they notice, what surprises them.

    This one was different on two counts. zyte.com/blog/not-the-same-dev

    #webscraping #webdata #data #web

  17. the next time you spin up a VPS to give it a persistent home, you spend the better part of an afternoon rebuilding from memory: installing Scrapy, wiring up Redis, configuring the systemd units, getting Playwright's Chromium dependencies in the right state. Here's a tool to help zyte.com/blog/flatcar-linux-fo

    #webscraping #webdata #data #web

  18. Discover how autonomous, agent-driven data pipelines are transforming web scraping in 2026, enabling self-healing systems, API discovery, and end-to-end automation. zyte.com/blog/dawn-of-the-auto

    #webscraping #webdata #data #web

  19. Marketers are giving up on the idea of plain-text pages - but llms.txt and Markdown are how we’ll get our docs in the hands of LLMs and developers. zyte.com/blog/how-we-put-dev-d

    #webscraping #webdata #data #web

  20. Discover the three best, most modern methods to access and harness web data for your projects. hackernoon.com/need-web-data-h #webdata