home.social

#scrape — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #scrape, aggregated by home.social.

fetched live
  1. Every #scrape, #site, range and page; every game, download, hack, song, movie and virrie on the Web. Everything on your #phone. Everything on your 'puta. Even the content directories of your cupboards. Almost every #system has been brute-forced; passwords cracked, firewalls breached. Nothing has been left untouched.

    L. Ashley Straker, Connected Infection

    #quote #quotes #cellphones #it #infotech #informationtechnology #code #coding #hack #hacks #hacking #privacy #invasionofprivacy #humanrights

  2. #Landlords Demand Tenants’ #Workplace #Logins to #Scrape Their #Paystubs

    Landlords are using a service that logs into a potential renter’s #employer systems and scrapes their paystubs and other information en masse, potentially in violation of U.S. #hacking laws, according to screenshots of the tool shared with 404 Media.
    #privacy #security #payroll #tenants

    404media.co/landlords-demand-t

  3. It bothers me when people and organizations back the "stealing content to train 'AI' is fair use" argument. To me, it seems pretty clearly *not* fair use. But these orgs frequently back their positions with something along the lines of "But AI can be a positive, ethical force, we have to be able to create ethical AI".

    I don't have a problem acknowledging that ethical uses of "AI" are possible.

    I have a problem with "AI" backers not acknowledging that the only people funding "AI" have no interest in ethical uses of it.

    #AI #ethics #ethical #FairUse #training #LLM #BigTech #scrape #slop #theft #copyright #funding

  4. When it comes to #AI, could you tf not with all that #scraping? Pay-per-packet could be the future now if some people can't control themselves.

    The more you #scrape, the more #developers have to pay, which should yield better and improved infrastructure to decrease the cost, but instead: it could turn the internet into a true "transactional" network.

    Suddenly the #internet is run on #microtransactions... hell hath arrived. Granted, this fringe scenario is a bit hyperbolic, but still.

  5. "Run!" yelled the boys, and all three took to their heels.
    "Look at you!" Riva's mother said to her. "I can't even let you wash clothes without you getting into a #scrape."
    "But I didn't do anything!" Riva cried.
    "If you didn't tease the boys, this wouldn't happen," her mother said, savagely. "I'm sure your father will have something to say to you when he gets home. Now quit just sitting around and finish the laundry!"
    Riva clenched her fists, but then, sullenly, went back to work.

    #wss366

  6. Hey, #webmasters ... just so you know.

    #Facebook's new-ish "meta-externalagent" #webcrawler, which they document is for stealing data for their Grand Theft Autocomplete (cough #AI cough), is ignoring robots.txt on my websites.

    developers.facebook.com/docs/s

    Is anyone surprised?

    #Meta #LLM #scrape #web #copyright #RobotsTXT

  7. Squirrel is rather persistent in fights. Perhaps too much so. He got into a fight with an alpha arcanine and proceeded to get bit around his midsection. So now he's recovering and rather bored!

    Art by furaffinity.net/user/haychel/

    Drackal is linktr.ee/Ra_Zim 

    #furry #furryart #wounded #dirty #scrape #scraped #injuries #lockpick

  8. Sagt mal, ein Webserver Endpoint bei dem man per GET eine Suche machen kann:
    Gibts da ne ausgefuchste Art ne Liste dessen unterstützter Parameter zu bekommen? 🤔

    Oder geht das nur über Browser-Bedienung der Filter und beobachten der Webrequests? 👀

    #webdevelopment #parser #scrape

  9. annas-archive.org/

    I can't believe that it took me almost two years to discover #Anna's #Archive. It is the largest open digital library. They #scrape and #opensource plenty of resources, and mirror the already existing ones, such as #SciHub and #LibGen.

    Their two goals:
    1. Backing up all knowledge and culture of humanity.
    2. Making this knowledge and culture available to anyone in the world.

    Huge respect for the digital pirates behind this project for their dedication to the community.

  10. Hat jemand Erfahrung mit dem #homeassistant Dienst #scrape? Habe einen Auftrag bei #Wachete laufen, dort klappt es super, bekomme es aber nicht in scrape hin.

  11. So you’re posting on Mastodon, enjoying your digital life away from most of the tech blunders of integrating AI into everything for no reason. Your words are your own, you’ve got your robots.txt files organized nicely in your root directories, and everything is peaceful.

    Life is good when corporations aren’t trying to plagiarize everything everyone’s ever written.

    Enter: Technology Worker Dude # 4,104,210, or, Maven.

    Sometime earlier today, on June 12th, 2024, an OpenAI backed project, or social network with no likes or faves and only people replying, that is also heavily infested with AI generated crap, used some code to ingest the whole ActivityPub fediverse. Just like that. Without mentioning it publicly, without uttering a single word.

    Just sucked it all up like a vacuum fed into a garbage disposal.

    Except that garbage disposal is the network feed of a website that is definitely not federating over ActivityPub.

    As quoted in this article:

    In addition to pulling in posts, the import process seems to be running AI sentiment analysis to add tags and relational data after content reaches Maven’s servers. This is a core part of Maven’s product: instead of follows or likes, a model trains itself on its own data in an attempt to surface unique content algorithmically.

    It’s worth mentioning that Maven received 2 million dollars in funding from former Twitter CEO Ev Williams and OpenAI CEO Sam Altman.

    Sean Tilley, wedistribute.org

    If you can make heads or tails of what exactly this guy’s AI was doing with all of our posts from Mastodon, I applaud you. Because it’s complete gibberish to me.

    Taking out the dopamine feeding parts of a social network is already kind of a weird decision, but building one based around violating consent, I mean, I guess we all saw this coming. As I’ve said before, a lot of the tech world hates consent.

    The people behind this have since halted ingestion and deleted everything that was scraped, for now. They haven’t said they weren’t going to do it again the second nobody’s looking.

    It’s clear from the feedback on this thread that even our experiments with the tech were confusing to users and didn’t fit with other people’s expectations of how it should work.

    We are currently pausing this integration, at least until we can better understand how Maven can fit in as a good citizen of the Fediverse.

    Jimmy Secretan, CTO, heymaven.com

    It takes a special kind of AI-muddled thought process to, instead of spending five minutes investigating how ActivityPub works, you just hook up some AI and download the whole damn thing. And then act surprised when people rightfully tell you that’s not how it works, and you can’t just shove an AI into an open space and expect a “Thank you.”

    But this isn’t really a surprise, since they advertise AI-scraped sludge directly on their homepage.

    “I believe that Midjourney is a great way to get started learning about generative art.”

    You mean, a great way to get started being a pariah that everyone hates. Sure. You put that on your main page as a focal point.

    I can’t emphasize enough how much I would love if all the data centers containing the code running these things, across every network, just suddenly exploded. Take it all back to zero, and then put up a digital wall, like in Cyberpunk 2077 when they built a whole new internet that isn’t infested with garbage.

    But it seems, this will just continue to be a constant fight against greed, the death of creativity, and Sam Altman.

    https://cmdr-nova.online/2024/06/13/hey-its-maven-whos-maven/

    #mastodon #maven #samAltman #scrape #socialMedia

  12. From a location on Google Maps and a metadata tag (e.g. restaurant) this service retrieves all the information present in the map view and exports it in JSON, CSV or Excel #scrape

    scrapetable.com/

  13. #DailyBloggingChallenge (153/200)

    There are two main ways to #scrape a #website, either actively or passively.

    Active scraping is the process of using a trigger to actively scrape the already loaded webpage.

    Passive scraping is the process of having the tool navigate to the webpage and scrape it.

    The main difference is how one is getting to the loaded #webpage.

    #WebsiteScraping

  14. Scrapers. Spiders. Crawlers. Whatever you call them, they're are extremely versatile tools for collecting website data quickly and efficiently. Here's how to crawl a site without getting blocked.

    #crawl #scrape #spider #website
    tchlp.com/3t0eRSw

  15. @MattBinder In #Canada recent legislation ( #billC18 ) that came into force basically said social media sites and search engines can't link directly to legitimate *Canadian* #news sources without paying or negotiating some #compensation directly to those news sources. Seems fair, I mean they #scrape the articles and shove them in your faces with their own ads, stripping off the news site's ads.

    Facebook and Google responded with a tantrum. "Pay? For news we scrape? FINE, NO NEWS FOR ANYONE". 1/

  16. Looking at practices in the wild is a ride. Trying to websites, just looking at headers I've seen:

    - Divs with id=header or id=site-header
    - Multiple header elements with duplicated logos etc., changing which one is visible with JavaScript based on screen size
    - Divs with absolutely no way to identify that it's the header, just CSS to put it at the top of the page
    - A search form that is visually positioned in the header but is actually outside the header in the DOM