MikeTrendsTrends right now

search

Scrape

Trends

  1. 1
    OpenAI and Microsoft researchers warn of AI 'doom loop' consuming the webβ–Όβ€˜Doom Loop’: OpenAI and Microsoft Admits LLMs Are Destroying the Web and Built on TheftMmastodonTechnologyAI855 d ago

    Researchers at OpenAI and Microsoft have published a paper describing a 'doom loop' in which large language models, trained on data scraped largely without consent from human creators, degrade the open web and eventually poison their own training data. The paper reportedly acknowledges that people will come to see the models' wholesale ingestion of creative work as an unprecedented act of theft, reigniting debate over copyright and the sustainability of generative AI.

  2. 2

    A Microsoft director has described AI training data scraping as 'the largest theft of labor in human history', while OpenAI's head of publisher strategy called ChatGPT an existential threat to publishers. The remarks were revealed in legal briefs filed in the New York Times' copyright lawsuit against OpenAI and Microsoft. The comments add to the debate over whether AI companies unfairly exploit journalists' and creators' work without permission or payment.

  3. 3
    OpenAI and Microsoft reportedly knew of 'doom loop' for the web●OpenAI and Microsoft knew they were starting a 'doom loop' for the webYhnTechnologyInternet357 d ago

    A new report claims OpenAI and Microsoft were aware that their AI products, trained partly on scraped web content, could trigger a 'doom loop' degrading the open web. According to The Verge, internal discussions suggested the companies understood that AI-generated content and content theft could crowd out human-made material, drawing comparisons to earlier warnings about Google's zero-click search results.

  4. 4
    OpenAI-linked AI agents scanned UN statistics site thousands of times●AI agents linked to OpenAI scanned a UN statistics website more than 16,500 times between April and June, using proxiesMmastodonWorld26 d ago

    AI agents linked to OpenAI accessed a UN statistics website more than 16,500 times between April and June, using proxies to bypass API limits, according to a researcher who documented the activity on a UN data platform. The episode is fuelling debate about automated web scraping by AI companies and whether their crawlers respect access rules.

  5. 5
    AI firm reportedly hammered UN data site to bypass limits●An AI company was hammering a UN data website with creative ways to get around the limitations of its API and interface.MmastodonWorldUnited Nations45 d ago

    An AI company reportedly hammered a United Nations data website, using creative methods to work around the restrictions of its API and interface. Researcher Rowan Howard-Jones pieced together the evidence, which had been published earlier, and the Wall Street Journal has since covered the story. The report adds to ongoing concerns about AI firms aggressively scraping public data sources.

  6. 6
    OpenAI agents scanned UNCTAD site 16,000 times, researcher says●Security researcher Rowan Howard-Jones says that OpenAI agents scanned the UN Conference on Trade and Development's (UNCMmastodonTechnologyAI36 d ago

    Security researcher Rowan Howard-Jones reports that OpenAI's automated agents hit the UN Conference on Trade and Development's statistics site more than 16,000 times between April and June. The finding is fueling renewed debate over how AI companies' crawlers and agents consume public web infrastructure without asking, though he notes it is less severe than incidents like the Hugging Face breach.

  7. 7
    Researcher links 16,000 scans of UN statistics portal to OpenAI agentsβ–ΌResearcher links 16,000 scans of a UN statistics portal to OpenAI agentsβœ‰newsWorldUnited Nations6 d ago

    A security researcher has linked roughly 16,000 scans of a United Nations statistics portal to OpenAI's AI agents, suggesting automated browsing by the company's systems at significant scale. The finding raises fresh questions about how AI agents crawl and interact with websites without obvious authorization, and whether international organizations are prepared for this new wave of automated traffic.

  8. 8
    Claims of Coordinated Campaign Against Glass Skyscrapers in India●I am seeing a concerted campaign against glass skyscrapers in India through multi channel targeting of different groups𝕏xIN4.3K2 d ago

    A claim is circulating that a coordinated, multi-channel effort is being run against glass skyscrapers in India, targeting environmentalists and animal lovers in particular. The person making the claim characterises it as a possible psyop intended to disrupt India's construction of tall buildings. No evidence is offered to substantiate the alleged campaign.

  9. 9
    Elephants mine minerals deep inside Kenya's Kitum Caveβ—πŸ˜ WATCH: In the pitch-black chambers of # MountElgon 's # Kitum Cave, African # savannah # elephants use their tusks toMmastodonScience182 d ago

    African savannah elephants venture into the pitch-black chambers of Kitum Cave on Mount Elgon in Kenya, using their tusks to scrape volcanic rock for sodium, calcium and magnesium. Generations of herds have hollowed out these deep mineral chambers, passing the unusual mining behaviour down through the herd. Footage of the elephants at work is drawing attention to one of nature's most remarkable learned traditions.

  10. 10
    Meta's Muse model praised for web scraping tasks●Meta's Muse is fantastic for web scrapingYhnBusinessLabor6041 min ago

    A new write-up argues that Meta's Muse model performs impressively well at web scraping, drawing attention among developers evaluating AI tools for automated data extraction. The piece has sparked discussion about how capable Meta's model is at practical, unglamorous tasks like parsing and extracting information from web pages.

  11. 11
    Substack accused of obfuscating text to break reader mode●Tell HN: Substack obfuscating text to break reading modeYhnBusinessCrypto204 d ago

    Substack is being accused of deliberately obfuscating its article text so that browser-based reader modes cannot cleanly parse its pages, forcing visitors to read on the site itself. Commenters are debating whether the practice is an anti-scraping measure or a hostile move against accessibility tools, with many calling it anti-user and comparing it to similar tactics on other publishing platforms.

  12. 12
    EU Commission accused of sacrificing children's privacy to AI firms●RE: https:// mstdn.social/@stux/11737340528 4945562 also, @EUCommission wants the world to offer their children’s privacMmastodonTechnologyAI81 d ago

    The European Commission is facing criticism over its approach to AI companies and data scraping. Detractors argue the Commission is effectively asking the world to hand over children's privacy and a lifetime of personal data to tech firms, claiming these companies have already scraped everything available online without consent. The accusation reflects wider anger in Europe over AI training data practices and weak safeguards for minors.

  13. 13
    Reddit is killing RSS feeds and public API access over AI botsβ–ΌReddit is killing RSS feeds and ending public API access because of AI bots | TechCrunchMmastodon1481 d ago

    Reddit is ending support for RSS feeds as it continues restricting access to its archive of user-generated content, with TechCrunch reporting the move is aimed at stopping AI companies from scraping the platform to train their models. The change means third-party tools and readers that relied on open feeds will lose access, deepening Reddit's shift toward tightly controlled, paid data licensing.

  14. 14
    Elephants mine minerals deep inside Mount Elgon's Kitum Caveβ—πŸ˜ ICYMI: In the pitch-black chambers of # MountElgon 's # Kitum Cave, African # savannah # elephants use their tusks toMmastodonHealth2328 min ago

    African savannah elephants venture into the pitch-black chambers of Kitum Cave on Mount Elgon, using their tusks to scrape volcanic rock rich in sodium, calcium and magnesium. Generations of herds have hollowed out these deep mineral chambers, passing the unusual mining behaviour down through the generations.

  15. 15

    Le Parisien has published a report on households in France who say they live 'day to day on a workaround', scraping by to the nearest euro as the cost of living erodes their budgets. The piece details how families cut spending, juggle bills and delay purchases to make ends meet, fuelling renewed public debate about purchasing power, one of the most sensitive issues in French politics.

  16. 16
    UK Media Mocked for Reporting on Someone's Lunch●Why is someone’s lunch newsworthy??? Slow day in the UK?𝕏xGB21 d ago

    People in the UK are questioning why a national news outlet ran a story about an ordinary person's lunch, with some calling it a sign of a slow news day. The reaction online has been largely mocking, as readers express disbelief that such a mundane topic warranted coverage, and debate whether the media is scraping the barrel for content.

  17. 17
    Chinese AI model accused of taking developers' code●A Chinese AI model is stealing your codeβ–ΆyoutubeTechnologySoftware94.6K3 d ago

    A claim is circulating among developers that a Chinese AI model is 'stealing' code, suggesting it may be trained on or reproduce programmers' work without permission. The warning has drawn large attention in the software community, where concerns about code provenance, licensing and data scraping by AI companies are already running high. No specific model or proof is named in the discussion, so the accusation remains unverified.

  18. 18

    A technical article from TrueSign examines common methods used to detect residential proxies and argues they frequently fail. The piece walks through flawed detection approaches and explains why they produce poor results, drawing attention from developers and anti-fraud engineers on Hacker News, where it reached a top-10 position amid ongoing debate over bot detection and proxy evasion techniques.

  19. 19
    Philadelphia Inquirer launches Scrape, an AI tool for hyperlocal newsβ–ΌThe Philadelphia Inquirer built Scrape, an AI tool to surface hyperlocal newsYhnEnvironment141 h ago

    The Philadelphia Inquirer has built Scrape, an artificial intelligence tool designed to surface hyperlocal news stories that might otherwise go unreported. Developed with support from the Lenfest Institute, the tool is aimed at helping the newsroom cover neighbourhood-level events and community issues more efficiently. The project is drawing attention as an example of how local news organisations are experimenting with AI to expand coverage despite shrinking resources.

  20. 20
    Reddit restricts Old Reddit access citing AI bots●Reddit says it has to cut back access to β€˜Old Reddit’ because of AI bots Reddit is further limiting who can use the "OldMmastodonTechnology33 d ago

    Reddit says it is further limiting access to its legacy 'Old Reddit' interface, citing efforts to combat scraping and automated traffic from AI bots. The company has recently begun requiring users to log in before using the classic experience. Many long-time users rely on Old Reddit for its simpler design, and the tightening rules have sparked frustration among those who see it as another step toward restricting free access to the platform.

  21. 21
    AI bots and scraping could push the open web out of reach●Anno 202X, il web Γ¨ cosΓ¬ impestato e saccheggiato da bot e scraping dell'AI che i costi per la gestione di siti e comuniMmastodonTechnologyInternet215 h ago

    Commenters warn that by 202X the web could be so infested with bots and AI scraping that running websites and online communities becomes affordable only for a handful of global billionaires, pushing humanity into isolated, self-contained systems. The post is drawing attention in Italian-language discussions about the rising costs of maintaining an open internet under pressure from automated AI traffic.

  22. 22
    Reddit ends RSS feeds and public API access, citing AI bots●Reddit is killing RSS feeds and ending public API access because of AI bots https:// lemmy.zip/post/72580572MmastodonTechnology421 h ago

    Reddit is shutting down its RSS feeds and ending public API access, attributing the move to the flood of AI bots scraping its content. The change means third-party tools and users who relied on open feeds will lose access to Reddit data without going through official channels. Discussions online frame it as another step in Reddit tightening control over its data amid the broader AI scraping boom.

  23. 23
    Google Maps Scraper MCP Debuts for Developers●Show HN: Google Maps Scraper MCPYhnBusiness951 min ago

    A new tool called Google Maps Scraper MCP has been introduced, designed to let AI assistants and applications extract business data from Google Maps, such as locations, contact details and reviews. The launch is drawing attention from developers and businesses interested in automated location data collection, as well as those weighing the legal and practical limits of scraping Google's mapping services.

  24. 24
    Reddit to shut down RSS feeds in 2026β–ΌReddit is Killing RSS Support https://www. reddit.com/r/modnews/s/kvHTs7I xpg "Because RSS is now a common surface for lMmastodonTechnology52 d ago

    Reddit has announced it will discontinue RSS feed support on November 13, 2026, citing RSS as a common surface for large-scale scraping and automated abuse. Moderation workflows that depend on RSS will be preserved. Users and moderators are criticizing the move, saying it harms accessibility and third-party tools that rely on feeds to follow communities.

  25. 25
    OpenAI accuses Chinese model of stealing its IP, drawing irony chargesβ–ΌIrony alert: OpenAI whines that Chinese model stole its special IP that it stole from everybody elseMmastodon532 d ago

    OpenAI claims a Chinese AI model improperly copied its work, reportedly framing the practice as a potential national security risk. Critics are pointing out the irony: OpenAI itself trained its models on vast amounts of web data taken without permission from creators and publishers. The dispute has reignited debate over double standards in the AI industry, with many arguing that US model makers want exclusive rights to scrape and distill data while denying the same to rivals.

  26. 26
    Reddit curbs 'Old Reddit' access citing AI bots●Reddit says it has to cut back access to 'Old Reddit' because of AI bots https://www.theverge.com/tech/1002788/old-reddiMmastodonTechnology43 d ago

    Reddit says it must restrict access to its legacy 'Old Reddit' interface because of AI scraping bots overwhelming the site. The company frames the move as necessary to protect the platform from automated traffic, though the change is likely to frustrate longtime users who prefer the simpler, faster classic layout over the modern redesign.

  27. 27
    Philadelphia Inquirer launches AI tool Scrape for hyperlocal news●The Philadelphia Inquirer built Scrape, an AI tool to surface hyperlocal news https://www.lenfestinstitute.org/solutionsMmastodonTechnology31 d ago

    The Philadelphia Inquirer has built Scrape, an artificial intelligence tool designed to surface hyperlocal news that might otherwise go unreported. The project is highlighted by the Lenfest Institute, which supports the paper and promotes it as a model for local journalism. Observers in tech circles are discussing whether AI can help struggling local outlets cover neighborhood-level stories at scale.

  28. 28

    Debates over how Multiple Listing Service data is licensed are intensifying as artificial intelligence raises new questions about who can use property listings and on what terms. Industry figures are weighing how to protect MLS data from unrestricted AI training and scraping while preserving its value for agents and brokerages. No specific policy changes have been announced yet.

  29. 29
    Skateboarding injuries send hundreds of thousands to emergency rooms●Skateboarding can be a dangerous activity with a high risk of injury, but the severity of these injuries can vary. β€’ InMmastodonCultureMusic211 h ago

    Skateboarding carries a high risk of injury, with severity varying widely depending on the accident. In 2022, an estimated 230,000 people were treated in emergency rooms for injuries related to skateboarding, scooters and hoverboards in the United States. Commenters are highlighting the dangers of these popular activities and the range of injuries, from minor scrapes to serious fractures and head trauma.

  30. 30
    Developer Launches Google Maps Scraper MCP Tool●Show HN: Google Maps Scraper MCP https://gmapscrawl.com/google-maps-scraper-mcp # HackerNews # Tech # OpenSourceMmastodonTechnologySoftware31 d ago

    A new open-source tool called Google Maps Scraper MCP has been released, allowing developers to extract location and business data from Google Maps through the Model Context Protocol. The launch is being shared and discussed in developer communities, where users are weighing its usefulness for AI-driven workflows against questions around scraping legality and terms of service.

  31. 31
    Google Hides Search Link Destinations to Deter Scrapers●Google's Goto Gambit: Search Links Now Hide Destinations to Thwart Scrapers https:// lemmy.dbzer0.com/post/76357045MmastodonTechnology21 d ago

    Google has changed how search result links work so that the destination web address is hidden until the user clicks, with the redirect routed through a Google 'goto' mechanism. The change is being discussed as a move to stop scrapers and automated tools from extracting search results, though users worry it reduces transparency about where links lead.

  32. 32
    Reddit kills RSS feeds and public API access over AI bots●Reddit is killing RSS feeds and ending public API access because of AI bots https://techcrunch.com/2026/09/30/reddit-is-MmastodonTechnology43 d ago

    Reddit is ending public API access and shutting down RSS feeds, citing abuse by AI bots scraping its content. The move follows the platform's earlier restrictions on data access and reflects a broader industry shift to lock down content that could train AI models. The change affects developers, researchers, and third-party tools that relied on free, open access to Reddit data.

  33. 33
    OpenAI Agents Accused of Scraping 55 Sites Amid Probe●OpenAI AI Agents Secretly Scraped Data from 55 Sites as Regulators Launch Probe𝕏xSE1302 d ago

    Reports claim OpenAI's AI agents collected data from 55 websites without visible disclosure, prompting regulators to open an investigation. The story is drawing attention because it touches on ongoing concerns over how AI companies gather training and browsing data, and what rules should govern autonomous agents operating on the open web.

  34. 34
    BitShot app promo video showcases crypto tracking process●And here is promo video for BitShot app, similar to an article I wrote earlier but showcasing the whole process with scrMmastodonBusinessCrypto03 d ago

    A security researcher has shared a promotional video for BitShot, an app demonstrating a complete screen-recorded process tied to bitcoin, OSINT and data scraping. The demo follows an earlier written article by the same person and highlights the tool's open-source approach to cryptocurrency monitoring and privacy-related security research.

  35. 35
    Google hides destination URLs in search links to deter scrapers●Google's Goto Gambit: Search Links Now Hide Destinations to Thwart Scrapers https:// lemmy.dbzer0.com/post/76357044MmastodonTechnology21 d ago

    Google has changed how search result links work, replacing visible destination URLs with intermediary redirects that conceal the actual target address. The move, dubbed the 'Goto' gambit after the redirect domain in use, is seen as an attempt to stop scrapers and automated tools from harvesting results. Users report it complicates link inspection and breaks some third-party tools, renewing criticism of Google's tightening grip on search data.

  36. 36
    Google hides search link destinations to block scrapers●Google's Goto Gambit: Search Links Now Hide Destinations to Thwart Scrapers https:// lemmy.dbzer0.com/post/76356943MmastodonTechnology21 d ago

    Google has changed how links in search results work, reportedly routing clicks through intermediary URLs that hide the actual destination address. The move is being read as an attempt to stop scrapers and automated tools from harvesting result URLs. Users discussing the change worry about transparency, privacy and the extra redirect hop before reaching a site.

  37. 37
    Vogue and Esquire publishers urge Congress to ban AI content theft●Vogue, Esquire publishers plead with Congress to ban AI theft of their contentβœ‰newsWorldUS Politics3 d ago

    The publishers behind Vogue and Esquire have appealed to Congress to pass legislation preventing AI companies from using their journalism and photography without permission or payment. The plea frames unlicensed scraping of magazine content as theft and adds major media brands to the growing fight over copyright and compensation in the age of generative AI.

  38. 38
    Post-Conference '$10 Gift Card' Emails Allegedly Harvest Professional Data●The "$10 Conference Review" Used to Harvest Your Professional Profile For Profit Post-conference "review for a $10 giftMmastodonTechnologyCybersecurity218 h ago

    Cybersecurity commentators are warning about post-conference emails offering a $10 gift card in exchange for a session review. According to the claims, organisers use LinkedIn-scraped attendee data and personalized tracking links, so anyone who clicks confirms their identity, employer and role, building a self-verified professional profile that can then be sold or reused for targeting.

  39. 39
    Lightpanda 1.0 launches as a browser built for machines●Lightpanda: A browser for machines instead of humans Lightpanda 1.0.0 brings the Classic WebDriver for automation tasks.MmastodonWorld01 d ago

    Lightpanda has released version 1.0.0 of its browser designed for automation rather than human users. The update brings Classic WebDriver support for automation tasks, enforces CORS by default, and adds new Web APIs. The browser is aimed at developers running AI agents, scraping, and testing workloads at scale, and coverage of the release is drawing attention to the growing demand for machine-facing web tooling.

  40. 40
    Scientists urged to sabotage AI training with junk data●If you are a scientist and have been asked to participate in AI crap, please feed the machines crazy ideas that are a) wMmastodonScience23 d ago

    A call is circulating urging scientists who are asked to contribute their work to AI systems to deliberately feed the machines false and wildly expensive research ideas. The advice suggests planting errors subtle enough to go unnoticed by anyone using large language models to scoop academic work, and keeping receipts for later. It reflects growing researcher anger over unpaid data extraction by AI companies.

Repos