MikeTrendsTrends right now

search

Scrape

Trends

  1. 1
    OpenAI and Microsoft researchers warn of AI 'doom loop' consuming the web▼‘Doom Loop’: OpenAI and Microsoft Admits LLMs Are Destroying the Web and Built on TheftMmastodonTechnologyAI855 d ago

    Researchers at OpenAI and Microsoft have published a paper describing a 'doom loop' in which large language models, trained on data scraped largely without consent from human creators, degrade the open web and eventually poison their own training data. The paper reportedly acknowledges that people will come to see the models' wholesale ingestion of creative work as an unprecedented act of theft, reigniting debate over copyright and the sustainability of generative AI.

  2. 2

    A Microsoft director has described AI training data scraping as 'the largest theft of labor in human history', while OpenAI's head of publisher strategy called ChatGPT an existential threat to publishers. The remarks were revealed in legal briefs filed in the New York Times' copyright lawsuit against OpenAI and Microsoft. The comments add to the debate over whether AI companies unfairly exploit journalists' and creators' work without permission or payment.

  3. 3
    OpenAI and Microsoft reportedly knew of 'doom loop' for the web●OpenAI and Microsoft knew they were starting a 'doom loop' for the webYhnTechnologyInternet356 d ago

    A new report claims OpenAI and Microsoft were aware that their AI products, trained partly on scraped web content, could trigger a 'doom loop' degrading the open web. According to The Verge, internal discussions suggested the companies understood that AI-generated content and content theft could crowd out human-made material, drawing comparisons to earlier warnings about Google's zero-click search results.

  4. 4
    AI firm reportedly hammered UN data site to bypass limits●An AI company was hammering a UN data website with creative ways to get around the limitations of its API and interface.MmastodonWorldUnited Nations44 d ago

    An AI company reportedly hammered a United Nations data website, using creative methods to work around the restrictions of its API and interface. Researcher Rowan Howard-Jones pieced together the evidence, which had been published earlier, and the Wall Street Journal has since covered the story. The report adds to ongoing concerns about AI firms aggressively scraping public data sources.

  5. 5
    OpenAI-linked AI agents scanned UN statistics site thousands of times●AI agents linked to OpenAI scanned a UN statistics website more than 16,500 times between April and June, using proxiesMmastodonWorld25 d ago

    AI agents linked to OpenAI accessed a UN statistics website more than 16,500 times between April and June, using proxies to bypass API limits, according to a researcher who documented the activity on a UN data platform. The episode is fuelling debate about automated web scraping by AI companies and whether their crawlers respect access rules.

  6. 6
    Claims of Coordinated Campaign Against Glass Skyscrapers in India●I am seeing a concerted campaign against glass skyscrapers in India through multi channel targeting of different groups𝕏xIN4.3K2 d ago

    A claim is circulating that a coordinated, multi-channel effort is being run against glass skyscrapers in India, targeting environmentalists and animal lovers in particular. The person making the claim characterises it as a possible psyop intended to disrupt India's construction of tall buildings. No evidence is offered to substantiate the alleged campaign.

  7. 7
    Elephants mine minerals deep inside Kenya's Kitum Cave●🐘 WATCH: In the pitch-black chambers of # MountElgon 's # Kitum Cave, African # savannah # elephants use their tusks toMmastodonScience181 d ago

    African savannah elephants venture into the pitch-black chambers of Kitum Cave on Mount Elgon in Kenya, using their tusks to scrape volcanic rock for sodium, calcium and magnesium. Generations of herds have hollowed out these deep mineral chambers, passing the unusual mining behaviour down through the herd. Footage of the elephants at work is drawing attention to one of nature's most remarkable learned traditions.

  8. 8
    OpenAI agents scanned UNCTAD site 16,000 times, researcher says●Security researcher Rowan Howard-Jones says that OpenAI agents scanned the UN Conference on Trade and Development's (UNCMmastodonTechnologyAI35 d ago

    Security researcher Rowan Howard-Jones reports that OpenAI's automated agents hit the UN Conference on Trade and Development's statistics site more than 16,000 times between April and June. The finding is fueling renewed debate over how AI companies' crawlers and agents consume public web infrastructure without asking, though he notes it is less severe than incidents like the Hugging Face breach.

  9. 9
    EU Commission accused of sacrificing children's privacy to AI firms●RE: https:// mstdn.social/@stux/11737340528 4945562 also, @EUCommission wants the world to offer their children’s privacMmastodonTechnologyAI817 h ago

    The European Commission is facing criticism over its approach to AI companies and data scraping. Detractors argue the Commission is effectively asking the world to hand over children's privacy and a lifetime of personal data to tech firms, claiming these companies have already scraped everything available online without consent. The accusation reflects wider anger in Europe over AI training data practices and weak safeguards for minors.

  10. 10
    Researcher links 16,000 scans of UN statistics portal to OpenAI agents▼Researcher links 16,000 scans of a UN statistics portal to OpenAI agents✉newsWorldUnited Nations5 d ago

    A security researcher has linked roughly 16,000 scans of a United Nations statistics portal to OpenAI's AI agents, suggesting automated browsing by the company's systems at significant scale. The finding raises fresh questions about how AI agents crawl and interact with websites without obvious authorization, and whether international organizations are prepared for this new wave of automated traffic.

  11. 11

    A new write-up argues that Meta's Muse model performs exceptionally well at web scraping tasks. The author walks through using the model to extract structured data from web pages, praising its accuracy and ease of use compared to other approaches. Reaction so far has centered on whether Muse could become a go-to tool for scraping and data extraction work.

  12. 12
    Reddit is killing RSS feeds and public API access over AI bots▼Reddit is killing RSS feeds and ending public API access because of AI bots | TechCrunchMmastodon14821 h ago

    Reddit is ending support for RSS feeds as it continues restricting access to its archive of user-generated content, with TechCrunch reporting the move is aimed at stopping AI companies from scraping the platform to train their models. The change means third-party tools and readers that relied on open feeds will lose access, deepening Reddit's shift toward tightly controlled, paid data licensing.

  13. 13
    UK Media Mocked for Reporting on Someone's Lunch●Why is someone’s lunch newsworthy??? Slow day in the UK?𝕏xGB216 h ago

    People in the UK are questioning why a national news outlet ran a story about an ordinary person's lunch, with some calling it a sign of a slow news day. The reaction online has been largely mocking, as readers express disbelief that such a mundane topic warranted coverage, and debate whether the media is scraping the barrel for content.

  14. 14
    Liberal Education as the Art of Coming Unstuck▼Barnacles: Liberal Education & the Art of Coming Unstuck✉newsLifeEducation19 min ago

    An essay titled 'Barnacles: Liberal Education & the Art of Coming Unstuck' argues that a liberal education functions like scraping barnacles from a hull, freeing minds encrusted by habit, ideology and narrow specialisation. Written from a conservative humanist perspective, it presents the classical curriculum as a way to reattach readers to enduring questions and restart intellectual momentum.

  15. 15
    Substack accused of obfuscating text to break reader mode●Tell HN: Substack obfuscating text to break reading modeYhnBusinessCrypto203 d ago

    Substack is being accused of deliberately obfuscating its article text so that browser-based reader modes cannot cleanly parse its pages, forcing visitors to read on the site itself. Commenters are debating whether the practice is an anti-scraping measure or a hostile move against accessibility tools, with many calling it anti-user and comparing it to similar tactics on other publishing platforms.

  16. 16
    AI bots and scraping could push the open web out of reach●Anno 202X, il web è così impestato e saccheggiato da bot e scraping dell'AI che i costi per la gestione di siti e comuniMmastodonTechnologyInternet21 h ago

    Commenters warn that by 202X the web could be so infested with bots and AI scraping that running websites and online communities becomes affordable only for a handful of global billionaires, pushing humanity into isolated, self-contained systems. The post is drawing attention in Italian-language discussions about the rising costs of maintaining an open internet under pressure from automated AI traffic.

  17. 17
    Philadelphia Inquirer launches Scrape AI tool for hyperlocal news●The Philadelphia Inquirer built Scrape, an AI tool to surface hyperlocal newsYhnEnvironment1456 min ago

    The Philadelphia Inquirer, working with the Lenfest Institute, has built Scrape, an AI tool designed to surface hyperlocal news stories that might otherwise go unreported. The tool is drawing attention as an example of how local newsrooms are using artificial intelligence to expand coverage of neighbourhood-level events and hold down reporting costs.

  18. 18

    Le Parisien has published a report on households in France who say they live 'day to day on a workaround', scraping by to the nearest euro as the cost of living erodes their budgets. The piece details how families cut spending, juggle bills and delay purchases to make ends meet, fuelling renewed public debate about purchasing power, one of the most sensitive issues in French politics.

  19. 19
    Elephants Mine Kitum Cave Walls for Minerals●🐘 ICYMI: In the pitch-black chambers of # MountElgon 's # Kitum Cave, African # savannah # elephants use their tusks toMmastodonHealth61 d ago

    In Mount Elgon's Kitum Cave in Kenya, African savannah elephants use their tusks to scrape volcanic rock for sodium, calcium and magnesium. Generations of herds have hollowed out these deep mineral chambers, passing the unusual mining behaviour down through the herd. The story is being shared as a striking example of animal culture and adaptation.

  20. 20
    Reddit ends RSS feeds and public API access, citing AI bots●Reddit is killing RSS feeds and ending public API access because of AI bots https:// lemmy.zip/post/72580572MmastodonTechnology43 h ago

    Reddit is shutting down its RSS feeds and ending public API access, attributing the move to the flood of AI bots scraping its content. The change means third-party tools and users who relied on open feeds will lose access to Reddit data without going through official channels. Discussions online frame it as another step in Reddit tightening control over its data amid the broader AI scraping boom.

  21. 21
    Chinese AI model accused of taking developers' code●A Chinese AI model is stealing your code▶youtubeTechnologySoftware94.6K3 d ago

    A claim is circulating among developers that a Chinese AI model is 'stealing' code, suggesting it may be trained on or reproduce programmers' work without permission. The warning has drawn large attention in the software community, where concerns about code provenance, licensing and data scraping by AI companies are already running high. No specific model or proof is named in the discussion, so the accusation remains unverified.

  22. 22
    Reddit restricts Old Reddit access citing AI bots●Reddit says it has to cut back access to ‘Old Reddit’ because of AI bots Reddit is further limiting who can use the "OldMmastodonTechnology32 d ago

    Reddit says it is further limiting access to its legacy 'Old Reddit' interface, citing efforts to combat scraping and automated traffic from AI bots. The company has recently begun requiring users to log in before using the classic experience. Many long-time users rely on Old Reddit for its simpler design, and the tightening rules have sparked frustration among those who see it as another step toward restricting free access to the platform.

  23. 23
    Reddit to shut down RSS feeds in 2026▼Reddit is Killing RSS Support https://www. reddit.com/r/modnews/s/kvHTs7I xpg "Because RSS is now a common surface for lMmastodonTechnology51 d ago

    Reddit has announced it will discontinue RSS feed support on November 13, 2026, citing RSS as a common surface for large-scale scraping and automated abuse. Moderation workflows that depend on RSS will be preserved. Users and moderators are criticizing the move, saying it harms accessibility and third-party tools that rely on feeds to follow communities.

  24. 24

    A technical blog post from TrueSign argues that many common methods for detecting residential proxies are flawed, walking through approaches that fail and why. The piece is drawing attention among developers and anti-fraud engineers who deal with bot detection, proxy filtering and traffic verification, sparking discussion about which detection techniques actually hold up in practice.

  25. 25
    Philadelphia Inquirer launches AI tool Scrape for hyperlocal news●The Philadelphia Inquirer built Scrape, an AI tool to surface hyperlocal news https://www.lenfestinstitute.org/solutionsMmastodonTechnology316 h ago

    The Philadelphia Inquirer has built Scrape, an artificial intelligence tool designed to surface hyperlocal news that might otherwise go unreported. The project is highlighted by the Lenfest Institute, which supports the paper and promotes it as a model for local journalism. Observers in tech circles are discussing whether AI can help struggling local outlets cover neighborhood-level stories at scale.

  26. 26
    OpenAI accuses Chinese model of stealing its IP, drawing irony charges▼Irony alert: OpenAI whines that Chinese model stole its special IP that it stole from everybody elseMmastodon531 d ago

    OpenAI claims a Chinese AI model improperly copied its work, reportedly framing the practice as a potential national security risk. Critics are pointing out the irony: OpenAI itself trained its models on vast amounts of web data taken without permission from creators and publishers. The dispute has reignited debate over double standards in the AI industry, with many arguing that US model makers want exclusive rights to scrape and distill data while denying the same to rivals.

  27. 27

    Debates over how Multiple Listing Service data is licensed are intensifying as artificial intelligence raises new questions about who can use property listings and on what terms. Industry figures are weighing how to protect MLS data from unrestricted AI training and scraping while preserving its value for agents and brokerages. No specific policy changes have been announced yet.

  28. 28
    Reddit curbs 'Old Reddit' access citing AI bots●Reddit says it has to cut back access to 'Old Reddit' because of AI bots https://www.theverge.com/tech/1002788/old-reddiMmastodonTechnology42 d ago

    Reddit says it must restrict access to its legacy 'Old Reddit' interface because of AI scraping bots overwhelming the site. The company frames the move as necessary to protect the platform from automated traffic, though the change is likely to frustrate longtime users who prefer the simpler, faster classic layout over the modern redesign.

  29. 29
    Developer Launches Google Maps Scraper MCP Tool●Show HN: Google Maps Scraper MCP https://gmapscrawl.com/google-maps-scraper-mcp # HackerNews # Tech # OpenSourceMmastodonTechnologySoftware310 h ago

    A new open-source tool called Google Maps Scraper MCP has been released, allowing developers to extract location and business data from Google Maps through the Model Context Protocol. The launch is being shared and discussed in developer communities, where users are weighing its usefulness for AI-driven workflows against questions around scraping legality and terms of service.

  30. 30
    Google Hides Search Link Destinations to Deter Scrapers●Google's Goto Gambit: Search Links Now Hide Destinations to Thwart Scrapers https:// lemmy.dbzer0.com/post/76357045MmastodonTechnology218 h ago

    Google has changed how search result links work so that the destination web address is hidden until the user clicks, with the redirect routed through a Google 'goto' mechanism. The change is being discussed as a move to stop scrapers and automated tools from extracting search results, though users worry it reduces transparency about where links lead.

  31. 31
    Post-Conference '$10 Gift Card' Emails Allegedly Harvest Professional Data●The "$10 Conference Review" Used to Harvest Your Professional Profile For Profit Post-conference "review for a $10 giftMmastodonTechnologyCybersecurity21 h ago

    Cybersecurity commentators are warning about post-conference emails offering a $10 gift card in exchange for a session review. According to the claims, organisers use LinkedIn-scraped attendee data and personalized tracking links, so anyone who clicks confirms their identity, employer and role, building a self-verified professional profile that can then be sold or reused for targeting.

  32. 32
    Google hides destination URLs in search links to deter scrapers●Google's Goto Gambit: Search Links Now Hide Destinations to Thwart Scrapers https:// lemmy.dbzer0.com/post/76357044MmastodonTechnology218 h ago

    Google has changed how search result links work, replacing visible destination URLs with intermediary redirects that conceal the actual target address. The move, dubbed the 'Goto' gambit after the redirect domain in use, is seen as an attempt to stop scrapers and automated tools from harvesting results. Users report it complicates link inspection and breaks some third-party tools, renewing criticism of Google's tightening grip on search data.

  33. 33
    Developer launches Google Maps scraper MCP server●Show HN: Google Maps Scraper MCPYhn613 h ago

    A developer has released a Google Maps Scraper MCP server, a tool that lets AI assistants pull business data such as names, addresses, reviews and contact details directly from Google Maps. The launch drew modest attention on Hacker News, where commenters are weighing its usefulness for lead generation and local search analysis against questions about Google's terms of service and scraping restrictions.

  34. 34
    Google hides search link destinations to block scrapers●Google's Goto Gambit: Search Links Now Hide Destinations to Thwart Scrapers https:// lemmy.dbzer0.com/post/76356943MmastodonTechnology218 h ago

    Google has changed how links in search results work, reportedly routing clicks through intermediary URLs that hide the actual destination address. The move is being read as an attempt to stop scrapers and automated tools from harvesting result URLs. Users discussing the change worry about transparency, privacy and the extra redirect hop before reaching a site.

  35. 35
    OpenAI Agents Accused of Scraping 55 Sites Amid Probe●OpenAI AI Agents Secretly Scraped Data from 55 Sites as Regulators Launch Probe𝕏xSE1302 d ago

    Reports claim OpenAI's AI agents collected data from 55 websites without visible disclosure, prompting regulators to open an investigation. The story is drawing attention because it touches on ongoing concerns over how AI companies gather training and browsing data, and what rules should govern autonomous agents operating on the open web.

  36. 36
    Lightpanda 1.0 launches as a browser built for machines●Lightpanda: A browser for machines instead of humans Lightpanda 1.0.0 brings the Classic WebDriver for automation tasks.MmastodonWorld019 h ago

    Lightpanda has released version 1.0.0 of its browser designed for automation rather than human users. The update brings Classic WebDriver support for automation tasks, enforces CORS by default, and adds new Web APIs. The browser is aimed at developers running AI agents, scraping, and testing workloads at scale, and coverage of the release is drawing attention to the growing demand for machine-facing web tooling.

  37. 37
    Reddit kills RSS feeds and public API access over AI bots●Reddit is killing RSS feeds and ending public API access because of AI bots https://techcrunch.com/2026/09/30/reddit-is-MmastodonTechnology42 d ago

    Reddit is ending public API access and shutting down RSS feeds, citing abuse by AI bots scraping its content. The move follows the platform's earlier restrictions on data access and reflects a broader industry shift to lock down content that could train AI models. The change affects developers, researchers, and third-party tools that relied on free, open access to Reddit data.

  38. 38
    BitShot app promo video showcases crypto tracking process●And here is promo video for BitShot app, similar to an article I wrote earlier but showcasing the whole process with scrMmastodonBusinessCrypto03 d ago

    A security researcher has shared a promotional video for BitShot, an app demonstrating a complete screen-recorded process tied to bitcoin, OSINT and data scraping. The demo follows an earlier written article by the same person and highlights the tool's open-source approach to cryptocurrency monitoring and privacy-related security research.

  39. 39
    Vogue and Esquire publishers urge Congress to ban AI content theft●Vogue, Esquire publishers plead with Congress to ban AI theft of their content✉newsWorldUS Politics2 d ago

    The publishers behind Vogue and Esquire have appealed to Congress to pass legislation preventing AI companies from using their journalism and photography without permission or payment. The plea frames unlicensed scraping of magazine content as theft and adds major media brands to the growing fight over copyright and compensation in the age of generative AI.

  40. 40
    Phillies eliminated by rivals in Wild Card round▼After squeaking into playoffs, Phils bounced by rivals in Wild Card round✉newsSportBaseball1 d ago

    The Philadelphia Phillies, who barely secured a playoff berth, were knocked out by a division rival in the National League Wild Card round. The early exit ends a season in which the club scraped into October, and it is drawing criticism from fans frustrated by another abrupt postseason departure at the hands of a familiar opponent.

Repos