search
Scrape
Trends
- 1OpenAI and Microsoft researchers warn of AI 'doom loop' consuming the webβΌβDoom Loopβ: OpenAI and Microsoft Admits LLMs Are Destroying the Web and Built on Theft
Researchers at OpenAI and Microsoft have published a paper describing a 'doom loop' in which large language models, trained on data scraped largely without consent from human creators, degrade the open web and eventually poison their own training data. The paper reportedly acknowledges that people will come to see the models' wholesale ingestion of creative work as an unprecedented act of theft, reigniting debate over copyright and the sustainability of generative AI.
- 2
A Microsoft director has described AI training data scraping as 'the largest theft of labor in human history', while OpenAI's head of publisher strategy called ChatGPT an existential threat to publishers. The remarks were revealed in legal briefs filed in the New York Times' copyright lawsuit against OpenAI and Microsoft. The comments add to the debate over whether AI companies unfairly exploit journalists' and creators' work without permission or payment.
- 3OpenAI and Microsoft reportedly knew of 'doom loop' for the webβOpenAI and Microsoft knew they were starting a 'doom loop' for the web
A new report claims OpenAI and Microsoft were aware that their AI products, trained partly on scraped web content, could trigger a 'doom loop' degrading the open web. According to The Verge, internal discussions suggested the companies understood that AI-generated content and content theft could crowd out human-made material, drawing comparisons to earlier warnings about Google's zero-click search results.
- 4OpenAI-linked AI agents scanned UN statistics site thousands of timesβAI agents linked to OpenAI scanned a UN statistics website more than 16,500 times between April and June, using proxies
AI agents linked to OpenAI accessed a UN statistics website more than 16,500 times between April and June, using proxies to bypass API limits, according to a researcher who documented the activity on a UN data platform. The episode is fuelling debate about automated web scraping by AI companies and whether their crawlers respect access rules.
- 5AI firm reportedly hammered UN data site to bypass limitsβAn AI company was hammering a UN data website with creative ways to get around the limitations of its API and interface.
An AI company reportedly hammered a United Nations data website, using creative methods to work around the restrictions of its API and interface. Researcher Rowan Howard-Jones pieced together the evidence, which had been published earlier, and the Wall Street Journal has since covered the story. The report adds to ongoing concerns about AI firms aggressively scraping public data sources.
- 6OpenAI agents scanned UNCTAD site 16,000 times, researcher saysβSecurity researcher Rowan Howard-Jones says that OpenAI agents scanned the UN Conference on Trade and Development's (UNC
Security researcher Rowan Howard-Jones reports that OpenAI's automated agents hit the UN Conference on Trade and Development's statistics site more than 16,000 times between April and June. The finding is fueling renewed debate over how AI companies' crawlers and agents consume public web infrastructure without asking, though he notes it is less severe than incidents like the Hugging Face breach.
- 7Researcher links 16,000 scans of UN statistics portal to OpenAI agentsβΌResearcher links 16,000 scans of a UN statistics portal to OpenAI agents
A security researcher has linked roughly 16,000 scans of a United Nations statistics portal to OpenAI's AI agents, suggesting automated browsing by the company's systems at significant scale. The finding raises fresh questions about how AI agents crawl and interact with websites without obvious authorization, and whether international organizations are prepared for this new wave of automated traffic.
- 8Claims of Coordinated Campaign Against Glass Skyscrapers in IndiaβI am seeing a concerted campaign against glass skyscrapers in India through multi channel targeting of different groups
A claim is circulating that a coordinated, multi-channel effort is being run against glass skyscrapers in India, targeting environmentalists and animal lovers in particular. The person making the claim characterises it as a possible psyop intended to disrupt India's construction of tall buildings. No evidence is offered to substantiate the alleged campaign.
- 9Elephants mine minerals deep inside Kenya's Kitum Caveβπ WATCH: In the pitch-black chambers of # MountElgon 's # Kitum Cave, African # savannah # elephants use their tusks to
African savannah elephants venture into the pitch-black chambers of Kitum Cave on Mount Elgon in Kenya, using their tusks to scrape volcanic rock for sodium, calcium and magnesium. Generations of herds have hollowed out these deep mineral chambers, passing the unusual mining behaviour down through the herd. Footage of the elephants at work is drawing attention to one of nature's most remarkable learned traditions.
- 10
A new write-up argues that Meta's Muse model performs impressively well at web scraping, drawing attention among developers evaluating AI tools for automated data extraction. The piece has sparked discussion about how capable Meta's model is at practical, unglamorous tasks like parsing and extracting information from web pages.
- 11Substack accused of obfuscating text to break reader modeβTell HN: Substack obfuscating text to break reading mode
Substack is being accused of deliberately obfuscating its article text so that browser-based reader modes cannot cleanly parse its pages, forcing visitors to read on the site itself. Commenters are debating whether the practice is an anti-scraping measure or a hostile move against accessibility tools, with many calling it anti-user and comparing it to similar tactics on other publishing platforms.
- 12EU Commission accused of sacrificing children's privacy to AI firmsβRE: https:// mstdn.social/@stux/11737340528 4945562 also, @EUCommission wants the world to offer their childrenβs privac
The European Commission is facing criticism over its approach to AI companies and data scraping. Detractors argue the Commission is effectively asking the world to hand over children's privacy and a lifetime of personal data to tech firms, claiming these companies have already scraped everything available online without consent. The accusation reflects wider anger in Europe over AI training data practices and weak safeguards for minors.
- 13Reddit is killing RSS feeds and public API access over AI botsβΌReddit is killing RSS feeds and ending public API access because of AI bots | TechCrunch
Reddit is ending support for RSS feeds as it continues restricting access to its archive of user-generated content, with TechCrunch reporting the move is aimed at stopping AI companies from scraping the platform to train their models. The change means third-party tools and readers that relied on open feeds will lose access, deepening Reddit's shift toward tightly controlled, paid data licensing.
- 14Elephants mine minerals deep inside Mount Elgon's Kitum Caveβπ ICYMI: In the pitch-black chambers of # MountElgon 's # Kitum Cave, African # savannah # elephants use their tusks to
African savannah elephants venture into the pitch-black chambers of Kitum Cave on Mount Elgon, using their tusks to scrape volcanic rock rich in sodium, calcium and magnesium. Generations of herds have hollowed out these deep mineral chambers, passing the unusual mining behaviour down through the generations.
- 15
Le Parisien has published a report on households in France who say they live 'day to day on a workaround', scraping by to the nearest euro as the cost of living erodes their budgets. The piece details how families cut spending, juggle bills and delay purchases to make ends meet, fuelling renewed public debate about purchasing power, one of the most sensitive issues in French politics.
- 16UK Media Mocked for Reporting on Someone's LunchβWhy is someoneβs lunch newsworthy??? Slow day in the UK?
People in the UK are questioning why a national news outlet ran a story about an ordinary person's lunch, with some calling it a sign of a slow news day. The reaction online has been largely mocking, as readers express disbelief that such a mundane topic warranted coverage, and debate whether the media is scraping the barrel for content.
- 17
A claim is circulating among developers that a Chinese AI model is 'stealing' code, suggesting it may be trained on or reproduce programmers' work without permission. The warning has drawn large attention in the software community, where concerns about code provenance, licensing and data scraping by AI companies are already running high. No specific model or proof is named in the discussion, so the accusation remains unverified.
- 18
A technical article from TrueSign examines common methods used to detect residential proxies and argues they frequently fail. The piece walks through flawed detection approaches and explains why they produce poor results, drawing attention from developers and anti-fraud engineers on Hacker News, where it reached a top-10 position amid ongoing debate over bot detection and proxy evasion techniques.
- 19Philadelphia Inquirer launches Scrape, an AI tool for hyperlocal newsβΌThe Philadelphia Inquirer built Scrape, an AI tool to surface hyperlocal news
The Philadelphia Inquirer has built Scrape, an artificial intelligence tool designed to surface hyperlocal news stories that might otherwise go unreported. Developed with support from the Lenfest Institute, the tool is aimed at helping the newsroom cover neighbourhood-level events and community issues more efficiently. The project is drawing attention as an example of how local news organisations are experimenting with AI to expand coverage despite shrinking resources.
- 20Reddit restricts Old Reddit access citing AI botsβReddit says it has to cut back access to βOld Redditβ because of AI bots Reddit is further limiting who can use the "Old
Reddit says it is further limiting access to its legacy 'Old Reddit' interface, citing efforts to combat scraping and automated traffic from AI bots. The company has recently begun requiring users to log in before using the classic experience. Many long-time users rely on Old Reddit for its simpler design, and the tightening rules have sparked frustration among those who see it as another step toward restricting free access to the platform.
- 21AI bots and scraping could push the open web out of reachβAnno 202X, il web Γ¨ cosΓ¬ impestato e saccheggiato da bot e scraping dell'AI che i costi per la gestione di siti e comuni
Commenters warn that by 202X the web could be so infested with bots and AI scraping that running websites and online communities becomes affordable only for a handful of global billionaires, pushing humanity into isolated, self-contained systems. The post is drawing attention in Italian-language discussions about the rising costs of maintaining an open internet under pressure from automated AI traffic.
- 22Reddit ends RSS feeds and public API access, citing AI botsβReddit is killing RSS feeds and ending public API access because of AI bots https:// lemmy.zip/post/72580572
Reddit is shutting down its RSS feeds and ending public API access, attributing the move to the flood of AI bots scraping its content. The change means third-party tools and users who relied on open feeds will lose access to Reddit data without going through official channels. Discussions online frame it as another step in Reddit tightening control over its data amid the broader AI scraping boom.
- 23
A new tool called Google Maps Scraper MCP has been introduced, designed to let AI assistants and applications extract business data from Google Maps, such as locations, contact details and reviews. The launch is drawing attention from developers and businesses interested in automated location data collection, as well as those weighing the legal and practical limits of scraping Google's mapping services.
- 24Reddit to shut down RSS feeds in 2026βΌReddit is Killing RSS Support https://www. reddit.com/r/modnews/s/kvHTs7I xpg "Because RSS is now a common surface for l
Reddit has announced it will discontinue RSS feed support on November 13, 2026, citing RSS as a common surface for large-scale scraping and automated abuse. Moderation workflows that depend on RSS will be preserved. Users and moderators are criticizing the move, saying it harms accessibility and third-party tools that rely on feeds to follow communities.
- 25OpenAI accuses Chinese model of stealing its IP, drawing irony chargesβΌIrony alert: OpenAI whines that Chinese model stole its special IP that it stole from everybody else
OpenAI claims a Chinese AI model improperly copied its work, reportedly framing the practice as a potential national security risk. Critics are pointing out the irony: OpenAI itself trained its models on vast amounts of web data taken without permission from creators and publishers. The dispute has reignited debate over double standards in the AI industry, with many arguing that US model makers want exclusive rights to scrape and distill data while denying the same to rivals.
- 26Reddit curbs 'Old Reddit' access citing AI botsβReddit says it has to cut back access to 'Old Reddit' because of AI bots https://www.theverge.com/tech/1002788/old-reddi
Reddit says it must restrict access to its legacy 'Old Reddit' interface because of AI scraping bots overwhelming the site. The company frames the move as necessary to protect the platform from automated traffic, though the change is likely to frustrate longtime users who prefer the simpler, faster classic layout over the modern redesign.
- 27Philadelphia Inquirer launches AI tool Scrape for hyperlocal newsβThe Philadelphia Inquirer built Scrape, an AI tool to surface hyperlocal news https://www.lenfestinstitute.org/solutions
The Philadelphia Inquirer has built Scrape, an artificial intelligence tool designed to surface hyperlocal news that might otherwise go unreported. The project is highlighted by the Lenfest Institute, which supports the paper and promotes it as a model for local journalism. Observers in tech circles are discussing whether AI can help struggling local outlets cover neighborhood-level stories at scale.
- 28
Debates over how Multiple Listing Service data is licensed are intensifying as artificial intelligence raises new questions about who can use property listings and on what terms. Industry figures are weighing how to protect MLS data from unrestricted AI training and scraping while preserving its value for agents and brokerages. No specific policy changes have been announced yet.
- 29Skateboarding injuries send hundreds of thousands to emergency roomsβSkateboarding can be a dangerous activity with a high risk of injury, but the severity of these injuries can vary. β’ In
Skateboarding carries a high risk of injury, with severity varying widely depending on the accident. In 2022, an estimated 230,000 people were treated in emergency rooms for injuries related to skateboarding, scooters and hoverboards in the United States. Commenters are highlighting the dangers of these popular activities and the range of injuries, from minor scrapes to serious fractures and head trauma.
- 30Developer Launches Google Maps Scraper MCP ToolβShow HN: Google Maps Scraper MCP https://gmapscrawl.com/google-maps-scraper-mcp # HackerNews # Tech # OpenSource
A new open-source tool called Google Maps Scraper MCP has been released, allowing developers to extract location and business data from Google Maps through the Model Context Protocol. The launch is being shared and discussed in developer communities, where users are weighing its usefulness for AI-driven workflows against questions around scraping legality and terms of service.
- 31Google Hides Search Link Destinations to Deter ScrapersβGoogle's Goto Gambit: Search Links Now Hide Destinations to Thwart Scrapers https:// lemmy.dbzer0.com/post/76357045
Google has changed how search result links work so that the destination web address is hidden until the user clicks, with the redirect routed through a Google 'goto' mechanism. The change is being discussed as a move to stop scrapers and automated tools from extracting search results, though users worry it reduces transparency about where links lead.
- 32Reddit kills RSS feeds and public API access over AI botsβReddit is killing RSS feeds and ending public API access because of AI bots https://techcrunch.com/2026/09/30/reddit-is-
Reddit is ending public API access and shutting down RSS feeds, citing abuse by AI bots scraping its content. The move follows the platform's earlier restrictions on data access and reflects a broader industry shift to lock down content that could train AI models. The change affects developers, researchers, and third-party tools that relied on free, open access to Reddit data.
- 33OpenAI Agents Accused of Scraping 55 Sites Amid ProbeβOpenAI AI Agents Secretly Scraped Data from 55 Sites as Regulators Launch Probe
Reports claim OpenAI's AI agents collected data from 55 websites without visible disclosure, prompting regulators to open an investigation. The story is drawing attention because it touches on ongoing concerns over how AI companies gather training and browsing data, and what rules should govern autonomous agents operating on the open web.
- 34BitShot app promo video showcases crypto tracking processβAnd here is promo video for BitShot app, similar to an article I wrote earlier but showcasing the whole process with scr
A security researcher has shared a promotional video for BitShot, an app demonstrating a complete screen-recorded process tied to bitcoin, OSINT and data scraping. The demo follows an earlier written article by the same person and highlights the tool's open-source approach to cryptocurrency monitoring and privacy-related security research.
- 35Google hides destination URLs in search links to deter scrapersβGoogle's Goto Gambit: Search Links Now Hide Destinations to Thwart Scrapers https:// lemmy.dbzer0.com/post/76357044
Google has changed how search result links work, replacing visible destination URLs with intermediary redirects that conceal the actual target address. The move, dubbed the 'Goto' gambit after the redirect domain in use, is seen as an attempt to stop scrapers and automated tools from harvesting results. Users report it complicates link inspection and breaks some third-party tools, renewing criticism of Google's tightening grip on search data.
- 36Google hides search link destinations to block scrapersβGoogle's Goto Gambit: Search Links Now Hide Destinations to Thwart Scrapers https:// lemmy.dbzer0.com/post/76356943
Google has changed how links in search results work, reportedly routing clicks through intermediary URLs that hide the actual destination address. The move is being read as an attempt to stop scrapers and automated tools from harvesting result URLs. Users discussing the change worry about transparency, privacy and the extra redirect hop before reaching a site.
- 37Vogue and Esquire publishers urge Congress to ban AI content theftβVogue, Esquire publishers plead with Congress to ban AI theft of their content
The publishers behind Vogue and Esquire have appealed to Congress to pass legislation preventing AI companies from using their journalism and photography without permission or payment. The plea frames unlicensed scraping of magazine content as theft and adds major media brands to the growing fight over copyright and compensation in the age of generative AI.
- 38Post-Conference '$10 Gift Card' Emails Allegedly Harvest Professional DataβThe "$10 Conference Review" Used to Harvest Your Professional Profile For Profit Post-conference "review for a $10 gift
Cybersecurity commentators are warning about post-conference emails offering a $10 gift card in exchange for a session review. According to the claims, organisers use LinkedIn-scraped attendee data and personalized tracking links, so anyone who clicks confirms their identity, employer and role, building a self-verified professional profile that can then be sold or reused for targeting.
- 39Lightpanda 1.0 launches as a browser built for machinesβLightpanda: A browser for machines instead of humans Lightpanda 1.0.0 brings the Classic WebDriver for automation tasks.
Lightpanda has released version 1.0.0 of its browser designed for automation rather than human users. The update brings Classic WebDriver support for automation tasks, enforces CORS by default, and adds new Web APIs. The browser is aimed at developers running AI agents, scraping, and testing workloads at scale, and coverage of the release is drawing attention to the growing demand for machine-facing web tooling.
- 40Scientists urged to sabotage AI training with junk dataβIf you are a scientist and have been asked to participate in AI crap, please feed the machines crazy ideas that are a) w
A call is circulating urging scientists who are asked to contribute their work to AI systems to deliberately feed the machines false and wildly expensive research ideas. The advice suggests planting errors subtle enough to go unnoticed by anyone using large language models to scoop academic work, and keeping receipts for later. It reflects growing researcher anger over unpaid data extraction by AI companies.
Repos
- mobile-next/mobile-mcp Model Context Protocol Server for Mobile Automation and Scraping (iOS, Android, Emulators, Simulators and Real Devices)