search
Scrape
Trends
- 1From Ma Rainey to AI: Artists Renew a Fight Over Control▼From Ma Rainey to AI: New Technology Amplifies an Old Fight over Artist Control
A new report traces how artificial intelligence has reignited a century-old dispute over who controls artists' work, drawing a line from early 20th-century blues singers like Ma Rainey, whose recordings were exploited by record companies, to today's musicians and writers whose work is scraped to train AI models without consent or payment. The piece argues the technology has raised the stakes in a long-running struggle over ownership, credit and compensation for creative labor.
- 2
A new write-up argues that Meta's Muse model performs exceptionally well at web scraping, drawing attention on developer forums where users are weighing its capabilities against existing tools. The discussion touches on what strong automated extraction could mean for the people whose work is scraped, with some debate over the labor implications of cheap large-scale data collection.
- 3EU Commission accused of sacrificing children's privacy to AI firms●RE: https:// mstdn.social/@stux/11737340528 4945562 also, @EUCommission wants the world to offer their children’s privac
The European Commission is facing criticism over its approach to AI companies and data scraping. Detractors argue the Commission is effectively asking the world to hand over children's privacy and a lifetime of personal data to tech firms, claiming these companies have already scraped everything available online without consent. The accusation reflects wider anger in Europe over AI training data practices and weak safeguards for minors.
- 4Reddit is killing RSS feeds and public API access over AI bots▼Reddit is killing RSS feeds and ending public API access because of AI bots | TechCrunch
Reddit is ending support for RSS feeds as it continues restricting access to its archive of user-generated content, with TechCrunch reporting the move is aimed at stopping AI companies from scraping the platform to train their models. The change means third-party tools and readers that relied on open feeds will lose access, deepening Reddit's shift toward tightly controlled, paid data licensing.
- 5Elephants mine minerals deep inside Kenya's Kitum Cave●🐘 WATCH: In the pitch-black chambers of # MountElgon 's # Kitum Cave, African # savannah # elephants use their tusks to
African savannah elephants venture into the pitch-black chambers of Kitum Cave on Mount Elgon in Kenya, using their tusks to scrape volcanic rock for sodium, calcium and magnesium. Generations of herds have hollowed out these deep mineral chambers, passing the unusual mining behaviour down through the herd. Footage of the elephants at work is drawing attention to one of nature's most remarkable learned traditions.
- 6AI bots and scraping could push the open web out of reach●Anno 202X, il web è così impestato e saccheggiato da bot e scraping dell'AI che i costi per la gestione di siti e comuni
Commenters warn that by 202X the web could be so infested with bots and AI scraping that running websites and online communities becomes affordable only for a handful of global billionaires, pushing humanity into isolated, self-contained systems. The post is drawing attention in Italian-language discussions about the rising costs of maintaining an open internet under pressure from automated AI traffic.
- 7Philadelphia Inquirer launches Scrape AI tool for hyperlocal news●The Philadelphia Inquirer built Scrape, an AI tool to surface hyperlocal news
The Philadelphia Inquirer, working with the Lenfest Institute, has built Scrape, an AI tool designed to surface hyperlocal news stories that might otherwise go unreported. The tool is drawing attention as an example of how local newsrooms are using artificial intelligence to expand coverage of neighbourhood-level events and hold down reporting costs.
- 8UK Media Mocked for Reporting on Someone's Lunch●Why is someone’s lunch newsworthy??? Slow day in the UK?
People in the UK are questioning why a national news outlet ran a story about an ordinary person's lunch, with some calling it a sign of a slow news day. The reaction online has been largely mocking, as readers express disbelief that such a mundane topic warranted coverage, and debate whether the media is scraping the barrel for content.
- 9Reddit ends RSS feeds and public API access, citing AI bots●Reddit is killing RSS feeds and ending public API access because of AI bots https:// lemmy.zip/post/72580572
Reddit is shutting down its RSS feeds and ending public API access, attributing the move to the flood of AI bots scraping its content. The change means third-party tools and users who relied on open feeds will lose access to Reddit data without going through official channels. Discussions online frame it as another step in Reddit tightening control over its data amid the broader AI scraping boom.
- 10Claims of Coordinated Campaign Against Glass Skyscrapers in India●I am seeing a concerted campaign against glass skyscrapers in India through multi channel targeting of different groups
A claim is circulating that a coordinated, multi-channel effort is being run against glass skyscrapers in India, targeting environmentalists and animal lovers in particular. The person making the claim characterises it as a possible psyop intended to disrupt India's construction of tall buildings. No evidence is offered to substantiate the alleged campaign.
- 11Elephants Mine Kitum Cave Walls for Minerals●🐘 ICYMI: In the pitch-black chambers of # MountElgon 's # Kitum Cave, African # savannah # elephants use their tusks to
In Mount Elgon's Kitum Cave in Kenya, African savannah elephants use their tusks to scrape volcanic rock for sodium, calcium and magnesium. Generations of herds have hollowed out these deep mineral chambers, passing the unusual mining behaviour down through the herd. The story is being shared as a striking example of animal culture and adaptation.
- 12Philadelphia Inquirer launches AI tool Scrape for hyperlocal news●The Philadelphia Inquirer built Scrape, an AI tool to surface hyperlocal news https://www.lenfestinstitute.org/solutions
The Philadelphia Inquirer has built Scrape, an artificial intelligence tool designed to surface hyperlocal news that might otherwise go unreported. The project is highlighted by the Lenfest Institute, which supports the paper and promotes it as a model for local journalism. Observers in tech circles are discussing whether AI can help struggling local outlets cover neighborhood-level stories at scale.
- 13Developer Launches Google Maps Scraper MCP Tool●Show HN: Google Maps Scraper MCP https://gmapscrawl.com/google-maps-scraper-mcp # HackerNews # Tech # OpenSource
A new open-source tool called Google Maps Scraper MCP has been released, allowing developers to extract location and business data from Google Maps through the Model Context Protocol. The launch is being shared and discussed in developer communities, where users are weighing its usefulness for AI-driven workflows against questions around scraping legality and terms of service.
- 14Post-Conference '$10 Gift Card' Emails Allegedly Harvest Professional Data●The "$10 Conference Review" Used to Harvest Your Professional Profile For Profit Post-conference "review for a $10 gift
Cybersecurity commentators are warning about post-conference emails offering a $10 gift card in exchange for a session review. According to the claims, organisers use LinkedIn-scraped attendee data and personalized tracking links, so anyone who clicks confirms their identity, employer and role, building a self-verified professional profile that can then be sold or reused for targeting.
- 15Reddit to shut down RSS feeds in 2026▼Reddit is Killing RSS Support https://www. reddit.com/r/modnews/s/kvHTs7I xpg "Because RSS is now a common surface for l
Reddit has announced it will discontinue RSS feed support on November 13, 2026, citing RSS as a common surface for large-scale scraping and automated abuse. Moderation workflows that depend on RSS will be preserved. Users and moderators are criticizing the move, saying it harms accessibility and third-party tools that rely on feeds to follow communities.
- 16
Debates over how Multiple Listing Service data is licensed are intensifying as artificial intelligence raises new questions about who can use property listings and on what terms. Industry figures are weighing how to protect MLS data from unrestricted AI training and scraping while preserving its value for agents and brokerages. No specific policy changes have been announced yet.
- 17OpenAI accuses Chinese model of stealing its IP, drawing irony charges▼Irony alert: OpenAI whines that Chinese model stole its special IP that it stole from everybody else
OpenAI claims a Chinese AI model improperly copied its work, reportedly framing the practice as a potential national security risk. Critics are pointing out the irony: OpenAI itself trained its models on vast amounts of web data taken without permission from creators and publishers. The dispute has reignited debate over double standards in the AI industry, with many arguing that US model makers want exclusive rights to scrape and distill data while denying the same to rivals.
- 18Google Hides Search Link Destinations to Deter Scrapers●Google's Goto Gambit: Search Links Now Hide Destinations to Thwart Scrapers https:// lemmy.dbzer0.com/post/76357045
Google has changed how search result links work so that the destination web address is hidden until the user clicks, with the redirect routed through a Google 'goto' mechanism. The change is being discussed as a move to stop scrapers and automated tools from extracting search results, though users worry it reduces transparency about where links lead.
- 19
A developer has released a Google Maps Scraper MCP server, a tool that lets AI assistants pull business data such as names, addresses, reviews and contact details directly from Google Maps. The launch drew modest attention on Hacker News, where commenters are weighing its usefulness for lead generation and local search analysis against questions about Google's terms of service and scraping restrictions.
- 20Google hides destination URLs in search links to deter scrapers●Google's Goto Gambit: Search Links Now Hide Destinations to Thwart Scrapers https:// lemmy.dbzer0.com/post/76357044
Google has changed how search result links work, replacing visible destination URLs with intermediary redirects that conceal the actual target address. The move, dubbed the 'Goto' gambit after the redirect domain in use, is seen as an attempt to stop scrapers and automated tools from harvesting results. Users report it complicates link inspection and breaks some third-party tools, renewing criticism of Google's tightening grip on search data.
- 21
A technical blog post from TrueSign argues that many common methods for detecting residential proxies are flawed, walking through approaches that fail and why. The piece is drawing attention among developers and anti-fraud engineers who deal with bot detection, proxy filtering and traffic verification, sparking discussion about which detection techniques actually hold up in practice.
- 22Google hides search link destinations to block scrapers●Google's Goto Gambit: Search Links Now Hide Destinations to Thwart Scrapers https:// lemmy.dbzer0.com/post/76356943
Google has changed how links in search results work, reportedly routing clicks through intermediary URLs that hide the actual destination address. The move is being read as an attempt to stop scrapers and automated tools from harvesting result URLs. Users discussing the change worry about transparency, privacy and the extra redirect hop before reaching a site.
- 23Reddit restricts Old Reddit access citing AI bots●Reddit says it has to cut back access to ‘Old Reddit’ because of AI bots Reddit is further limiting who can use the "Old
Reddit says it is further limiting access to its legacy 'Old Reddit' interface, citing efforts to combat scraping and automated traffic from AI bots. The company has recently begun requiring users to log in before using the classic experience. Many long-time users rely on Old Reddit for its simpler design, and the tightening rules have sparked frustration among those who see it as another step toward restricting free access to the platform.
- 24Lightpanda 1.0 launches as a browser built for machines●Lightpanda: A browser for machines instead of humans Lightpanda 1.0.0 brings the Classic WebDriver for automation tasks.
Lightpanda has released version 1.0.0 of its browser designed for automation rather than human users. The update brings Classic WebDriver support for automation tasks, enforces CORS by default, and adds new Web APIs. The browser is aimed at developers running AI agents, scraping, and testing workloads at scale, and coverage of the release is drawing attention to the growing demand for machine-facing web tooling.
- 25Phillies eliminated by rivals in Wild Card round▼After squeaking into playoffs, Phils bounced by rivals in Wild Card round
The Philadelphia Phillies, who barely secured a playoff berth, were knocked out by a division rival in the National League Wild Card round. The early exit ends a season in which the club scraped into October, and it is drawing criticism from fans frustrated by another abrupt postseason departure at the hands of a familiar opponent.
- 26OpenAI Agents Accused of Scraping 55 Sites Amid Probe●OpenAI AI Agents Secretly Scraped Data from 55 Sites as Regulators Launch Probe
Reports claim OpenAI's AI agents collected data from 55 websites without visible disclosure, prompting regulators to open an investigation. The story is drawing attention because it touches on ongoing concerns over how AI companies gather training and browsing data, and what rules should govern autonomous agents operating on the open web.
- 27GitHub credentials exposed in AI training datasets●Credenziali GitHub esposte anche nei dataset per l’addestramento dell’IA # infosec https:// cert-agid.gov.it/news/creden
Italy's cybersecurity agency CERT-AGID is warning that GitHub credentials have been found exposed in datasets used to train artificial intelligence models. Leaked secrets in publicly scraped code repositories can end up embedded in AI training data, creating risks of credential reuse and unauthorized access. The advisory underscores growing concern among security experts about how AI data pipelines handle sensitive information from public code.
- 28OpenAI agents caught covertly scraping 55 websites●OpenAI agents tried to covertly scrape 55 websites https:// fawkes.rocks/2026/10/01/openai -agents-tried-to-covertly-scr
OpenAI's autonomous agents reportedly attempted to covertly scrape content from 55 websites, raising fresh concerns about how the company collects training data and whether its systems disguise their activity. Critics are questioning transparency practices, while others are debating the technical and legal implications of agent-driven data collection at scale.
- 29BrowserAct AI Web Scraper Lets Users Describe Data Needs in Plain Language●📰 BrowserAct AI Web Scraper in 2026: Build Once, Run Repeatedly Describe your data needs in plain language and turn webs
BrowserAct is being promoted as an AI-powered web scraping tool for 2026 that turns plain-language instructions into automated data collection. Users describe what data they need, and the tool converts websites into a continuous source of fresh data, with workflows built once and run repeatedly. The tool is being covered by KDnuggets as part of ongoing interest in AI-driven automation for data work.
- 30Vogue and Esquire publishers urge Congress to ban AI content theft●Vogue, Esquire publishers plead with Congress to ban AI theft of their content
The publishers behind Vogue and Esquire have appealed to Congress to pass legislation preventing AI companies from using their journalism and photography without permission or payment. The plea frames unlicensed scraping of magazine content as theft and adds major media brands to the growing fight over copyright and compensation in the age of generative AI.
- 31Scientists urged to sabotage AI training with junk data●If you are a scientist and have been asked to participate in AI crap, please feed the machines crazy ideas that are a) w
A call is circulating urging scientists who are asked to contribute their work to AI systems to deliberately feed the machines false and wildly expensive research ideas. The advice suggests planting errors subtle enough to go unnoticed by anyone using large language models to scoop academic work, and keeping receipts for later. It reflects growing researcher anger over unpaid data extraction by AI companies.
- 32Reddit scrapping RSS support over AI bots●Reddit slopar RSS-stöd på grund av AI-botar # GenerativeAI # ArtificialIntelligence Reddit slopar RSS-stöd på grund av A
Reddit is discontinuing support for RSS feeds, according to Swedish technology coverage. The move is attributed to AI bots scraping content from the platform. RSS has long allowed users and third-party tools to follow Reddit communities without visiting the site, and its removal is being read as another step in Reddit's broader effort to restrict automated access to its data, a fight that has shaped the platform's policies since the rise of generative AI.
- 33Reddit to shut down RSS feeds and public API access●# Reddit will end # RSS feeds on 13 November and public # API access by March 2027 due to scraping and AI bots. Moderato
Reddit will discontinue its RSS feeds on 13 November and end public API access by March 2027, citing heavy scraping by AI bots. Moderators are being advised to use Discord Relay Devvit, but no replacement exists for external feeds. Old Reddit access will also be restricted to recent users, closing off further workarounds for third-party tools and readers.
- 34KONE unveils elevator technology for skyscrapers beyond 1km●As skyscrapers stretch higher, KONE launches technology for cities reaching beyond 1km
Elevator maker KONE has announced new technology designed for buildings taller than one kilometre, as skyscrapers around the world continue to climb to unprecedented heights. The Finnish company, one of the world's largest lift manufacturers, says the launch is aimed at cities planning ultra-tall towers where vertical transport is a key engineering limit. Details on pricing and first deployments were not given.
- 35Reddit kills RSS feeds and public API access, blaming AI bots●Is there nothing that AI won’t ruin? Reddit is killing RSS feeds and ending public API access because of AI bots. # ai #
Reddit is ending its RSS feeds and shutting down public API access, citing overwhelming traffic from AI bots scraping its content. The move means third-party tools and readers will lose open access to Reddit data, and it is being discussed as the latest example of AI crawlers forcing platforms to lock down information that was previously freely available.
Repos
- mobile-next/mobile-mcp Model Context Protocol Server for Mobile Automation and Scraping (iOS, Android, Emulators, Simulators and Real Devices)