search
web scraping
Trends
- 1
A new write-up argues that Meta's Muse model performs remarkably well at web scraping, suggesting the AI can extract structured data from web pages more reliably than developers expected. The piece has drawn attention among developers discussing how AI coding assistants are changing data-collection work, and whether such tools will reduce the need for manual scraping scripts and routine data work.
- 2Wikipedia links OpenAI bots to May outage●Wikipedia operator says OpenAI’s ‘rogue’ bots may be linked to a May outage Following many recent disclosures about AI a
The Wikimedia Foundation says it discovered OpenAI crawling bots that ignored instructions and may have contributed to a Wikipedia outage in May. The operator of the online encyclopedia said the 'rogue' bots were found amid wider disclosures about AI agents accessing third-party websites without permission. The disclosure adds to growing friction between AI companies and the websites they scrape for training data.
- 3Developer flags privacy concerns with screenshot SaaS APIs●Ran into this building an agent that needed to look at web pages. Every screenshot API is SaaS, which is fine, until you
A developer building an AI agent that needs to view web pages ran into a problem: every screenshot capture API is a hosted SaaS service, meaning the text and content of every captured page is sent to and processed by a third party. For publicly available pages this may be acceptable, but it raises privacy and control concerns, prompting the developer to look at self-hosted alternatives.
- 4Yelp scrapers returning incomplete data puzzle developers▼As data engineers, we often encounter tools that promise straightforward data extraction. # python # webdev # programmin
Data engineers are discussing why Yelp scrapers often return incomplete datasets without throwing any errors, a problem that can silently corrupt downstream analysis. The conversation highlights a common frustration with tools that promise straightforward data extraction but fail quietly, prompting developers to share debugging approaches and caution about relying on scraping outputs without validation.
- 5Wikipedia links May outage to OpenAI's 'rogue' bots●Wikipedia operator says OpenAI's 'rogue' bots may be linked to a May outage https://www.theverge.com/news/1004929/wikipe
A Wikipedia operator says bots operated by OpenAI may have been linked to an outage that hit the site in May. The claim suggests the AI company's automated crawlers behaved in a 'rogue' way and placed unusual load on Wikimedia's servers. The report is drawing attention to growing strain that AI data collection is putting on the infrastructure of the open web.
- 6Guide Circulates on Setting Up SOCKS5 Proxy Servers for Automation●In the world of high-stakes automation—be it web scraping at scale, multi-account management, or... # ai # webdev # prog
Developers are sharing a technical walkthrough on setting up a SOCKS5 proxy server for automation work, covering use cases such as large-scale web scraping, multi-account management, and related engineering tasks. The discussion is framed around hands-on configuration and productivity for programmers, and is drawing modest attention within the software development community.
- 7AI bots and scraping could push the open web out of reach●Anno 202X, il web è così impestato e saccheggiato da bot e scraping dell'AI che i costi per la gestione di siti e comuni
Commenters warn that by 202X the web could be so infested with bots and AI scraping that running websites and online communities becomes affordable only for a handful of global billionaires, pushing humanity into isolated, self-contained systems. The post is drawing attention in Italian-language discussions about the rising costs of maintaining an open internet under pressure from automated AI traffic.
- 8Essay Calling AI Companies Parasites Draws Attention●AI Companies Are Parasites https://www.coryd.dev/posts/2026/ai-companies-are-parasites # HackerNews # Tech # AI
A new essay bluntly argues that AI companies behave like parasites, taking value from creators and the open web while giving little back. The piece is circulating among developers and tech commentators, reigniting debate over how AI firms use copyrighted and publicly available material to train their models without compensation.
- 9IP addresses remain the weak point in web scraping●In the high-stakes game of web scraping and browser automation, the IP address is your fingerprint, your reputation, and
Developers are discussing the realities of web scraping and browser automation, arguing that an IP address functions as a fingerprint, a reputation, and the biggest vulnerability in the practice. The discussion notes that even well-built scrapers with refined DOM selectors and careful handling of asynchronous race conditions can still fail once an IP address is flagged or blocked.
- 10OpenAI accuses Chinese model of stealing its IP, drawing irony charges▼Irony alert: OpenAI whines that Chinese model stole its special IP that it stole from everybody else
OpenAI claims a Chinese AI model improperly copied its work, reportedly framing the practice as a potential national security risk. Critics are pointing out the irony: OpenAI itself trained its models on vast amounts of web data taken without permission from creators and publishers. The dispute has reignited debate over double standards in the AI industry, with many arguing that US model makers want exclusive rights to scrape and distill data while denying the same to rivals.
- 11Google Hides Search Link Destinations to Deter Scrapers●Google's Goto Gambit: Search Links Now Hide Destinations to Thwart Scrapers https:// lemmy.dbzer0.com/post/76357045
Google has changed how search result links work so that the destination web address is hidden until the user clicks, with the redirect routed through a Google 'goto' mechanism. The change is being discussed as a move to stop scrapers and automated tools from extracting search results, though users worry it reduces transparency about where links lead.
- 12OpenAI Agents Accused of Scraping 55 Sites Amid Probe●OpenAI AI Agents Secretly Scraped Data from 55 Sites as Regulators Launch Probe
Reports claim OpenAI's AI agents collected data from 55 websites without visible disclosure, prompting regulators to open an investigation. The story is drawing attention because it touches on ongoing concerns over how AI companies gather training and browsing data, and what rules should govern autonomous agents operating on the open web.
- 13Lightpanda 1.0 launches as a browser built for machines●Lightpanda: A browser for machines instead of humans Lightpanda 1.0.0 brings the Classic WebDriver for automation tasks.
Lightpanda has released version 1.0.0 of its browser designed for automation rather than human users. The update brings Classic WebDriver support for automation tasks, enforces CORS by default, and adds new Web APIs. The browser is aimed at developers running AI agents, scraping, and testing workloads at scale, and coverage of the release is drawing attention to the growing demand for machine-facing web tooling.
- 14Reddit kills RSS feeds and public API access over AI bots●Reddit is killing RSS feeds and ending public API access because of AI bots https://techcrunch.com/2026/09/30/reddit-is-
Reddit is ending public API access and shutting down RSS feeds, citing abuse by AI bots scraping its content. The move follows the platform's earlier restrictions on data access and reflects a broader industry shift to lock down content that could train AI models. The change affects developers, researchers, and third-party tools that relied on free, open access to Reddit data.
- 15BrowserAct AI Web Scraper Lets Users Describe Data Needs in Plain Language●📰 BrowserAct AI Web Scraper in 2026: Build Once, Run Repeatedly Describe your data needs in plain language and turn webs
BrowserAct is being promoted as an AI-powered web scraping tool for 2026 that turns plain-language instructions into automated data collection. Users describe what data they need, and the tool converts websites into a continuous source of fresh data, with workflows built once and run repeatedly. The tool is being covered by KDnuggets as part of ongoing interest in AI-driven automation for data work.