search
AI evaluators
Trends
- 1NSA reportedly spending billions testing AI models●Classified estimates show the NSA is paying billions to test AI models
Classified budget estimates indicate the National Security Agency is directing billions of dollars toward testing artificial intelligence models, according to a report drawing on leaked or declassified figures. The scale of the spending suggests US intelligence agencies are moving aggressively to evaluate and adopt AI capabilities. Observers are weighing the national security implications and questions of oversight for such secretive AI procurement.
- 278% of U.S. homebuyers now use AI tools when buying a home●78% of U.S. homebuyers now use AI to help buy a home https://www.fastcompany.com/91618486/homebuying-real-estate-agent-s
A new survey indicates that 78% of U.S. homebuyers now use artificial intelligence to help with the home-buying process, from researching listings to evaluating offers. The figure highlights how quickly AI tools have entered real estate, raising questions about the evolving role of traditional agents and how buyers make one of life's biggest financial decisions.
- 3Graphene toolkit brings data analysis to coding agents●Show HN: Graphene – Data analysis toolkit for your coding agent
A new open-source project called Graphene has been launched, offering a data analysis toolkit designed to work with AI coding agents. The toolkit was shared as a launch announcement on Hacker News under its Show HN format, drawing early engagement from the developer community. Details beyond the GitHub repository are limited so far, with users likely evaluating its features and usefulness.
- 4
A developer has published a write-up after spending a month coding with GLM 5.3 Flash, a model from Chinese AI lab Zhipu. The post is drawing attention among developers weighing cheaper, faster models for everyday programming work, and it is being discussed on Hacker News, where readers are sharing their own experiences with budget coding models.
- 5Anthropic flags Chinese AI model over hacking abilities●Anthropic claims popular Chinese AI model has Mythos-class hacking abilities
Anthropic says a popular Chinese open-weight AI model shows 'Mythos-class' hacking capabilities, according to its frontier red-teaming report. The company's findings also detail weak safeguards on open-weight AI systems, raising concerns that such models could be misused for cyberattacks. The report is prompting debate over how open model releases should be evaluated and restricted.
- 6OpenAI publishes mathematical manuscripts and proof artifacts●Mathematical manuscripts and supporting proof artifacts produced by OpenAI
OpenAI has released a repository containing mathematical manuscripts along with supporting proof artifacts produced by its models. The collection gives researchers and mathematicians material to examine how AI systems formulate and verify formal proofs. The release is drawing attention from the AI and mathematics communities, who are assessing the quality and significance of the generated work.
- 7AI evaluators face pressure to keep advanced systems safe▼‘Evaluators’ are supposed to keep AI from killing us all. No pressure
A new report highlights the role of AI evaluators, the specialists tasked with assessing whether advanced AI systems are safe before deployment. Their work is framed as a critical safeguard against catastrophic AI risks, yet it is largely voluntary and under-resourced. Commenters are debating whether this handful of assessors can realistically hold back powerful technology developed by some of the world's richest companies.
- 8Study probes whether AI models judge code morally●Ask a model if code is malicious and it reaches for its morals https://www.manifold.security/blog/do-models-consider-mor
Security firm Manifold Security published research asking whether AI models factor morality into their judgments about malicious code. The finding: when asked to assess whether code is malware, language models appear to bring moral reasoning into their analysis rather than relying purely on technical criteria. The report is circulating among developers and security researchers interested in how AI tools evaluate potentially harmful software.
- 9
Companies are pouring money into artificial intelligence, but measuring the return on that investment is proving elusive across the industry. Many firms struggle to link AI spending to concrete business outcomes, leaving executives uncertain whether the technology is delivering value. The issue is prompting debate over how AI projects should be evaluated and whether current spending levels are justified.
- 10Banana Pi launches RK3576 computing module for Edge AI▼Banana Pi BPI-CM5 Pro is a computing module powered by the Rockchip RK3576, 6 TOPS computing power NPU,8-32G RAM and 8-1
Banana Pi has introduced the BPI-CM5 Pro, a compute module built on Rockchip's RK3576 chip with a 6 TOPS NPU, 8-32GB of RAM and 8-128GB of eMMC storage. The maker positions it as a board for Edge AI projects and an alternative to the Raspberry Pi Compute Module 4, drawing interest from hobbyists and robotics developers.
- 11OpenAI releases 722 math manuscripts●OpenAI releases 722 math manuscripts https://github.com/openai/math/blob/main/CONTENTS.md # HackerNews # Tech # AI
OpenAI has made 722 math manuscripts publicly available through a repository on GitHub, opening up a collection of mathematical texts to researchers and the public. The release is being discussed among developers and AI enthusiasts, who see it as a potentially useful resource for training and evaluating mathematical reasoning in AI models.
- 12Karpathy Offers Tips for Understanding AI Outputs Clearly●Karpathy Shares Tips for Understanding AI Outputs Clearly
Andrej Karpathy, the AI researcher and OpenAI co-founder, has shared practical advice on how to interpret and evaluate the outputs of AI language models more clearly. His guidance, circulated widely on X, covers ways users can check whether model responses are accurate rather than taking them at face value. Readers are discussing the tips as interest grows in how everyday users can judge the reliability of AI-generated answers.
- 13UN University Seeks to Build Independent AI Evaluators Community▼Building a Community of Independent AI Evaluators
United Nations University is working to create a community of independent AI evaluators, aiming to strengthen global oversight of artificial intelligence systems. The initiative points to growing concern that AI development is outpacing reliable, impartial assessment of its risks and impacts. The UN University's role suggests an effort to coordinate expertise across countries and disciplines for credible, independent evaluation capacity.
- 14Frontier AI Models Show Uneven Skill from Web Browsing to Robotics●Astra, Opus 5.5 Demonstrate Jagged Performance on Web to Robotics Tasks
A new evaluation from Fig reports that frontier AI models, including Astra and Opus 5.5, perform unevenly — 'jagged' — across agentic tasks, doing well on some web browsing challenges while struggling on robotics tasks. The findings suggest that even top models lack consistent reliability, with strong results in one domain not predicting competence in another.
- 15TechCrunch Disrupt 2026 reveals Startup Battlefield 200 judges●Meet the Startup Battlefield 200 judges who'll decide the winner at TechCrunch Disrupt 2026 https://techcrunch.com/2026/
TechCrunch has announced the panel of judges for the Startup Battlefield 200 at TechCrunch Disrupt 2026. The judges will evaluate the competing startups and decide which company takes the top prize at the conference. The lineup leans heavily toward technology and AI-focused investors and operators, drawing attention from the startup community ahead of the event.
- 16Developer pits AI coding agents against each other in races●Every few days someone tells me which coding agent is "obviously the best". They always have a... # ai # opensource # sh
A developer has built races between AI coding agents to test which one performs best, challenging the frequent claims that any single agent is "obviously the best". The finding, shared widely in software and open-source communities, was that the fastest agent was not the winner, prompting discussion about how AI coding tools should actually be evaluated.
- 17OpenAI floods mathematics with hundreds of new results●OpenAI unleashes hundreds more math results upon a field already in shock https://www.scientificamerican.com/article/ope
OpenAI has released hundreds of new mathematical results, adding to a wave of AI-generated output that has already unsettled researchers in the field. The announcement, covered by Scientific American, is drawing attention because of the sheer volume of results and ongoing debate about how AI-generated mathematics should be evaluated and verified.
- 18New Proxy Lets AI Models Train Inside Real Coding Harnesses●New Proxy Trains AI Models Inside Real Coding Harnesses Without Changes
A new open-source tool called Proxy allows AI coding models to be trained and evaluated inside real coding harnesses without any modifications to the existing setup. The project aims to bridge the gap between benchmark testing and practical use, letting developers plug models directly into their workflows. Developer communities are discussing its potential to speed up model iteration and testing.
- 19OpenAI Publishes Findings on 377 Math Problems▼OpenAI Releases Findings on 377 Math Problems, Further Roiling Field
OpenAI has released findings based on 377 math problems, an announcement reported by The New York Times as further roiling the artificial intelligence field. The release has drawn attention within the AI research community, where debate continues over model capabilities, evaluation methods and the pace of progress on mathematical reasoning benchmarks.
- 20Developer tests AI help to speed up video compositing●So... asked the AI to evaluate performance on r/t video compositing trying to get processing per frame under 50ms next-u
A developer is working on real-time video compositing and trying to bring processing time per frame under 50 milliseconds, a threshold needed for smooth real-time performance. They asked an AI assistant to evaluate performance and are now working on a strip-parallel compositor that splits the frame across threads, sacrificing four threads in the process. The work is being done while also livestreaming and gaming, and the developer has shared the experiment publicly.
- 21Foundation AI Highlights VLoc Bench and Cyber-Capability Safety▼Foundation AI in September: VLoc Bench and Cyber-Capability Safety
Cisco's Foundation AI team published its September update, spotlighting the VLoc Bench and new work on cyber-capability safety. The release covers how the team benchmarks and evaluates AI models for security-relevant capabilities, part of a broader push to make sure advanced AI systems do not amplify cyber threats. Readers in the AI security community are discussing what the benchmarks mean for safety testing.
- 22Federal AI portal scores 9 out of 12 on Bitcoin policy quiz●🤖 America .gov, el portal federal con # IA de # Google y # SpaceXAI , acertó 9 de 12 preguntas sobre política de Bitcoin
Bitcoin.com News evaluated America.gov, a federal portal using AI from Google and SpaceX, on Bitcoin policy questions. The portal answered 9 of 12 questions correctly, performing less accurately than ChatGPT and Claude in the same test. The comparison highlights how leading commercial AI chatbots stack up on cryptocurrency policy knowledge, with a government-backed tool trailing behind.
- 23AI Ready Roanoke Assesses New Technology Use in the Roanoke Valley▼AI Ready Roanoke evaluates use for new technology in Roanoke Valley
AI Ready Roanoke is evaluating how new technology, particularly artificial intelligence, could be used across the Roanoke Valley in Virginia. The initiative, covered by the Roanoke Times, is examining practical applications for local businesses, institutions, and residents as AI adoption spreads. Details on specific projects, partners, and timelines have not yet been reported.
- 24Innodata Opens Robot Data Lab, Investors Watch Closely▼How Investors May Respond To Innodata (INOD) Opening Robot Data Lab
Data engineering company Innodata has opened a new lab focused on robot training data, drawing attention from investors assessing what the move means for the firm's growth strategy. The company, known for AI data services, is positioning itself in the emerging market for data used to train robotics systems. Market watchers are weighing whether the lab can translate into revenue and how it might affect the stock's performance going forward.
- 25White House forms task force on AI risks▼White House forms AI task force to assess technology risks - WSJ
The White House has created a task force to assess the risks of artificial intelligence, according to the Wall Street Journal. The group will evaluate potential dangers from rapidly advancing AI technology, part of a broader effort by the US administration to develop policy on AI safety and regulation.
- 26Boston VA workshop explores AI for veteran care▼Boston VA workshop tests AI’s potential for Veteran care
The Department of Veterans Affairs held a workshop in Boston examining how artificial intelligence could improve care for veterans. The event tested potential applications of AI within VA healthcare services. It reflects the agency's broader effort to evaluate whether AI tools can support clinical decisions, administration, and patient outcomes for former service members.
- 27
Call centre and service workers are increasingly evaluated by AI systems that score their calls automatically, while the criteria behind those scores remain hidden from employees. Commentators argue this opaque algorithmic management gives employers sweeping power over performance reviews and discipline without workers being able to challenge or even understand how they are judged, fuelling demands for transparency rules and stronger workplace AI regulation.
- 28
The American Hospital Association has published an overview of how nurses are using artificial intelligence in their work, presenting the available data in chart form. The roundup highlights adoption patterns and practical applications of AI tools in nursing practice, a topic of growing interest as hospitals weigh efficiency gains against concerns over patient safety and job roles.
- 29As A.I. Agents Begin Shopping, Brands Rethink Their Sales Pitch▼As A.I. Agents Begin Shopping, Brands Are Changing Their Sales Pitch
Brands are adjusting how they market products as artificial intelligence agents increasingly shop on behalf of consumers. Instead of pitching to human shoppers with emotional advertising, companies are optimizing product data and descriptions so AI assistants can find, evaluate and recommend their goods, a shift with significant implications for e-commerce and digital marketing.
- 30Darwin-180B-RSI tops three AI benchmarks●Darwin-180B-RSI suma el puesto # 1 en MDPBench, ExtractBench e IFStruct y alcanza 10 primeros puestos oficiales. Y ZTC j
The open-source Spanish-language model Darwin-180B-RSI has reached first place on the MDPBench, ExtractBench and IFStruct leaderboards, bringing its total to ten official top rankings. Its ZTC component is noted for evaluating output without generating text itself. Commenters in the AI and open-source community are highlighting the model's benchmark sweep and its relevance for Spanish-language software development.
- 31
A debate is under way in the scientific community over whether artificial intelligence should be used in the peer review of research papers. Supporters see potential to speed up reviews and ease reviewer shortages, while critics worry about bias, confidentiality of unpublished manuscripts, and the risk of undermining trust in scholarly publishing.
- 32Researchers propose fixing GRPO's credit assignment problem●Fixing GRPO's credit assignment problem without evaluating every step
A new paper on arXiv proposes a way to fix the credit assignment problem in GRPO, a reinforcement learning method used for training language models, without evaluating every step of a response. GRPO currently assigns the same reward to all tokens in a completion, making it hard to identify which parts of an output earned the reward. The proposed approach aims to improve this at lower computational cost.
- 33OpenAI partners with Ironclad on AI contracting agents●🤖 Advancing computer use with Ironclad Learn how OpenAI and Ironclad are training and evaluating AI agents on complex co
OpenAI announced a partnership with contract management company Ironclad to train and evaluate AI agents on complex contracting workflows. The collaboration aims to advance computer-use AI for professional work, using Ironclad's real-world legal and contracting tasks as a testing ground. It signals a push toward deploying AI agents in specialized business environments beyond general browsing.
- 34Your company's AI needs a scoreboard●Your company's AI needs a scoreboard https://www.fastcompany.com/91614445/your-companys-ai-needs-a-scoreboard # AI # Bus
A Fast Company article argues that companies adopting AI need clear measurement systems—scoreboards—to track whether their AI investments are actually delivering results. The piece taps into a wider business debate about how firms should evaluate AI tools beyond hype, using metrics and benchmarks to judge performance and justify spending.
- 35Venture Investors Turn to Judgment Where AI Data Falls Short▼Once AI Reads the Deck, Venture Investors Test What the Data Cannot Explain
Venture capital investors are reportedly exploring what AI cannot explain when evaluating startups, even as AI tools increasingly analyse pitch decks and company data. The discussion centres on whether human judgment still matters in investment decisions once machines can process the numbers. Details of specific firms or funds involved remain limited, and reaction from the wider venture community is not yet clear.
- 36Companies begin allowing AI agents in coding interviews●More companies now let candidates use an AI agent during coding interviews. That sounds like it makes... # ai # career #
A growing number of employers are permitting job candidates to use AI coding agents during technical interviews. The shift changes what is being assessed: rather than writing code unaided, applicants must show they can direct and work with AI tools effectively. Commenters note this mirrors real-world engineering work, where AI assistance is now standard, and raises questions about how to evaluate core programming ability.
- 37
A widely shared essay argues that in the era of AI agents, the harness—the scaffolding of tools, prompts, evaluation and workflow code wrapped around a model—is where a company's actual value sits, not the underlying model itself. As models become commoditised and interchangeable, the author contends the harness is the durable product, and effectively the company's true identity.
- 38AI models ranked by normative score across twelve paradigms●Models ranked by normative score across twelve paradigms, once neutral and once human-primed https://cognit.rajtilak.tec
A new write-up ranks AI language models by their normative score, tested across twelve different paradigms and repeated twice — once under neutral conditions and once with human priming. The comparison aims to show how priming changes model behavior and judgments, offering a benchmark-style look at consistency across evaluation setups.
- 39New benchmark tests AI agents on messy company knowledge●Benchmarking retrieval for agents on messy real-world company knowledge
Kapa.ai has published a benchmark for evaluating how well retrieval systems let AI agents work with messy, real-world company knowledge bases. The release is drawing attention among developers and AI practitioners, who are discussing how enterprise search and agent performance should be measured outside clean, curated datasets, where documentation is inconsistent, outdated or scattered across tools.
- 40Law firm publishes guide to AI insurance coverage●Keep Your AI On The Ball: A Policyholder’s Guide To Artificial Intelligence Insurance Coverage
New York law firm Pryor Cashman has published a policyholder's guide to insurance coverage for artificial intelligence, explaining how existing policies may or may not respond to AI-related risks and losses. The guidance walks businesses through evaluating whether their current coverage extends to AI tools, what exclusions to watch for, and how to negotiate better protection. It reflects growing demand from companies using AI who are unsure whether their insurers will cover AI-related liability.
Repos
- archestra-ai/OpenAPPA Deterministic guardrails that don't break agents
- pallavi-shekhar/ai-engineering-interview-questions-company-wise Your Cheat Sheet For AI Engineering Interviews at Top AI Companies - Questions and Answers.
- amitshekhariitbhu/ai-system-design AI System Design - Learn how to design AI systems built on LLMs, RAG, and AI Agents step by step.
- zhengkid/Dream-RSI The offical repo for "Dream-RSI: Recursive Self-Improvement through Evolving Worlds"