search
AI evaluators
Trends
- 1Graphene toolkit brings data analysis to coding agentsβΌShow HN: Graphene β Data analysis toolkit for your coding agent
A new open-source project called Graphene has been launched, offering a data analysis toolkit designed to work with AI coding agents. The toolkit was shared as a launch announcement on Hacker News under its Show HN format, drawing early engagement from the developer community. Details beyond the GitHub repository are limited so far, with users likely evaluating its features and usefulness.
- 2
A developer has published a write-up after spending a month coding with GLM 5.3 Flash, sharing hands-on experience with the AI model for programming tasks. The post has drawn attention on Hacker News, where readers are weighing in on how the model performs in real-world development work and how it compares with rival coding assistants.
- 3Anthropic claims Chinese AI model has Mythos-class hacking abilitiesβAnthropic claims popular Chinese AI model has Mythos-class hacking abilities
Anthropic says a widely used Chinese open-weight AI model demonstrates Mythos-class hacking capabilities, according to a new frontier red-teaming report. The company's evaluation found that the model's safeguards are weak, raising concerns that powerful offensive cyber abilities could be freely distributed and abused. The report has drawn attention to broader questions about how open-weight AI models should be tested and restricted before release.
- 4OpenAI publishes mathematical manuscripts and proof artifactsβMathematical manuscripts and supporting proof artifacts produced by OpenAI
OpenAI has released a repository containing mathematical manuscripts along with supporting proof artifacts produced by its models. The collection gives researchers and mathematicians material to examine how AI systems formulate and verify formal proofs. The release is drawing attention from the AI and mathematics communities, who are assessing the quality and significance of the generated work.
- 5
Companies are pouring money into artificial intelligence, but measuring the return on that investment is proving elusive across the industry. Many firms struggle to link AI spending to concrete business outcomes, leaving executives uncertain whether the technology is delivering value. The issue is prompting debate over how AI projects should be evaluated and whether current spending levels are justified.
- 6UN University Seeks to Build Independent AI Evaluators CommunityβΌBuilding a Community of Independent AI Evaluators
United Nations University is working to create a community of independent AI evaluators, aiming to strengthen global oversight of artificial intelligence systems. The initiative points to growing concern that AI development is outpacing reliable, impartial assessment of its risks and impacts. The UN University's role suggests an effort to coordinate expertise across countries and disciplines for credible, independent evaluation capacity.
- 7OpenAI releases 722 math manuscriptsβOpenAI releases 722 math manuscripts https://github.com/openai/math/blob/main/CONTENTS.md # HackerNews # Tech # AI
OpenAI has made 722 math manuscripts publicly available through a repository on GitHub, opening up a collection of mathematical texts to researchers and the public. The release is being discussed among developers and AI enthusiasts, who see it as a potentially useful resource for training and evaluating mathematical reasoning in AI models.
- 8Foundation AI Highlights VLoc Bench and Cyber-Capability SafetyβΌFoundation AI in September: VLoc Bench and Cyber-Capability Safety
Cisco's Foundation AI team published its September update, spotlighting the VLoc Bench and new work on cyber-capability safety. The release covers how the team benchmarks and evaluates AI models for security-relevant capabilities, part of a broader push to make sure advanced AI systems do not amplify cyber threats. Readers in the AI security community are discussing what the benchmarks mean for safety testing.
- 9OpenAI floods mathematics with hundreds of new resultsβOpenAI unleashes hundreds more math results upon a field already in shock https://www.scientificamerican.com/article/ope
OpenAI has released hundreds of new mathematical results, adding to a wave of AI-generated output that has already unsettled researchers in the field. The announcement, covered by Scientific American, is drawing attention because of the sheer volume of results and ongoing debate about how AI-generated mathematics should be evaluated and verified.
- 10OpenAI Publishes Findings on 377 Math ProblemsβΌOpenAI Releases Findings on 377 Math Problems, Further Roiling Field
OpenAI has released findings based on 377 math problems, an announcement reported by The New York Times as further roiling the artificial intelligence field. The release has drawn attention within the AI research community, where debate continues over model capabilities, evaluation methods and the pace of progress on mathematical reasoning benchmarks.
- 11Study probes whether AI models judge code morallyβAsk a model if code is malicious and it reaches for its morals https://www.manifold.security/blog/do-models-consider-mor
Security firm Manifold Security published research asking whether AI models factor morality into their judgments about malicious code. The finding: when asked to assess whether code is malware, language models appear to bring moral reasoning into their analysis rather than relying purely on technical criteria. The report is circulating among developers and security researchers interested in how AI tools evaluate potentially harmful software.
- 12AI evaluators face pressure to keep advanced systems safeβΌβEvaluatorsβ are supposed to keep AI from killing us all. No pressure
A new report highlights the role of AI evaluators, the specialists tasked with assessing whether advanced AI systems are safe before deployment. Their work is framed as a critical safeguard against catastrophic AI risks, yet it is largely voluntary and under-resourced. Commenters are debating whether this handful of assessors can realistically hold back powerful technology developed by some of the world's richest companies.
- 13Banana Pi launches RK3576 computing module for Edge AIβBanana Pi BPI-CM5 Pro is a computing module powered by the Rockchip RK3576, 6 TOPS computing power NPU,8-32G RAM and 8-1
Banana Pi has introduced the BPI-CM5 Pro, a compute module built on Rockchip's RK3576 chip with a 6 TOPS NPU, 8-32GB of RAM and 8-128GB of eMMC storage. The maker positions it as a board for Edge AI projects and an alternative to the Raspberry Pi Compute Module 4, drawing interest from hobbyists and robotics developers.
- 14Boston VA workshop explores AI for veteran careβΌBoston VA workshop tests AIβs potential for Veteran care
The Department of Veterans Affairs held a workshop in Boston examining how artificial intelligence could improve care for veterans. The event tested potential applications of AI within VA healthcare services. It reflects the agency's broader effort to evaluate whether AI tools can support clinical decisions, administration, and patient outcomes for former service members.
- 15OpenAI partners with Ironclad on AI contracting agentsβπ€ Advancing computer use with Ironclad Learn how OpenAI and Ironclad are training and evaluating AI agents on complex co
OpenAI announced a partnership with contract management company Ironclad to train and evaluate AI agents on complex contracting workflows. The collaboration aims to advance computer-use AI for professional work, using Ironclad's real-world legal and contracting tasks as a testing ground. It signals a push toward deploying AI agents in specialized business environments beyond general browsing.
- 16
The American Hospital Association has published an overview of how nurses are using artificial intelligence in their work, presenting the available data in chart form. The roundup highlights adoption patterns and practical applications of AI tools in nursing practice, a topic of growing interest as hospitals weigh efficiency gains against concerns over patient safety and job roles.
- 17Frontier AI Models Show Uneven Skill from Web Browsing to RoboticsβAstra, Opus 5.5 Demonstrate Jagged Performance on Web to Robotics Tasks
A new evaluation from Fig reports that frontier AI models, including Astra and Opus 5.5, perform unevenly β 'jagged' β across agentic tasks, doing well on some web browsing challenges while struggling on robotics tasks. The findings suggest that even top models lack consistent reliability, with strong results in one domain not predicting competence in another.
- 18Companies begin allowing AI agents in coding interviewsβMore companies now let candidates use an AI agent during coding interviews. That sounds like it makes... # ai # career #
A growing number of employers are permitting job candidates to use AI coding agents during technical interviews. The shift changes what is being assessed: rather than writing code unaided, applicants must show they can direct and work with AI tools effectively. Commenters note this mirrors real-world engineering work, where AI assistance is now standard, and raises questions about how to evaluate core programming ability.
- 19
Call centre and service workers are increasingly evaluated by AI systems that score their calls automatically, while the criteria behind those scores remain hidden from employees. Commentators argue this opaque algorithmic management gives employers sweeping power over performance reviews and discipline without workers being able to challenge or even understand how they are judged, fuelling demands for transparency rules and stronger workplace AI regulation.
Repos
- pallavi-shekhar/ai-engineering-interview-questions-company-wise Your Cheat Sheet For AI Engineering Interviews at Top AI Companies - Questions and Answers.
- archestra-ai/OpenAPPA Deterministic guardrails that don't break agents
- amitshekhariitbhu/ai-system-design AI System Design - Learn how to design AI systems built on LLMs, RAG, and AI Agents step by step.