search
AI evaluators
Trends
- 1
The trycua/cua project, written in Rust, offers open-source drivers for scaling 'computer-use 2.0' agents across operating systems, with tools for building fleets of machines and benchmarks for training, evaluation and data generation. Developers are sharing and engaging with the repository, reflecting growing interest in infrastructure for AI agents that control computers.
- 2AI agents find two room-temperature magnetic semiconductor candidates●Opus 5.5 agents discover two room-temperature magnetic semiconductor candidates
Anthropic's Opus 5.5 model, running as autonomous research agents, has identified two candidate materials for room-temperature magnetic semiconductors, according to a report by evaluation firm vals.ai. The materials could enable spintronic devices that work without cryogenic cooling. Discussion is focused on whether AI-driven discovery can genuinely accelerate materials science or whether the candidates will survive experimental validation.
- 3
RSM is asking businesses whether their AI investments are delivering genuine growth rather than just experimentation. The piece prompts companies to evaluate whether artificial intelligence initiatives are producing measurable returns, reflecting a broader debate about moving from AI pilots to tangible business value.
- 4CDT Signs Coalition Comments on NIST AI Evaluation Framework●CDT Joins Coalition Comments on NIST’s AI Evaluation Framework
The Center for Democracy and Technology has joined a coalition in submitting comments to NIST on its AI evaluation framework. The filing reflects growing pressure on the US standards body to shape how artificial intelligence systems are tested and measured. Details of the coalition's specific recommendations were not disclosed, but the move underscores mounting civil society engagement in federal AI policy work.
- 5New Evaluation Says ChatGPT Poses Risk for Teens●ChatGPT for Teens Poses ‘Unacceptable Risk’ for Children Under 18, Says New Evaluation
A new evaluation concludes that ChatGPT's teen offering poses an 'unacceptable risk' for children under 18. The assessment adds to growing scrutiny of how AI chatbots handle younger users, with parents and child-safety advocates raising concerns about exposure to harmful content and inadequate age protections. OpenAI has faced repeated pressure to strengthen safeguards for minors using its products.
- 6Newsroom AI emotional optimization needs behavioural evaluation▼Emotional optimization by newsroom AI needs behavioural evaluation
Researchers writing in Nature argue that artificial intelligence systems used in newsrooms to optimize emotional responses in audiences require systematic behavioural evaluation. The call highlights growing concern that AI tools shaping how news makes readers feel are being deployed without rigorous testing of their psychological and societal effects, and the authors urge news organizations and regulators to establish evaluation frameworks before these systems become more widespread.
- 7Hermes Agent Rolls Out Major Developer Tool Updates●Hermes Agent Doubles Down on Developer Tools with Major Updates
Hermes Agent has announced a substantial set of updates to its developer tools, signaling a renewed focus on serving the developer community. The release expands the platform's capabilities for building and deploying agentic AI applications. Developers are discussing the changes online, with many weighing how the updates compare to competing agent frameworks and what they mean for ongoing projects.
- 8UT San Antonio launches certificate on critical AI use for educators●New UT San Antonio certificate helps educators approach AI with a critical eye
The University of Texas at San Antonio has introduced a new certificate program designed to help educators engage with artificial intelligence critically. The program aims to equip teachers with the skills to evaluate AI tools and their classroom implications rather than adopt them uncritically. It reflects a broader push in higher education to prepare teachers for the rapid spread of AI in schools.
Repos
- trycua/cua Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data gene
- amitshekhariitbhu/ai-system-design AI System Design - Learn how to design AI systems built on LLMs, RAG, and AI Agents step by step.
- archestra-ai/OpenAPPA Deterministic guardrails that don't break agents
- pallavi-shekhar/ai-engineering-interview-questions-company-wise Your Cheat Sheet For AI Engineering Interviews at Top AI Companies - Questions and Answers.