MikeTrendsTrends right now

search

AI evaluators

Trends

  1. 1
    NSA reportedly spending billions testing AI models●Classified estimates show the NSA is paying billions to test AI modelsYhnTechnology1784 d ago

    Classified budget estimates indicate the National Security Agency is directing billions of dollars toward testing artificial intelligence models, according to a report drawing on leaked or declassified figures. The scale of the spending suggests US intelligence agencies are moving aggressively to evaluate and adopt AI capabilities. Observers are weighing the national security implications and questions of oversight for such secretive AI procurement.

  2. 2
    78% of U.S. homebuyers now use AI tools when buying a home●78% of U.S. homebuyers now use AI to help buy a home https://www.fastcompany.com/91618486/homebuying-real-estate-agent-sMmastodonBusinessReal Estate319 min ago

    A new survey indicates that 78% of U.S. homebuyers now use artificial intelligence to help with the home-buying process, from researching listings to evaluating offers. The figure highlights how quickly AI tools have entered real estate, raising questions about the evolving role of traditional agents and how buyers make one of life's biggest financial decisions.

  3. 3
    Graphene toolkit brings data analysis to coding agents●Show HN: Graphene – Data analysis toolkit for your coding agentYhnEnvironmentOceans3251 min ago

    A new open-source project called Graphene has been launched, offering a data analysis toolkit designed to work with AI coding agents. The toolkit was shared as a launch announcement on Hacker News under its Show HN format, drawing early engagement from the developer community. Details beyond the GitHub repository are limited so far, with users likely evaluating its features and usefulness.

  4. 4

    A developer has published a write-up after spending a month coding with GLM 5.3 Flash, a model from Chinese AI lab Zhipu. The post is drawing attention among developers weighing cheaper, faster models for everyday programming work, and it is being discussed on Hacker News, where readers are sharing their own experiences with budget coding models.

  5. 5
    Anthropic flags Chinese AI model over hacking abilities●Anthropic claims popular Chinese AI model has Mythos-class hacking abilitiesYhnBusinessRetail618 min ago

    Anthropic says a popular Chinese open-weight AI model shows 'Mythos-class' hacking capabilities, according to its frontier red-teaming report. The company's findings also detail weak safeguards on open-weight AI systems, raising concerns that such models could be misused for cyberattacks. The report is prompting debate over how open model releases should be evaluated and restricted.

  6. 6
    OpenAI publishes mathematical manuscripts and proof artifacts●Mathematical manuscripts and supporting proof artifacts produced by OpenAIYhnCultureArt4256 min ago

    OpenAI has released a repository containing mathematical manuscripts along with supporting proof artifacts produced by its models. The collection gives researchers and mathematicians material to examine how AI systems formulate and verify formal proofs. The release is drawing attention from the AI and mathematics communities, who are assessing the quality and significance of the generated work.

  7. 7
    AI evaluators face pressure to keep advanced systems safe▼‘Evaluators’ are supposed to keep AI from killing us all. No pressure✉newsTechnology9 h ago

    A new report highlights the role of AI evaluators, the specialists tasked with assessing whether advanced AI systems are safe before deployment. Their work is framed as a critical safeguard against catastrophic AI risks, yet it is largely voluntary and under-resourced. Commenters are debating whether this handful of assessors can realistically hold back powerful technology developed by some of the world's richest companies.

  8. 8
    Study probes whether AI models judge code morally●Ask a model if code is malicious and it reaches for its morals https://www.manifold.security/blog/do-models-consider-morMmastodonTechnology47 h ago

    Security firm Manifold Security published research asking whether AI models factor morality into their judgments about malicious code. The finding: when asked to assess whether code is malware, language models appear to bring moral reasoning into their analysis rather than relying purely on technical criteria. The report is circulating among developers and security researchers interested in how AI tools evaluate potentially harmful software.

  9. 9

    Companies are pouring money into artificial intelligence, but measuring the return on that investment is proving elusive across the industry. Many firms struggle to link AI spending to concrete business outcomes, leaving executives uncertain whether the technology is delivering value. The issue is prompting debate over how AI projects should be evaluated and whether current spending levels are justified.

  10. 10
    Banana Pi launches RK3576 computing module for Edge AI▼Banana Pi BPI-CM5 Pro is a computing module powered by the Rockchip RK3576, 6 TOPS computing power NPU,8-32G RAM and 8-1MmastodonTechnologyRobotics211 h ago

    Banana Pi has introduced the BPI-CM5 Pro, a compute module built on Rockchip's RK3576 chip with a 6 TOPS NPU, 8-32GB of RAM and 8-128GB of eMMC storage. The maker positions it as a board for Edge AI projects and an alternative to the Raspberry Pi Compute Module 4, drawing interest from hobbyists and robotics developers.

  11. 11
    OpenAI releases 722 math manuscripts●OpenAI releases 722 math manuscripts https://github.com/openai/math/blob/main/CONTENTS.md # HackerNews # Tech # AIMmastodonTechnology414 min ago

    OpenAI has made 722 math manuscripts publicly available through a repository on GitHub, opening up a collection of mathematical texts to researchers and the public. The release is being discussed among developers and AI enthusiasts, who see it as a potentially useful resource for training and evaluating mathematical reasoning in AI models.

  12. 12
    Karpathy Offers Tips for Understanding AI Outputs Clearly●Karpathy Shares Tips for Understanding AI Outputs Clearly𝕏xSE2K3 d ago

    Andrej Karpathy, the AI researcher and OpenAI co-founder, has shared practical advice on how to interpret and evaluate the outputs of AI language models more clearly. His guidance, circulated widely on X, covers ways users can check whether model responses are accurate rather than taking them at face value. Readers are discussing the tips as interest grows in how everyday users can judge the reliability of AI-generated answers.

  13. 13
    UN University Seeks to Build Independent AI Evaluators Community▼Building a Community of Independent AI Evaluators✉newsWorldUnited Nations39 min ago

    United Nations University is working to create a community of independent AI evaluators, aiming to strengthen global oversight of artificial intelligence systems. The initiative points to growing concern that AI development is outpacing reliable, impartial assessment of its risks and impacts. The UN University's role suggests an effort to coordinate expertise across countries and disciplines for credible, independent evaluation capacity.

  14. 14
    Frontier AI Models Show Uneven Skill from Web Browsing to Robotics●Astra, Opus 5.5 Demonstrate Jagged Performance on Web to Robotics TasksYhnTechnologyRobotics816 h ago

    A new evaluation from Fig reports that frontier AI models, including Astra and Opus 5.5, perform unevenly — 'jagged' — across agentic tasks, doing well on some web browsing challenges while struggling on robotics tasks. The findings suggest that even top models lack consistent reliability, with strong results in one domain not predicting competence in another.

  15. 15
    TechCrunch Disrupt 2026 reveals Startup Battlefield 200 judges●Meet the Startup Battlefield 200 judges who'll decide the winner at TechCrunch Disrupt 2026 https://techcrunch.com/2026/MmastodonBusinessStartups31 d ago

    TechCrunch has announced the panel of judges for the Startup Battlefield 200 at TechCrunch Disrupt 2026. The judges will evaluate the competing startups and decide which company takes the top prize at the conference. The lineup leans heavily toward technology and AI-focused investors and operators, drawing attention from the startup community ahead of the event.

  16. 16
    Developer pits AI coding agents against each other in races●Every few days someone tells me which coding agent is "obviously the best". They always have a... # ai # opensource # shMmastodonTechnologySoftware41 d ago

    A developer has built races between AI coding agents to test which one performs best, challenging the frequent claims that any single agent is "obviously the best". The finding, shared widely in software and open-source communities, was that the fastest agent was not the winner, prompting discussion about how AI coding tools should actually be evaluated.

  17. 17
    OpenAI floods mathematics with hundreds of new results●OpenAI unleashes hundreds more math results upon a field already in shock https://www.scientificamerican.com/article/opeMmastodonScience41 h ago

    OpenAI has released hundreds of new mathematical results, adding to a wave of AI-generated output that has already unsettled researchers in the field. The announcement, covered by Scientific American, is drawing attention because of the sheer volume of results and ongoing debate about how AI-generated mathematics should be evaluated and verified.

  18. 18
    New Proxy Lets AI Models Train Inside Real Coding Harnesses●New Proxy Trains AI Models Inside Real Coding Harnesses Without Changes𝕏xSE1381 d ago

    A new open-source tool called Proxy allows AI coding models to be trained and evaluated inside real coding harnesses without any modifications to the existing setup. The project aims to bridge the gap between benchmark testing and practical use, letting developers plug models directly into their workflows. Developer communities are discussing its potential to speed up model iteration and testing.

  19. 19
    OpenAI Publishes Findings on 377 Math Problems▼OpenAI Releases Findings on 377 Math Problems, Further Roiling Field✉newsTechnologyAI1 h ago

    OpenAI has released findings based on 377 math problems, an announcement reported by The New York Times as further roiling the artificial intelligence field. The release has drawn attention within the AI research community, where debate continues over model capabilities, evaluation methods and the pace of progress on mathematical reasoning benchmarks.

  20. 20
    Developer tests AI help to speed up video compositing●So... asked the AI to evaluate performance on r/t video compositing trying to get processing per frame under 50ms next-uMmastodonCultureGaming11 d ago

    A developer is working on real-time video compositing and trying to bring processing time per frame under 50 milliseconds, a threshold needed for smooth real-time performance. They asked an AI assistant to evaluate performance and are now working on a strip-parallel compositor that splits the frame across threads, sacrificing four threads in the process. The work is being done while also livestreaming and gaming, and the developer has shared the experiment publicly.

  21. 21
    Foundation AI Highlights VLoc Bench and Cyber-Capability Safety▼Foundation AI in September: VLoc Bench and Cyber-Capability Safety✉newsTechnologyAI14 min ago

    Cisco's Foundation AI team published its September update, spotlighting the VLoc Bench and new work on cyber-capability safety. The release covers how the team benchmarks and evaluates AI models for security-relevant capabilities, part of a broader push to make sure advanced AI systems do not amplify cyber threats. Readers in the AI security community are discussing what the benchmarks mean for safety testing.

  22. 22
    Federal AI portal scores 9 out of 12 on Bitcoin policy quiz●🤖 America .gov, el portal federal con # IA de # Google y # SpaceXAI , acertó 9 de 12 preguntas sobre política de BitcoinMmastodonBusinessCrypto14 d ago

    Bitcoin.com News evaluated America.gov, a federal portal using AI from Google and SpaceX, on Bitcoin policy questions. The portal answered 9 of 12 questions correctly, performing less accurately than ChatGPT and Claude in the same test. The comparison highlights how leading commercial AI chatbots stack up on cryptocurrency policy knowledge, with a government-backed tool trailing behind.

  23. 23
    AI Ready Roanoke Assesses New Technology Use in the Roanoke Valley▼AI Ready Roanoke evaluates use for new technology in Roanoke Valley✉newsTechnology2 d ago

    AI Ready Roanoke is evaluating how new technology, particularly artificial intelligence, could be used across the Roanoke Valley in Virginia. The initiative, covered by the Roanoke Times, is examining practical applications for local businesses, institutions, and residents as AI adoption spreads. Details on specific projects, partners, and timelines have not yet been reported.

  24. 24
    Innodata Opens Robot Data Lab, Investors Watch Closely▼How Investors May Respond To Innodata (INOD) Opening Robot Data Lab✉newsTechnologyRobotics1 d ago

    Data engineering company Innodata has opened a new lab focused on robot training data, drawing attention from investors assessing what the move means for the firm's growth strategy. The company, known for AI data services, is positioning itself in the emerging market for data used to train robotics systems. Market watchers are weighing whether the lab can translate into revenue and how it might affect the stock's performance going forward.

  25. 25
    White House forms task force on AI risks▼White House forms AI task force to assess technology risks - WSJ✉newsTechnology2 d ago

    The White House has created a task force to assess the risks of artificial intelligence, according to the Wall Street Journal. The group will evaluate potential dangers from rapidly advancing AI technology, part of a broader effort by the US administration to develop policy on AI safety and regulation.

  26. 26
    Boston VA workshop explores AI for veteran care▼Boston VA workshop tests AI’s potential for Veteran care✉newsTechnologyAI12 h ago

    The Department of Veterans Affairs held a workshop in Boston examining how artificial intelligence could improve care for veterans. The event tested potential applications of AI within VA healthcare services. It reflects the agency's broader effort to evaluate whether AI tools can support clinical decisions, administration, and patient outcomes for former service members.

  27. 27

    Call centre and service workers are increasingly evaluated by AI systems that score their calls automatically, while the criteria behind those scores remain hidden from employees. Commentators argue this opaque algorithmic management gives employers sweeping power over performance reviews and discipline without workers being able to challenge or even understand how they are judged, fuelling demands for transparency rules and stronger workplace AI regulation.

  28. 28
    New Data Shows How Nurses Are Using AI▼How Nurses Are Using AI: The Data at a Glance✉newsTechnologyAI13 h ago

    The American Hospital Association has published an overview of how nurses are using artificial intelligence in their work, presenting the available data in chart form. The roundup highlights adoption patterns and practical applications of AI tools in nursing practice, a topic of growing interest as hospitals weigh efficiency gains against concerns over patient safety and job roles.

  29. 29
    As A.I. Agents Begin Shopping, Brands Rethink Their Sales Pitch▼As A.I. Agents Begin Shopping, Brands Are Changing Their Sales Pitch✉newsTechnologyAI2 d ago

    Brands are adjusting how they market products as artificial intelligence agents increasingly shop on behalf of consumers. Instead of pitching to human shoppers with emotional advertising, companies are optimizing product data and descriptions so AI assistants can find, evaluate and recommend their goods, a shift with significant implications for e-commerce and digital marketing.

  30. 30
    Darwin-180B-RSI tops three AI benchmarks●Darwin-180B-RSI suma el puesto # 1 en MDPBench, ExtractBench e IFStruct y alcanza 10 primeros puestos oficiales. Y ZTC jMmastodonTechnologySoftware31 d ago

    The open-source Spanish-language model Darwin-180B-RSI has reached first place on the MDPBench, ExtractBench and IFStruct leaderboards, bringing its total to ten official top rankings. Its ZTC component is noted for evaluating output without generating text itself. Commenters in the AI and open-source community are highlighting the model's benchmark sweep and its relevance for Spanish-language software development.

  31. 31

    A debate is under way in the scientific community over whether artificial intelligence should be used in the peer review of research papers. Supporters see potential to speed up reviews and ease reviewer shortages, while critics worry about bias, confidentiality of unpublished manuscripts, and the risk of undermining trust in scholarly publishing.

  32. 32
    Researchers propose fixing GRPO's credit assignment problem●Fixing GRPO's credit assignment problem without evaluating every stepYhnBusinessEconomy233 d ago

    A new paper on arXiv proposes a way to fix the credit assignment problem in GRPO, a reinforcement learning method used for training language models, without evaluating every step of a response. GRPO currently assigns the same reward to all tokens in a completion, making it hard to identify which parts of an output earned the reward. The proposed approach aims to improve this at lower computational cost.

  33. 33
    OpenAI partners with Ironclad on AI contracting agents●🤖 Advancing computer use with Ironclad Learn how OpenAI and Ironclad are training and evaluating AI agents on complex coMmastodonTechnologyAI011 h ago

    OpenAI announced a partnership with contract management company Ironclad to train and evaluate AI agents on complex contracting workflows. The collaboration aims to advance computer-use AI for professional work, using Ironclad's real-world legal and contracting tasks as a testing ground. It signals a push toward deploying AI agents in specialized business environments beyond general browsing.

  34. 34
    Your company's AI needs a scoreboard●Your company's AI needs a scoreboard https://www.fastcompany.com/91614445/your-companys-ai-needs-a-scoreboard # AI # BusMmastodonBusiness31 d ago

    A Fast Company article argues that companies adopting AI need clear measurement systems—scoreboards—to track whether their AI investments are actually delivering results. The piece taps into a wider business debate about how firms should evaluate AI tools beyond hype, using metrics and benchmarks to judge performance and justify spending.

  35. 35
    Venture Investors Turn to Judgment Where AI Data Falls Short▼Once AI Reads the Deck, Venture Investors Test What the Data Cannot Explain✉newsBusinessStartups2 d ago

    Venture capital investors are reportedly exploring what AI cannot explain when evaluating startups, even as AI tools increasingly analyse pitch decks and company data. The discussion centres on whether human judgment still matters in investment decisions once machines can process the numbers. Details of specific firms or funds involved remain limited, and reaction from the wider venture community is not yet clear.

  36. 36
    Companies begin allowing AI agents in coding interviews●More companies now let candidates use an AI agent during coding interviews. That sounds like it makes... # ai # career #MmastodonTechnologySoftware322 h ago

    A growing number of employers are permitting job candidates to use AI coding agents during technical interviews. The shift changes what is being assessed: rather than writing code unaided, applicants must show they can direct and work with AI tools effectively. Commenters note this mirrors real-world engineering work, where AI assistance is now standard, and raises questions about how to evaluate core programming ability.

  37. 37

    A widely shared essay argues that in the era of AI agents, the harness—the scaffolding of tools, prompts, evaluation and workflow code wrapped around a model—is where a company's actual value sits, not the underlying model itself. As models become commoditised and interchangeable, the author contends the harness is the durable product, and effectively the company's true identity.

  38. 38
    AI models ranked by normative score across twelve paradigms●Models ranked by normative score across twelve paradigms, once neutral and once human-primed https://cognit.rajtilak.tecMmastodonTechnologyAI32 d ago

    A new write-up ranks AI language models by their normative score, tested across twelve different paradigms and repeated twice — once under neutral conditions and once with human priming. The comparison aims to show how priming changes model behavior and judgments, offering a benchmark-style look at consistency across evaluation setups.

  39. 39
    New benchmark tests AI agents on messy company knowledge●Benchmarking retrieval for agents on messy real-world company knowledgeYhn224 d ago

    Kapa.ai has published a benchmark for evaluating how well retrieval systems let AI agents work with messy, real-world company knowledge bases. The release is drawing attention among developers and AI practitioners, who are discussing how enterprise search and agent performance should be measured outside clean, curated datasets, where documentation is inconsistent, outdated or scattered across tools.

  40. 40
    Law firm publishes guide to AI insurance coverage●Keep Your AI On The Ball: A Policyholder’s Guide To Artificial Intelligence Insurance Coverage✉newsTechnologyAI1 d ago

    New York law firm Pryor Cashman has published a policyholder's guide to insurance coverage for artificial intelligence, explaining how existing policies may or may not respond to AI-related risks and losses. The guidance walks businesses through evaluating whether their current coverage extends to AI tools, what exclusions to watch for, and how to negotiate better protection. It reflects growing demand from companies using AI who are unsure whether their insurers will cover AI-related liability.

Repos