MikeTrendsTrends right now

search

AI evaluators

Trends

  1. 1
    78% of U.S. homebuyers now use AI tools when buying a home●78% of U.S. homebuyers now use AI to help buy a home https://www.fastcompany.com/91618486/homebuying-real-estate-agent-sMmastodonBusinessReal Estate332 min ago

    A new survey indicates that 78% of U.S. homebuyers now use artificial intelligence to help with the home-buying process, from researching listings to evaluating offers. The figure highlights how quickly AI tools have entered real estate, raising questions about the evolving role of traditional agents and how buyers make one of life's biggest financial decisions.

  2. 2
    Graphene toolkit brings data analysis to coding agents▼Show HN: Graphene – Data analysis toolkit for your coding agentYhnEnvironmentOceans329 min ago

    A new open-source project called Graphene has been launched, offering a data analysis toolkit designed to work with AI coding agents. The toolkit was shared as a launch announcement on Hacker News under its Show HN format, drawing early engagement from the developer community. Details beyond the GitHub repository are limited so far, with users likely evaluating its features and usefulness.

  3. 3

    A developer has published a write-up after spending a month coding with GLM 5.3 Flash, sharing hands-on experience with the AI model for programming tasks. The post has drawn attention on Hacker News, where readers are weighing in on how the model performs in real-world development work and how it compares with rival coding assistants.

  4. 4
    Anthropic claims Chinese AI model has Mythos-class hacking abilities●Anthropic claims popular Chinese AI model has Mythos-class hacking abilitiesYhnBusinessRetail631 min ago

    Anthropic says a widely used Chinese open-weight AI model demonstrates Mythos-class hacking capabilities, according to a new frontier red-teaming report. The company's evaluation found that the model's safeguards are weak, raising concerns that powerful offensive cyber abilities could be freely distributed and abused. The report has drawn attention to broader questions about how open-weight AI models should be tested and restricted before release.

  5. 5
    OpenAI publishes mathematical manuscripts and proof artifacts●Mathematical manuscripts and supporting proof artifacts produced by OpenAIYhnCultureArt421 h ago

    OpenAI has released a repository containing mathematical manuscripts along with supporting proof artifacts produced by its models. The collection gives researchers and mathematicians material to examine how AI systems formulate and verify formal proofs. The release is drawing attention from the AI and mathematics communities, who are assessing the quality and significance of the generated work.

  6. 6

    Companies are pouring money into artificial intelligence, but measuring the return on that investment is proving elusive across the industry. Many firms struggle to link AI spending to concrete business outcomes, leaving executives uncertain whether the technology is delivering value. The issue is prompting debate over how AI projects should be evaluated and whether current spending levels are justified.

  7. 7
    AI evaluators face pressure to keep advanced systems safe▼‘Evaluators’ are supposed to keep AI from killing us all. No pressure✉newsTechnology11 h ago

    A new report highlights the role of AI evaluators, the specialists tasked with assessing whether advanced AI systems are safe before deployment. Their work is framed as a critical safeguard against catastrophic AI risks, yet it is largely voluntary and under-resourced. Commenters are debating whether this handful of assessors can realistically hold back powerful technology developed by some of the world's richest companies.

  8. 8
    OpenAI releases 722 math manuscripts●OpenAI releases 722 math manuscripts https://github.com/openai/math/blob/main/CONTENTS.md # HackerNews # Tech # AIMmastodonTechnology41 h ago

    OpenAI has made 722 math manuscripts publicly available through a repository on GitHub, opening up a collection of mathematical texts to researchers and the public. The release is being discussed among developers and AI enthusiasts, who see it as a potentially useful resource for training and evaluating mathematical reasoning in AI models.

  9. 9
    Study probes whether AI models judge code morally●Ask a model if code is malicious and it reaches for its morals https://www.manifold.security/blog/do-models-consider-morMmastodonTechnology48 h ago

    Security firm Manifold Security published research asking whether AI models factor morality into their judgments about malicious code. The finding: when asked to assess whether code is malware, language models appear to bring moral reasoning into their analysis rather than relying purely on technical criteria. The report is circulating among developers and security researchers interested in how AI tools evaluate potentially harmful software.

  10. 10
    UN University Seeks to Build Independent AI Evaluators Community▼Building a Community of Independent AI Evaluators✉newsWorldUnited Nations52 min ago

    United Nations University is working to create a community of independent AI evaluators, aiming to strengthen global oversight of artificial intelligence systems. The initiative points to growing concern that AI development is outpacing reliable, impartial assessment of its risks and impacts. The UN University's role suggests an effort to coordinate expertise across countries and disciplines for credible, independent evaluation capacity.

  11. 11
    Banana Pi launches RK3576 computing module for Edge AI▼Banana Pi BPI-CM5 Pro is a computing module powered by the Rockchip RK3576, 6 TOPS computing power NPU,8-32G RAM and 8-1MmastodonTechnologyRobotics212 h ago

    Banana Pi has introduced the BPI-CM5 Pro, a compute module built on Rockchip's RK3576 chip with a 6 TOPS NPU, 8-32GB of RAM and 8-128GB of eMMC storage. The maker positions it as a board for Edge AI projects and an alternative to the Raspberry Pi Compute Module 4, drawing interest from hobbyists and robotics developers.

  12. 12
    OpenAI floods mathematics with hundreds of new results●OpenAI unleashes hundreds more math results upon a field already in shock https://www.scientificamerican.com/article/opeMmastodonScience43 h ago

    OpenAI has released hundreds of new mathematical results, adding to a wave of AI-generated output that has already unsettled researchers in the field. The announcement, covered by Scientific American, is drawing attention because of the sheer volume of results and ongoing debate about how AI-generated mathematics should be evaluated and verified.

  13. 13
    OpenAI Publishes Findings on 377 Math Problems▼OpenAI Releases Findings on 377 Math Problems, Further Roiling Field✉newsTechnologyAI3 h ago

    OpenAI has released findings based on 377 math problems, an announcement reported by The New York Times as further roiling the artificial intelligence field. The release has drawn attention within the AI research community, where debate continues over model capabilities, evaluation methods and the pace of progress on mathematical reasoning benchmarks.

  14. 14
    Foundation AI Highlights VLoc Bench and Cyber-Capability Safety▼Foundation AI in September: VLoc Bench and Cyber-Capability Safety✉newsTechnologyAI1 h ago

    Cisco's Foundation AI team published its September update, spotlighting the VLoc Bench and new work on cyber-capability safety. The release covers how the team benchmarks and evaluates AI models for security-relevant capabilities, part of a broader push to make sure advanced AI systems do not amplify cyber threats. Readers in the AI security community are discussing what the benchmarks mean for safety testing.

  15. 15
    TechCrunch Disrupt 2026 reveals Startup Battlefield 200 judges●Meet the Startup Battlefield 200 judges who'll decide the winner at TechCrunch Disrupt 2026 https://techcrunch.com/2026/MmastodonBusinessStartups31 d ago

    TechCrunch has announced the panel of judges for the Startup Battlefield 200 at TechCrunch Disrupt 2026. The judges will evaluate the competing startups and decide which company takes the top prize at the conference. The lineup leans heavily toward technology and AI-focused investors and operators, drawing attention from the startup community ahead of the event.

  16. 16
    Developer tests AI help to speed up video compositing●So... asked the AI to evaluate performance on r/t video compositing trying to get processing per frame under 50ms next-uMmastodonCultureGaming11 d ago

    A developer is working on real-time video compositing and trying to bring processing time per frame under 50 milliseconds, a threshold needed for smooth real-time performance. They asked an AI assistant to evaluate performance and are now working on a strip-parallel compositor that splits the frame across threads, sacrificing four threads in the process. The work is being done while also livestreaming and gaming, and the developer has shared the experiment publicly.

  17. 17
    Boston VA workshop explores AI for veteran care▼Boston VA workshop tests AI’s potential for Veteran care✉newsTechnologyAI14 h ago

    The Department of Veterans Affairs held a workshop in Boston examining how artificial intelligence could improve care for veterans. The event tested potential applications of AI within VA healthcare services. It reflects the agency's broader effort to evaluate whether AI tools can support clinical decisions, administration, and patient outcomes for former service members.

  18. 18
    New Data Shows How Nurses Are Using AI▼How Nurses Are Using AI: The Data at a Glance✉newsTechnologyAI15 h ago

    The American Hospital Association has published an overview of how nurses are using artificial intelligence in their work, presenting the available data in chart form. The roundup highlights adoption patterns and practical applications of AI tools in nursing practice, a topic of growing interest as hospitals weigh efficiency gains against concerns over patient safety and job roles.

  19. 19
    Frontier AI Models Show Uneven Skill from Web Browsing to Robotics●Astra, Opus 5.5 Demonstrate Jagged Performance on Web to Robotics TasksYhnTechnologyRobotics818 h ago

    A new evaluation from Fig reports that frontier AI models, including Astra and Opus 5.5, perform unevenly — 'jagged' — across agentic tasks, doing well on some web browsing challenges while struggling on robotics tasks. The findings suggest that even top models lack consistent reliability, with strong results in one domain not predicting competence in another.

  20. 20
    OpenAI partners with Ironclad on AI contracting agents●🤖 Advancing computer use with Ironclad Learn how OpenAI and Ironclad are training and evaluating AI agents on complex coMmastodonTechnologyAI012 h ago

    OpenAI announced a partnership with contract management company Ironclad to train and evaluate AI agents on complex contracting workflows. The collaboration aims to advance computer-use AI for professional work, using Ironclad's real-world legal and contracting tasks as a testing ground. It signals a push toward deploying AI agents in specialized business environments beyond general browsing.

  21. 21

    Call centre and service workers are increasingly evaluated by AI systems that score their calls automatically, while the criteria behind those scores remain hidden from employees. Commentators argue this opaque algorithmic management gives employers sweeping power over performance reviews and discipline without workers being able to challenge or even understand how they are judged, fuelling demands for transparency rules and stronger workplace AI regulation.

  22. 22
    New Proxy Lets AI Models Train Inside Real Coding Harnesses●New Proxy Trains AI Models Inside Real Coding Harnesses Without Changes𝕏xSE1381 d ago

    A new open-source tool called Proxy allows AI coding models to be trained and evaluated inside real coding harnesses without any modifications to the existing setup. The project aims to bridge the gap between benchmark testing and practical use, letting developers plug models directly into their workflows. Developer communities are discussing its potential to speed up model iteration and testing.

  23. 23
    Developer pits AI coding agents against each other in races●Every few days someone tells me which coding agent is "obviously the best". They always have a... # ai # opensource # shMmastodonTechnologySoftware41 d ago

    A developer has built races between AI coding agents to test which one performs best, challenging the frequent claims that any single agent is "obviously the best". The finding, shared widely in software and open-source communities, was that the fastest agent was not the winner, prompting discussion about how AI coding tools should actually be evaluated.

  24. 24
    Darwin-180B-RSI tops three AI benchmarks●Darwin-180B-RSI suma el puesto # 1 en MDPBench, ExtractBench e IFStruct y alcanza 10 primeros puestos oficiales. Y ZTC jMmastodonTechnologySoftware31 d ago

    The open-source Spanish-language model Darwin-180B-RSI has reached first place on the MDPBench, ExtractBench and IFStruct leaderboards, bringing its total to ten official top rankings. Its ZTC component is noted for evaluating output without generating text itself. Commenters in the AI and open-source community are highlighting the model's benchmark sweep and its relevance for Spanish-language software development.

  25. 25
    Companies begin allowing AI agents in coding interviews●More companies now let candidates use an AI agent during coding interviews. That sounds like it makes... # ai # career #MmastodonTechnologySoftware323 h ago

    A growing number of employers are permitting job candidates to use AI coding agents during technical interviews. The shift changes what is being assessed: rather than writing code unaided, applicants must show they can direct and work with AI tools effectively. Commenters note this mirrors real-world engineering work, where AI assistance is now standard, and raises questions about how to evaluate core programming ability.

  26. 26
    Your company's AI needs a scoreboard●Your company's AI needs a scoreboard https://www.fastcompany.com/91614445/your-companys-ai-needs-a-scoreboard # AI # BusMmastodonBusiness31 d ago

    A Fast Company article argues that companies adopting AI need clear measurement systems—scoreboards—to track whether their AI investments are actually delivering results. The piece taps into a wider business debate about how firms should evaluate AI tools beyond hype, using metrics and benchmarks to judge performance and justify spending.

  27. 27
    Innodata Opens Robot Data Lab, Investors Watch Closely▼How Investors May Respond To Innodata (INOD) Opening Robot Data Lab✉newsTechnologyRobotics2 d ago

    Data engineering company Innodata has opened a new lab focused on robot training data, drawing attention from investors assessing what the move means for the firm's growth strategy. The company, known for AI data services, is positioning itself in the emerging market for data used to train robotics systems. Market watchers are weighing whether the lab can translate into revenue and how it might affect the stock's performance going forward.

  28. 28
    AI Ready Roanoke Assesses New Technology Use in the Roanoke Valley▼AI Ready Roanoke evaluates use for new technology in Roanoke Valley✉newsTechnology2 d ago

    AI Ready Roanoke is evaluating how new technology, particularly artificial intelligence, could be used across the Roanoke Valley in Virginia. The initiative, covered by the Roanoke Times, is examining practical applications for local businesses, institutions, and residents as AI adoption spreads. Details on specific projects, partners, and timelines have not yet been reported.

  29. 29
    White House forms task force on AI risks▼White House forms AI task force to assess technology risks - WSJ✉newsTechnology2 d ago

    The White House has created a task force to assess the risks of artificial intelligence, according to the Wall Street Journal. The group will evaluate potential dangers from rapidly advancing AI technology, part of a broader effort by the US administration to develop policy on AI safety and regulation.

  30. 30
    Venture Investors Turn to Judgment Where AI Data Falls Short▼Once AI Reads the Deck, Venture Investors Test What the Data Cannot Explain✉newsBusinessStartups2 d ago

    Venture capital investors are reportedly exploring what AI cannot explain when evaluating startups, even as AI tools increasingly analyse pitch decks and company data. The discussion centres on whether human judgment still matters in investment decisions once machines can process the numbers. Details of specific firms or funds involved remain limited, and reaction from the wider venture community is not yet clear.

  31. 31

    A debate is under way in the scientific community over whether artificial intelligence should be used in the peer review of research papers. Supporters see potential to speed up reviews and ease reviewer shortages, while critics worry about bias, confidentiality of unpublished manuscripts, and the risk of undermining trust in scholarly publishing.

  32. 32
    Law firm publishes guide to AI insurance coverage●Keep Your AI On The Ball: A Policyholder’s Guide To Artificial Intelligence Insurance Coverage✉newsTechnologyAI1 d ago

    New York law firm Pryor Cashman has published a policyholder's guide to insurance coverage for artificial intelligence, explaining how existing policies may or may not respond to AI-related risks and losses. The guidance walks businesses through evaluating whether their current coverage extends to AI tools, what exclusions to watch for, and how to negotiate better protection. It reflects growing demand from companies using AI who are unsure whether their insurers will cover AI-related liability.

  33. 33
    AI safety researchers push reliability over raw capability●Intelligence without reliability is not enough. As AI systems become more autonomous, we need better approaches to evaluMmastodonTechnologyAI12 d ago

    Researchers at Antralabs are arguing that raw intelligence in AI systems is not enough as models become more autonomous. The company says the field needs stronger approaches to evaluation, alignment, reliability and system-level safety, and it says it is working on foundations for safer intelligent systems. The argument feeds a wider debate about whether current AI evaluation methods can keep pace with increasingly autonomous systems.

  34. 34
    New VA-Bench Shows AI Models Struggle With Robot Tasks●Dalian University of Technology's VA-Bench: Top Multimodal Models Finish Only Half of Robot Tasks✉newsTechnologyRobotics1 d ago

    Dalian University of Technology has introduced VA-Bench, a benchmark for evaluating multimodal AI models on robotics tasks. Results show that even the top-performing models complete only about half of the tested tasks, highlighting a significant gap between current multimodal capabilities and practical robot control. The benchmark is drawing attention as a measure of how far AI still is from reliable real-world robotics.

  35. 35
    AI models ranked by normative score across twelve paradigms●Models ranked by normative score across twelve paradigms, once neutral and once human-primed https://cognit.rajtilak.tecMmastodonTechnologyAI32 d ago

    A new write-up ranks AI language models by their normative score, tested across twelve different paradigms and repeated twice — once under neutral conditions and once with human priming. The comparison aims to show how priming changes model behavior and judgments, offering a benchmark-style look at consistency across evaluation setups.

  36. 36
    As A.I. Agents Begin Shopping, Brands Rethink Their Sales Pitch▼As A.I. Agents Begin Shopping, Brands Are Changing Their Sales Pitch✉newsTechnologyAI2 d ago

    Brands are adjusting how they market products as artificial intelligence agents increasingly shop on behalf of consumers. Instead of pitching to human shoppers with emotional advertising, companies are optimizing product data and descriptions so AI assistants can find, evaluate and recommend their goods, a shift with significant implications for e-commerce and digital marketing.

  37. 37
    NFL Next Gen Stats Turns AI On The Running Game●How NFL Next Gen Stats And AI Are Learning The Running Game✉newsTechnologyAI2 d ago

    The NFL's Next Gen Stats platform is applying artificial intelligence to better analyze the running game, using player tracking data to break down runs in new ways. The work shows how machine learning is moving beyond pass-focused analytics in football, and it is drawing attention as teams and fans look for smarter ways to evaluate rushing performance across the league.

  38. 38
    White House Moves to Assess Growing AI Risks●Watch White House Moves to Assess Growing AI Risks✉newsTechnologyAI2 d ago

    The White House is taking steps to assess the growing risks posed by artificial intelligence, according to Bloomberg. The move signals increasing concern at the highest level of US government about the safety and societal impact of rapidly advancing AI systems. Details of the assessment and any resulting policy measures have not yet been outlined.

Repos