MikeTrendsTrends right now

search

AI evaluators

Trends

  1. 1

    A developer has published a write-up after spending a month coding with GLM 5.3 Flash, sharing hands-on experience with the AI model for programming tasks. The post has drawn attention on Hacker News, where readers are weighing in on how the model performs in real-world development work and how it compares with rival coding assistants.

  2. 2
    Anthropic claims Chinese AI model has Mythos-class hacking abilitiesโ—Anthropic claims popular Chinese AI model has Mythos-class hacking abilitiesYhnBusinessRetail61 h ago

    Anthropic says a widely used Chinese open-weight AI model demonstrates Mythos-class hacking capabilities, according to a new frontier red-teaming report. The company's evaluation found that the model's safeguards are weak, raising concerns that powerful offensive cyber abilities could be freely distributed and abused. The report has drawn attention to broader questions about how open-weight AI models should be tested and restricted before release.

  3. 3
    OpenAI publishes mathematical manuscripts and proof artifactsโ—Mathematical manuscripts and supporting proof artifacts produced by OpenAIYhnCultureArt421 h ago

    OpenAI has released a repository containing mathematical manuscripts along with supporting proof artifacts produced by its models. The collection gives researchers and mathematicians material to examine how AI systems formulate and verify formal proofs. The release is drawing attention from the AI and mathematics communities, who are assessing the quality and significance of the generated work.

  4. 4

    Companies are pouring money into artificial intelligence, but measuring the return on that investment is proving elusive across the industry. Many firms struggle to link AI spending to concrete business outcomes, leaving executives uncertain whether the technology is delivering value. The issue is prompting debate over how AI projects should be evaluated and whether current spending levels are justified.

  5. 5
    UN University Seeks to Build Independent AI Evaluators Communityโ–ผBuilding a Community of Independent AI Evaluatorsโœ‰newsWorldUnited Nations1 h ago

    United Nations University is working to create a community of independent AI evaluators, aiming to strengthen global oversight of artificial intelligence systems. The initiative points to growing concern that AI development is outpacing reliable, impartial assessment of its risks and impacts. The UN University's role suggests an effort to coordinate expertise across countries and disciplines for credible, independent evaluation capacity.

  6. 6
    UT San Antonio launches certificate on critical AI use for educatorsโ—New UT San Antonio certificate helps educators approach AI with a critical eyeโœ‰newsLifeEducation20 min ago

    The University of Texas at San Antonio has introduced a new certificate program designed to help educators engage with artificial intelligence critically. The program aims to equip teachers with the skills to evaluate AI tools and their classroom implications rather than adopt them uncritically. It reflects a broader push in higher education to prepare teachers for the rapid spread of AI in schools.

  7. 7
    OpenAI releases 722 math manuscriptsโ—OpenAI releases 722 math manuscripts https://github.com/openai/math/blob/main/CONTENTS.md # HackerNews # Tech # AIMmastodonTechnology43 h ago

    OpenAI has made 722 math manuscripts publicly available through a repository on GitHub, opening up a collection of mathematical texts to researchers and the public. The release is being discussed among developers and AI enthusiasts, who see it as a potentially useful resource for training and evaluating mathematical reasoning in AI models.

  8. 8
    Foundation AI Highlights VLoc Bench and Cyber-Capability Safetyโ–ผFoundation AI in September: VLoc Bench and Cyber-Capability Safetyโœ‰newsTechnologyAI3 h ago

    Cisco's Foundation AI team published its September update, spotlighting the VLoc Bench and new work on cyber-capability safety. The release covers how the team benchmarks and evaluates AI models for security-relevant capabilities, part of a broader push to make sure advanced AI systems do not amplify cyber threats. Readers in the AI security community are discussing what the benchmarks mean for safety testing.

  9. 9
    OpenAI floods mathematics with hundreds of new resultsโ—OpenAI unleashes hundreds more math results upon a field already in shock https://www.scientificamerican.com/article/opeMmastodonScience45 h ago

    OpenAI has released hundreds of new mathematical results, adding to a wave of AI-generated output that has already unsettled researchers in the field. The announcement, covered by Scientific American, is drawing attention because of the sheer volume of results and ongoing debate about how AI-generated mathematics should be evaluated and verified.

  10. 10
    Study probes whether AI models judge code morallyโ—Ask a model if code is malicious and it reaches for its morals https://www.manifold.security/blog/do-models-consider-morMmastodonTechnology410 h ago

    Security firm Manifold Security published research asking whether AI models factor morality into their judgments about malicious code. The finding: when asked to assess whether code is malware, language models appear to bring moral reasoning into their analysis rather than relying purely on technical criteria. The report is circulating among developers and security researchers interested in how AI tools evaluate potentially harmful software.

Repos