MikeTrendsTrends right now

search

AI benchmarks

Trends

  1. 1
    Running Qwen 3.8 Flash Next 125B on an RTX 4090 at 100 tokens per second●Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/sYhnSportFootball9351 h ago

    A new open-source project called Strata claims it can run the Qwen 3.8 Flash Next 125-billion-parameter model on a single consumer RTX 4090 graphics card at roughly 100 tokens per second. If the benchmarks hold up, it would let hobbyists and small teams run a frontier-scale language model locally without expensive data-center hardware. The project has drawn attention on developer forums, where users are questioning the memory techniques behind the speed claims and awaiting independent replication.

  2. 2
    AI agents find two room-temperature magnetic semiconductor candidates●Opus 5.5 agents discover two room-temperature magnetic semiconductor candidatesYhnLifeHome & Garden49347 min ago

    Anthropic's Opus 5.5 model, running as autonomous research agents, has reportedly identified two candidate materials for room-temperature magnetic semiconductors, a long-sought class of materials that could enable spintronic devices without cryogenic cooling. The work, described by AI benchmark firm Vals AI, is drawing attention for showing AI agents contributing genuine materials-science discoveries rather than incremental analysis.

  3. 3
    AI decision models tested by playing Pac-Man▼Show HN: Jevman – AI decision models play Pac-ManYhnLifeFood7620 min ago

    A new benchmark called Jevman has been launched that evaluates AI decision-making models by having them play Pac-Man. The project, presented by Opper AI, uses the classic arcade game as a testing ground for how well language models plan, weigh risks and make sequential decisions. It is drawing attention from developers and AI researchers discussing whether game-based benchmarks meaningfully measure model reasoning ability.

  4. 4

    AIMS Lab at Stanford has made available a textbook on AI Measurement Science, a field focused on rigorously evaluating and measuring the performance of artificial intelligence systems. The free online textbook is drawing attention among technologists discussing how AI capabilities should be quantified, benchmarked and validated as systems grow more complex and their evaluation methods come under scrutiny.

  5. 5
    Recent AI models struggle to match a human algorithmic breakthrough●Recent AI models struggled to match a human algorithmic innovationYhn3211 min ago

    New research from Epoch AI finds that frontier AI models fell short of reproducing a human algorithmic innovation, highlighting gaps between benchmark performance and genuine inventive capability. The study examines how well current models can independently rediscover improvements humans devised, and commentators are debating what the results say about AI's real research abilities.

  6. 6
    Robotera's VPP2 model tops RoboDojo robotics benchmark▼Robotera's VPP2 World Action Model Tops RoboDojo, Scores 58.5% Zero-Shot on Real ALOHA Arms✉newsTechnologySoftware10 min ago

    Robotera says its VPP2 world action model has taken the top spot on the RoboDojo leaderboard, scoring 58.5% zero-shot on real ALOHA robot arms. The result, reported by Pandaily, suggests the model can transfer directly to physical hardware without task-specific fine-tuning, a key test for general-purpose robot control.

  7. 7
    AP Stylebook's expanded AI guidance draws journalist attention●"The AP Stylebook first added its dedicated artificial intelligence chapter in 2023, then broadly expanded and updated iMmastodonLifeFashion449 min ago

    The Associated Press Stylebook added a dedicated artificial intelligence chapter in 2023 and expanded it earlier this year. The guidance goes beyond terminology, urging journalists to critically evaluate AI tools and their limitations. Commenters are highlighting how the AP's advice is ethically grounded, framing it as a benchmark for responsible AI use in newsrooms.

Repos