MikeTrendsTrends right now

search

AI benchmarks

Trends

  1. 1
    Qwen 125B model runs on a single RTX 4090 at high speed●Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/sYhnSportFootball9358 min ago

    A new open-source project called Strata claims to run the Qwen 3.8 Flash Next model, a 125-billion-parameter language model, on consumer hardware like an Nvidia RTX 4090, reportedly reaching around 100 tokens per second. The claim has drawn attention from developers discussing whether such performance on a single consumer GPU is realistic and what it could mean for local AI inference.

  2. 2
    OpenAI details its progress on AI in mathematics●Sharing AI progress in mathematicsYhnTechnologyAI1.3K28 min ago

    OpenAI has published a new post outlining how its AI systems are advancing on mathematical reasoning and problem-solving, and how it evaluates and shares that progress. The announcement is drawing attention on Hacker News, where it is among the most engaged items, with readers debating how significant the mathematical capabilities are and how honestly such progress is being reported.

  3. 3
    Jevman benchmark puts AI decision models in Pac-Man●Show HN: Jevman – AI decision models play Pac-ManYhnLifeFood785 min ago

    A new project called Jevman has been launched, using the classic arcade game Pac-Man as a benchmark environment for testing AI decision-making models. Announced on Hacker News, the tool from Opper AI invites developers and researchers to see how well language and decision models navigate the game's maze, weighing trade-offs and planning under uncertainty. Early commenters are discussing the approach and what game-based benchmarks reveal about reasoning ability.

  4. 4
    Anthropic cuts off its internal evaluations from the internet●Anthropic is cutting off its internal evaluations from the internet✉newsTechnologyInternet24 min ago

    Anthropic has restricted public internet access to its internal evaluations, according to a report by The Verge. The move means the AI company's benchmark testing and safety assessments will no longer be reachable online, drawing attention as AI firms face scrutiny over transparency of how they measure model safety and performance.

  5. 5
    AI models fall short of human algorithmic innovation●Recent AI models struggled to match a human algorithmic innovationYhn851 h ago

    A new study by Epoch AI finds that recent AI models struggled to match a human algorithmic innovation when tested on reproducing novel research ideas. The findings suggest frontier models still lag behind humans on tasks requiring genuine creative problem-solving, even as they excel at routine coding and reasoning benchmarks.

Repos