search
AI benchmarks
Trends
- 1Qwen 125B model runs on a single RTX 4090 at high speedâRun Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
A new open-source project called Strata claims to run the Qwen 3.8 Flash Next model, a 125-billion-parameter language model, on consumer hardware like an Nvidia RTX 4090, reportedly reaching around 100 tokens per second. The claim has drawn attention from developers discussing whether such performance on a single consumer GPU is realistic and what it could mean for local AI inference.
- 2
OpenAI has published a new post outlining how its AI systems are advancing on mathematical reasoning and problem-solving, and how it evaluates and shares that progress. The announcement is drawing attention on Hacker News, where it is among the most engaged items, with readers debating how significant the mathematical capabilities are and how honestly such progress is being reported.
- 3Jevman benchmark puts AI decision models in Pac-ManâShow HN: Jevman â AI decision models play Pac-Man
A new project called Jevman has been launched, using the classic arcade game Pac-Man as a benchmark environment for testing AI decision-making models. Announced on Hacker News, the tool from Opper AI invites developers and researchers to see how well language and decision models navigate the game's maze, weighing trade-offs and planning under uncertainty. Early commenters are discussing the approach and what game-based benchmarks reveal about reasoning ability.
- 4Anthropic cuts off its internal evaluations from the internetâAnthropic is cutting off its internal evaluations from the internet
Anthropic has restricted public internet access to its internal evaluations, according to a report by The Verge. The move means the AI company's benchmark testing and safety assessments will no longer be reachable online, drawing attention as AI firms face scrutiny over transparency of how they measure model safety and performance.
- 5AI models fall short of human algorithmic innovationâRecent AI models struggled to match a human algorithmic innovation
A new study by Epoch AI finds that recent AI models struggled to match a human algorithmic innovation when tested on reproducing novel research ideas. The findings suggest frontier models still lag behind humans on tasks requiring genuine creative problem-solving, even as they excel at routine coding and reasoning benchmarks.
Repos
- DietrichGebert/ponytail Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
- QingYunA/answer-me-with-html Answer me with HTML â an agent skill that answers hard questions with a one-page HTML you can actually read. 莊 AI Agent
- trycua/cua Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data gene
- elstongun/leviathan Deep memory for agents over large datasets. Leviathan is a single static binary that turns your records (JSONL, JSON, CS
- archestra-ai/OpenAPPA Deterministic guardrails that don't break agents
- edgedelta/project-arena A vendor-neutral Kubernetes incident benchmark for testing AI investigation tools across detection, diagnosis, and mitig
- JoasASantos/Offensive-Security-AI-Models Uncensored AI models or those fine-tuned for cybersecurity tasks.
- ninjahawk/livenerf Benchmark for tracking model capability after release.
- Rizzo-AI-Academy/rizzo-flow The open, local take on Jev: typed decisions from an LLM, without generating a single token