search
AI benchmarks
Trends
- 1Running Qwen 3.8 Flash Next 125B on an RTX 4090 at 100 tokens per second●Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
A new open-source project called Strata claims it can run the Qwen 3.8 Flash Next 125-billion-parameter model on a single consumer RTX 4090 graphics card at roughly 100 tokens per second. If the benchmarks hold up, it would let hobbyists and small teams run a frontier-scale language model locally without expensive data-center hardware. The project has drawn attention on developer forums, where users are questioning the memory techniques behind the speed claims and awaiting independent replication.
- 2AI agents find two room-temperature magnetic semiconductor candidates●Opus 5.5 agents discover two room-temperature magnetic semiconductor candidates
Anthropic's Opus 5.5 model, running as autonomous research agents, has reportedly identified two candidate materials for room-temperature magnetic semiconductors, a long-sought class of materials that could enable spintronic devices without cryogenic cooling. The work, described by AI benchmark firm Vals AI, is drawing attention for showing AI agents contributing genuine materials-science discoveries rather than incremental analysis.
- 3
A new benchmark called Jevman has been launched that evaluates AI decision-making models by having them play Pac-Man. The project, presented by Opper AI, uses the classic arcade game as a testing ground for how well language models plan, weigh risks and make sequential decisions. It is drawing attention from developers and AI researchers discussing whether game-based benchmarks meaningfully measure model reasoning ability.
- 4
AIMS Lab at Stanford has made available a textbook on AI Measurement Science, a field focused on rigorously evaluating and measuring the performance of artificial intelligence systems. The free online textbook is drawing attention among technologists discussing how AI capabilities should be quantified, benchmarked and validated as systems grow more complex and their evaluation methods come under scrutiny.
- 5Recent AI models struggle to match a human algorithmic breakthrough●Recent AI models struggled to match a human algorithmic innovation
New research from Epoch AI finds that frontier AI models fell short of reproducing a human algorithmic innovation, highlighting gaps between benchmark performance and genuine inventive capability. The study examines how well current models can independently rediscover improvements humans devised, and commentators are debating what the results say about AI's real research abilities.
- 6Robotera's VPP2 model tops RoboDojo robotics benchmark▼Robotera's VPP2 World Action Model Tops RoboDojo, Scores 58.5% Zero-Shot on Real ALOHA Arms
Robotera says its VPP2 world action model has taken the top spot on the RoboDojo leaderboard, scoring 58.5% zero-shot on real ALOHA robot arms. The result, reported by Pandaily, suggests the model can transfer directly to physical hardware without task-specific fine-tuning, a key test for general-purpose robot control.
- 7AP Stylebook's expanded AI guidance draws journalist attention●"The AP Stylebook first added its dedicated artificial intelligence chapter in 2023, then broadly expanded and updated i
The Associated Press Stylebook added a dedicated artificial intelligence chapter in 2023 and expanded it earlier this year. The guidance goes beyond terminology, urging journalists to critically evaluate AI tools and their limitations. Commenters are highlighting how the AP's advice is ethically grounded, framing it as a benchmark for responsible AI use in newsrooms.
Repos
- DietrichGebert/ponytail Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
- QingYunA/answer-me-with-html Answer me with HTML — an agent skill that answers hard questions with a one-page HTML you can actually read. 让 AI Agent
- trycua/cua Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data gene
- elstongun/leviathan Deep memory for agents over large datasets. Leviathan is a single static binary that turns your records (JSONL, JSON, CS
- archestra-ai/OpenAPPA Deterministic guardrails that don't break agents
- JoasASantos/Offensive-Security-AI-Models Uncensored AI models or those fine-tuned for cybersecurity tasks.
- edgedelta/project-arena A vendor-neutral Kubernetes incident benchmark for testing AI investigation tools across detection, diagnosis, and mitig
- ninjahawk/livenerf Benchmark for tracking model capability after release.
- Rizzo-AI-Academy/rizzo-flow The open, local take on Jev: typed decisions from an LLM, without generating a single token