search
coding benchmark
Trends
- 1
Cursor, the AI-powered code editor, has added GLM 5.3 models, which are reported to top open-weight benchmarks. Developers following the AI coding tools space are discussing what the new model option means for performance and competition among coding assistants.
- 2Quantized 27B LLM claimed to match frontier models on one coding benchmarkโA 27B Quantized LLM Is Said To Match Frontier AI Models In Just One Task From A Coding Benchmark, Making It A More Believable Claim
A 27-billion-parameter quantized large language model is reported to match frontier AI models on a single task from a coding benchmark. Observers note that a narrow claim like this is more believable than sweeping performance comparisons, since smaller quantized models can reach parity in isolated tasks while still trailing frontier systems overall across coding and reasoning evaluations.
- 3DoGBench launches as first docs generation benchmark, AI falls shortโDoGBench: The first user-facing docs generation benchmark. No model scores >50%
DoGBench has been introduced as the first benchmark aimed at evaluating how well AI models generate user-facing documentation. Early results show that no model scores above 50%, a surprisingly low ceiling that is drawing attention. Developers on Hacker News are discussing what the weak performance says about the gap between coding assistants and genuinely usable documentation output.
- 4Open-Source Model Routing Claims Astra-Level Coding Agent PerformanceโShow HN: Open-source model routing for coding agents at Astra-level performance https://news.ycombinator.com/item?id=499
A developer has shared an open-source project on Hacker News that provides model routing for coding agents, claiming it reaches Astra-level performance. The tool routes requests between AI models to balance quality and cost for coding tasks. It is being showcased to the developer community, where feedback on the benchmark claims is likely to follow.
Repos
- ninjahawk/livenerf Benchmark for tracking model capability after release.
- DietrichGebert/ponytail Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
- archestra-ai/OpenAPPA Deterministic guardrails that don't break agents
- EverMind-AI/Raven The Harness of Harnesses โข built for RSI: a trusted, persistent, self-evolving multi-agent ecosystem for all-domain coll
- awlevin/typesafe-computer-use Computer use for about $0.0002 a step: OCR the screen, classify the next action with TypeSafe, click. macOS.
- PostHog/jeeves Jeeves โ Reasoning improves Jev-like decision models
- Rizzo-AI-Academy/rizzo-flow The open, local take on Jev: typed decisions from an LLM, without generating a single token
- ivankovic/codediff Fast, robust, accurate diffing
- ethanplusai/astra-flash-orchestrator Astra plans and reviews; DeepSeek Flash builds. A native Codex workflow with phased tasks, verification, safe installati
- andreylukin/where-next Ask your repo "where is X?" and get the 2โ3 files to open. A local model that learns from your git history, fo