search
coding benchmark
Trends
Nothing in this window yet.
Repos
- DietrichGebert/ponytail Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
- ninjahawk/livenerf Benchmark for tracking model capability after release.
- archestra-ai/OpenAPPA Deterministic guardrails that don't break agents
- EverMind-AI/Raven The Harness of Harnesses • built for RSI: a trusted, persistent, self-evolving multi-agent ecosystem for all-domain coll
- PostHog/jeeves Jeeves – Reasoning improves Jev-like decision models
- ivankovic/codediff Fast, robust, accurate diffing
- Rizzo-AI-Academy/rizzo-flow The open, local take on Jev: typed decisions from an LLM, without generating a single token
- andreylukin/where-next Ask your repo "where is X?" and get the 2–3 files to open. A local model that learns from your git history, fo
- ethanplusai/astra-flash-orchestrator Astra plans and reviews; DeepSeek Flash builds. A native Codex workflow with phased tasks, verification, safe installati
- awlevin/typesafe-computer-use Computer use for about $0.0002 a step: OCR the screen, classify the next action with TypeSafe, click. macOS.