Showcase · claims ship with receipts
One sentence in,
one merged PR out.
Write-ups, with receipts, from the team building RunAI Coder — an autonomous coding agent. Every claim traces back to a machine-recorded ledger or a reproducible script.
968
merged squash PRs in 7 days
98.8%
first-pass completion
2.83×
median context compression
$0.26/M
effective input price
Measured, not projected — one developer, one machine, 2026-07-14 → 07-21 PDT. Formulas, sample windows and known limitations are in the performance report.
Posts
- 008 · Aug 6, 2026Nobody reads your architecture overview. Including the agent.Why an instruction file is rent, not documentation: three conflicting evals reconciled, a five-lane probe on Flask, and the harness that never read our AGENTS.md at all
- 007 · Aug 5, 202633k tokens before hello: the anatomy of your coding agent's preambleWhat the viral token-overhead study really found: configuration swamps defaults, cache churn beats size, and a two-minute way to measure your own
- 006 · Aug 4, 2026Why your coding agent forgets: a field guide to context rotThe two-budget model of a context window, what compaction quietly costs, and what helps if you just use these tools
- 005 · Aug 3, 2026The bill is 99% input tokens: cost engineering for a coding agentOne ledger day, three levers from $10/M to an effective $0.26/M, and the SWE-bench A/B that keeps it honest
- 002 · Jul 31, 2026Our house rules for letting an AI agent commit codeFive boring-on-purpose guardrails for giving an agent write access, and what they don't solve
- 001 · Jul 30, 2026One sentence in, one merged PR outWhat RunAI Coder is, a week of dogfooding numbers with the caveats attached, and why a merged PR costs $10