Over ~12 weeks, two people driving an AI Developer Workflow shipped a full production work-order SaaS for Mars — 34 specs, a FastAPI backend, a React front end, two external integrations, and a self-testing QA fleet. The metered AI cost is measured. What a conventional team would have charged to build the same thing is estimated. This page puts both on one axis.
A whole-system tally from git — your commits are the traceable spine, but the figures cover the entire Mars build.
This is not a prototype. The surface includes Entra external-identity auth, a guarded work-order state machine, compliance-document uploads, hauler notifications, two live external integrations, real-time status alerts, API + prod-DB network hardening, an observability layer, a bespoke ops CLI, and an agent fleet that regression-, load-, security-, and UAT-tests the system on its own. The defining metric: 85% of every commit in the repo carries a Co-Authored-By: Claude trailer — this was an AI-driven build end to end, with humans owning specs, review, and sign-off.
Billed USD for the mars-adw-ci-github workspace, pulled live from the Anthropic Admin Cost API for the exact build window.
| Period (2026) | Billed $ | Note |
|---|---|---|
| June | $0.00 | No spend under this workspace — early-June usage sits in shared/personal keys |
| July | $5,201.49 | Core backend + integrations build-out |
| August | $12,930.13 | QA-agent fleet, hardening, UAT, prod deploy |
| Full window (Jun 9 – Sep 1) | $18,131.63 | 170.1M tokens · ~99.997% cache-read · 1.86M output |
$18,131.63 is the defensible floor — the billed spend for the ADW CI workspace over the whole build. Two honest caveats push the true AI cost somewhat higher, but unmeasurably so: the workspace shows zero June spend even though the build started June 9 (that early usage lands in an unattributable shared/default bucket), and the two humans' interactive Claude Code sessions ran on separate keys the workspace total doesn't capture. There is no per-task or per-project tag in the billing data — only workspace and key — so a cleaner isolate isn't reconstructable. Every conclusion below holds even if you double this number.
Bottom-up effort estimate to build the identical system by hand, by workstream. Figures are estimates; assumptions stated.
| Workstream | Scope | Person-days |
|---|---|---|
| Backend / API | State machine, auth, customer API, cietrade + QuickBooks integrations, compliance uploads, hardening | 130 |
| Front end | ~37.6K LOC React/TSX — WO UI, auth flows, status chips, real-time alerts, code-splitting | 95 |
| QA program | Human equivalent of the agent fleet: functional, regression, integration, load, red-team, golden-path, UAT (~249 scenarios) | 130 |
| DevOps / infra | Bicep IaC, dual-env CI/CD, DB migration & release process, ops tooling, observability | 55 |
| Architecture & specs | 34 specs designed to ADR depth | 35 |
| Documentation | Knowledge graph, runbooks, published docs | 20 |
| Management & review | ~15% coordination / code-review overhead | 70 |
| Total effort | ≈ 107 person-weeks → a team of ~5–6 over ~4–5 months | 535 |
Costed two ways, per your ask:
| Baseline | Rate / productivity assumption | Cost |
|---|---|---|
| Traditional, no-AI, US in-house | 535 pd × 8h × $80/hr fully-loaded blended (mid QA + SDET + devs) | $342K |
| ↳ rate sensitivity | $65/hr → $278K · $100/hr → $428K | $278–428K |
| AI-assisted / blended-shore | AI compresses codeable effort ~35% (≈350 effective pd) at a lower blended rate | $150–225K |
| Human column (band) | conservative AI-assisted → traditional no-AI | $150K–430K |
Agency/contractor delivery at $120–180/hr would sit at the top of — or above — this band; an offshore-only build could undercut it, but typically without the code-traceability, IaC, and self-maintaining QA fleet this system shipped with. The mid-point most leaders would recognize is ~$340K.
Total delivered cost for the identical scope. Green = how it was actually built; amber = the counterfactual.
The all-in ADW figure (~$78K) is the honest total for how Mars was really built: the measured $18.1K of LLM spend plus the two humans' oversight — spec authoring, PR review, sign-off — estimated at ~$45–65K of part-time labor across the window. The LLM-only figure ($18.1K) is the pure marginal cost of the "labor" the automation replaced. Bars share a common axis; time compressed too — ~12 weeks with 2 people vs ~4–5 months with a team of 5–6.
Three independent readings of the same result — cost, speed, and leverage.
The two cost multiples aren't in conflict — they answer different questions. ~4.4× is what the CFO saved versus commissioning the build conventionally. ~19× is the raw automation dividend: the marginal machine cost of the authorship itself. The truth sits between them and depends on how you price the human oversight, which is why both are shown. On the non-dollar axes, the story is just as strong: the same output that conventionally needs a five-to-six-person team came from two people, delivered in roughly half the calendar time — and it left behind a standing QA agent fleet that keeps testing the system for free as it changes.
The advantage is real and large. A credible case study still names its own limits.
Sources. Measured: git history of mars_workorder_system (commits, PRs, LOC, co-authorship) and the Anthropic Admin Cost/Usage API for the mars-adw-ci-github workspace, 2026-06-09 → 2026-09-01. Estimated: bottom-up person-day model at industry-standard full-lifecycle productivity and US fully-loaded blended labor cost ($65–100/hr), assumptions stated inline. The ADW column is measured; the human column is an estimate presented as a band. Ratios: all-in ~$78K vs human ~$340K → ~4.4×; LLM-only $18.1K → ~19×.