ADW ROI · Mars Work Order System

The same production build, priced two ways.

Over ~12 weeks, two people driving an AI Developer Workflow shipped a full production work-order SaaS for Mars — 34 specs, a FastAPI backend, a React front end, two external integrations, and a self-testing QA fleet. The metered AI cost is measured. What a conventional team would have charged to build the same thing is estimated. This page puts both on one axis.

$18.1K
Measured ADW pipeline spend, full build window
Measured
$150K–430K
Human-team equivalent for the same scope
Estimated
~4×–19×
Cost advantage, all-in vs LLM-only framing
Derived
234
PRs merged · 34 specs · ~12 weeks · 2 humans
Measured
Measured = from git & the Anthropic billing API Estimated = bottom-up model, assumptions stated
01

What was actually delivered

A whole-system tally from git — your commits are the traceable spine, but the figures cover the entire Mars build.

34
specs shipped (33 numbered + smoke), planned → validated → deployed
~88.7K
lines of app source (45.7K Python · 37.6K TS/TSX · 5.5K docs)
~16.1K
lines of tests · 81 Python suites · 31 Playwright/e2e specs
234
pull requests merged (167 under your identity)
664
commits total, of which 566 (85%) are Claude co-authored
14 / 14
Bicep IaC files / CI-CD workflows to sandbox & prod
19 / 7
custom slash-commands / standing QA & ops subagents
~12 wks
calendar: 2026-06-09 → 2026-09-01 · 2 humans + ADW

This is not a prototype. The surface includes Entra external-identity auth, a guarded work-order state machine, compliance-document uploads, hauler notifications, two live external integrations, real-time status alerts, API + prod-DB network hardening, an observability layer, a bespoke ops CLI, and an agent fleet that regression-, load-, security-, and UAT-tests the system on its own. The defining metric: 85% of every commit in the repo carries a Co-Authored-By: Claude trailer — this was an AI-driven build end to end, with humans owning specs, review, and sign-off.

Entra external identityWO state machineCompliance uploadsHauler notificationscietrade syncQuickBooks billingReal-time alertsLoad & resilienceRed-team QAOrbit ops CLILogfire observabilityProd-DB hardening
02

The measured ADW cost

Billed USD for the mars-adw-ci-github workspace, pulled live from the Anthropic Admin Cost API for the exact build window.

Period (2026)Billed $Note
June$0.00No spend under this workspace — early-June usage sits in shared/personal keys
July$5,201.49Core backend + integrations build-out
August$12,930.13QA-agent fleet, hardening, UAT, prod deploy
Full window (Jun 9 – Sep 1)$18,131.63170.1M tokens · ~99.997% cache-read · 1.86M output

$18,131.63 is the defensible floor — the billed spend for the ADW CI workspace over the whole build. Two honest caveats push the true AI cost somewhat higher, but unmeasurably so: the workspace shows zero June spend even though the build started June 9 (that early usage lands in an unattributable shared/default bucket), and the two humans' interactive Claude Code sessions ran on separate keys the workspace total doesn't capture. There is no per-task or per-project tag in the billing data — only workspace and key — so a cleaner isolate isn't reconstructable. Every conclusion below holds even if you double this number.

03

The human-team equivalent

Bottom-up effort estimate to build the identical system by hand, by workstream. Figures are estimates; assumptions stated.

WorkstreamScopePerson-days
Backend / APIState machine, auth, customer API, cietrade + QuickBooks integrations, compliance uploads, hardening130
Front end~37.6K LOC React/TSX — WO UI, auth flows, status chips, real-time alerts, code-splitting95
QA programHuman equivalent of the agent fleet: functional, regression, integration, load, red-team, golden-path, UAT (~249 scenarios)130
DevOps / infraBicep IaC, dual-env CI/CD, DB migration & release process, ops tooling, observability55
Architecture & specs34 specs designed to ADR depth35
DocumentationKnowledge graph, runbooks, published docs20
Management & review~15% coordination / code-review overhead70
Total effort≈ 107 person-weeks → a team of ~5–6 over ~4–5 months535

Costed two ways, per your ask:

BaselineRate / productivity assumptionCost
Traditional, no-AI, US in-house535 pd × 8h × $80/hr fully-loaded blended (mid QA + SDET + devs)$342K
 ↳ rate sensitivity$65/hr → $278K · $100/hr → $428K$278–428K
AI-assisted / blended-shoreAI compresses codeable effort ~35% (≈350 effective pd) at a lower blended rate$150–225K
Human column (band)conservative AI-assisted → traditional no-AI$150K–430K

Agency/contractor delivery at $120–180/hr would sit at the top of — or above — this band; an offshore-only build could undercut it, but typically without the code-traceability, IaC, and self-maintaining QA fleet this system shipped with. The mid-point most leaders would recognize is ~$340K.

04

Side by side

Total delivered cost for the identical scope. Green = how it was actually built; amber = the counterfactual.

Human — traditional no-AIest. $278–428K
~$340K
Human — AI-assisted / blendedest. $150–225K
~$185K
ADW build — all-inLLM + human oversight
~$78K
ADW pipeline — LLM onlymeasured, billed
$18.1K

The all-in ADW figure (~$78K) is the honest total for how Mars was really built: the measured $18.1K of LLM spend plus the two humans' oversight — spec authoring, PR review, sign-off — estimated at ~$45–65K of part-time labor across the window. The LLM-only figure ($18.1K) is the pure marginal cost of the "labor" the automation replaced. Bars share a common axis; time compressed too — ~12 weeks with 2 people vs ~4–5 months with a team of 5–6.

05

How much ADW improved development

Three independent readings of the same result — cost, speed, and leverage.

~4.4×
cheaper, all-in vs a traditional team
~$78K vs ~$340K
~19×
cheaper on LLM-marginal cost alone
$18.1K vs ~$340K
~3×
headcount leverage — 2 people did a 5–6-person job
~1.7× faster calendar too

The two cost multiples aren't in conflict — they answer different questions. ~4.4× is what the CFO saved versus commissioning the build conventionally. ~19× is the raw automation dividend: the marginal machine cost of the authorship itself. The truth sits between them and depends on how you price the human oversight, which is why both are shown. On the non-dollar axes, the story is just as strong: the same output that conventionally needs a five-to-six-person team came from two people, delivered in roughly half the calendar time — and it left behind a standing QA agent fleet that keeps testing the system for free as it changes.

06

Reading it honestly

The advantage is real and large. A credible case study still names its own limits.

  • Different currencies. The ADW number is marginal API cost; the human number is fully-loaded labor. That gap is the automation dividend — but it's the honest framing, not a like-for-like invoice.
  • The counterfactual is an estimate. No one built this system by hand to compare against. The 535 person-days is a defensible bottom-up model, but it's a model — the human column is a band, not a quote.
  • Attribution has soft edges. $18.1K is the ADW-CI workspace only; early-June and interactive human sessions aren't captured and can't be cleanly isolated from the billing data. Treat it as a measured floor.
  • Oversight isn't free, and it isn't zero-skill. Two experienced engineers scoped, steered, and reviewed everything. The model doesn't run itself — the ~$78K all-in figure already prices that in.
  • Throughput, not a quality-equivalence claim. Agents and humans miss different things. The build passed its own regression, load, red-team, and UAT gates, but this is a cost-and-speed case, not a proof of identical quality to a hand-built system.

Sources. Measured: git history of mars_workorder_system (commits, PRs, LOC, co-authorship) and the Anthropic Admin Cost/Usage API for the mars-adw-ci-github workspace, 2026-06-09 → 2026-09-01. Estimated: bottom-up person-day model at industry-standard full-lifecycle productivity and US fully-loaded blended labor cost ($65–100/hr), assumptions stated inline. The ADW column is measured; the human column is an estimate presented as a band. Ratios: all-in ~$78K vs human ~$340K → ~4.4×; LLM-only $18.1K → ~19×.