Arc theory.est. 2023

From experiment to infrastructure

A study of the practice · 2024 – 2026

Here's a look at how we measured whether AI actually worked for us — a retroactive study of our own adoption, session by session, and a method any team could run on its own.

−48%
shorter median development session, first tracked quarter to latest
1.7×
labor margin on identical fixed-scope work, price unchanged
1,412
client work sessions tracked, 2024–2026

The hypothesis is simple: AI produces optimal results when it's incorporated in a structured, systematic way — built into the work, not bolted on. The corollary is that the difference is measurable. So we went back through our own practice and measured it: every logged session, before and after the tooling matured.

The unit is the median duration of each logged session; the category view aggregates by quarter.

  • Correlation, not proof — the decline overlaps a shifting client mix, so the time data can't isolate AI's contribution on its own.
  • The one number that controls for all of it is the fixed-scope margin.

It ran in three phases. Manual work was the baseline. Then came experimentation — AI in chats and in-tool, but not yet in the workflow: faster on average, still volatile, and carrying a real token-and-overhead tax. The turn came when the work got structured — triaging, insights, and agents built into the workflow. That's where the time settled lowest and, more tellingly, steadiest.

0306090Pre-AI→ Mar 2024 · manual45mspread 60ExperimentingMar 2024 – Aug 2025 · in-tool & chat36mspread 46AdoptingAug 2025 → · structured24mspread 37

Done by hand. Fast some days, slow others, and built around one person. The baseline.

In-tool AI features and chat-based prompting — used alongside the work, not yet built into it. Faster on average, but volatile, and carrying a real token + overhead tax. Worked; won't scale.

AI built into the workflow itself — structured inputs and outputs, with triaging, insights, and agents running as repeatable steps. Lowest time and the tightest spread: stable and lasting.

Fig. 01 Median session length (line) and the middle-50% spread, p25–p75 (bar), by phase. Internal builds excluded; pre-AI includes 2023 (rougher tracking). The experimentation tax was in token spend and build overhead — noted here, not drawn.
experimentingadopting03060202420252026
Development−48% · 46→24mBuilding and shipping features or infrastructure.
Planning−90% · 60→6mScoping the work before building.
Analysis−74% · 31→8mInvestigating data & problems. Directional.
Reporting≈16m · steadyShort & steady once the reporting agent shipped (2025).
Fig. 02 Median minutes per session, by quarter and category, 2024–2026 — one shared scale. Bands mark the phases; the dashed line marks structured tooling (Aug 2025). Reporting begins when its agent shipped. Some quarters are directional (small samples).

The wins came from building AI into the workflow on one repeatable pattern: prepare the data, let a webhook fire the right agent, and always return a structured output. The more disciplined the inputs and outputs, the more reliable the result.

Data preparationclean, structured inputs
Webhook-triggered agentsthe right agent fires
JSON payload / MCPby task type
Structured outputsalways
Meeting recap & actionstranscript action items

Turns a transcript into triaged, assigned action items.

Weekly client reportthe week's work report + insights

Structures the week's work into a report the team reads — and an AI rates for signal.

Documentation draftsdecisions & changes records & docs

Drafts decision records and docs from the decisions and changes as they happen.

Estimation & ticketsa goal scoped tickets

Scopes a goal into user stories with requirements, written out as tickets.

Workflow diagramsa described process a diagram

Generates a process diagram from a plain-language description.

Everything above can be argued with — except this. Across identical contracts delivered at the same fixed price, the hours to deliver fell. That's a 1.7× labor margin, earned purely in execution, with nothing in the pricing or the scope to explain it away.

1.7×
labor margin on fixed-scope work

Structure AI around repeatable processes — with structured inputs and structured outputs — and it compounds into infrastructure. Skip the structure and it stays an expensive experiment that won't scale. The measuring is the point: any team can run this, and most should.

Source: practice time-tracking — n=1,412 client sessions tracked across different categories, 2024–2026. Median session length was used.