From experiment to infrastructure
A study of the practice · 2024 – 2026
Here's a look at how we measured whether AI actually worked for us — a retroactive study of our own adoption, session by session, and a method any team could run on its own.
The hypothesis is simple: AI produces optimal results when it's incorporated in a structured, systematic way — built into the work, not bolted on. The corollary is that the difference is measurable. So we went back through our own practice and measured it: every logged session, before and after the tooling matured.
The unit is the median duration of each logged session; the category view aggregates by quarter.
- Correlation, not proof — the decline overlaps a shifting client mix, so the time data can't isolate AI's contribution on its own.
- The one number that controls for all of it is the fixed-scope margin.
It ran in three phases. Manual work was the baseline. Then came experimentation — AI in chats and in-tool, but not yet in the workflow: faster on average, still volatile, and carrying a real token-and-overhead tax. The turn came when the work got structured — triaging, insights, and agents built into the workflow. That's where the time settled lowest and, more tellingly, steadiest.
Done by hand. Fast some days, slow others, and built around one person. The baseline.
In-tool AI features and chat-based prompting — used alongside the work, not yet built into it. Faster on average, but volatile, and carrying a real token + overhead tax. Worked; won't scale.
AI built into the workflow itself — structured inputs and outputs, with triaging, insights, and agents running as repeatable steps. Lowest time and the tightest spread: stable and lasting.
The wins came from building AI into the workflow on one repeatable pattern: prepare the data, let a webhook fire the right agent, and always return a structured output. The more disciplined the inputs and outputs, the more reliable the result.
Turns a transcript into triaged, assigned action items.
Structures the week's work into a report the team reads — and an AI rates for signal.
Drafts decision records and docs from the decisions and changes as they happen.
Scopes a goal into user stories with requirements, written out as tickets.
Generates a process diagram from a plain-language description.
Everything above can be argued with — except this. Across identical contracts delivered at the same fixed price, the hours to deliver fell. That's a 1.7× labor margin, earned purely in execution, with nothing in the pricing or the scope to explain it away.
Structure AI around repeatable processes — with structured inputs and structured outputs — and it compounds into infrastructure. Skip the structure and it stays an expensive experiment that won't scale. The measuring is the point: any team can run this, and most should.
Source: practice time-tracking — n=1,412 client sessions tracked across different categories, 2024–2026. Median session length was used.