04 · Measuring the Impact of AI-Assisted Work¶
"AI is helping" is not a measurement — it's an impression, and impressions are exactly what led to the "looks right" trap in Level 3, Module 9, applied now to the impact of a whole workflow rather than one piece of output. This module covers how to measure whether AI-assisted work is actually delivering value, not just activity.
Why measurement usually gets skipped¶
Measuring impact is harder than measuring adoption (how many people used the tool, how many prompts were run), so organizations often report adoption numbers as if they were impact numbers. Adoption is necessary but says nothing about whether the work got better, faster, or worse — a team can run thousands of prompts a month and be net slower once review and correction time is counted.
Metrics that actually indicate impact¶
| Category | What to measure | Why it beats a vaguer proxy |
|---|---|---|
| Time | Time from task start to accepted final output, including review/correction | Captures the whole cost, not just the drafting speed |
| Quality | Error rate caught before vs. after external release | Adoption doesn't matter if quality drops |
| Rework | How often AI-assisted output needs a second pass vs. non-AI output | Surfaces workflows where AI is adding, not removing, work |
| Consistency | Variance in output quality across people doing the same task | A shared template should shrink this; if it doesn't, the template isn't being used correctly |
| Actual usage of what was built | Whether the shared templates/processes from Module 8 (Level 3) are still used weeks later | Distinguishes a real capability from a one-time training event |
Baselines matter more than the after-number¶
A time-savings claim is only meaningful against a real baseline measured the same way, on comparable work, before the change. "This used to take about half a day" from memory is not a baseline — memory of "before" tends to compress once the "after" feels easier. Where possible, measure a handful of representative tasks the old way and the new way close in time, before drawing a conclusion.
Reporting impact honestly¶
Report the full picture, including the costs: time spent learning, instances where AI-assisted output needed significant correction, and tasks where it didn't help. A report that only lists wins will be right once and then lose credibility the first time someone finds a counterexample; a report that's honest about tradeoffs is what actually earns continued investment.
How It Actually Works¶
Measuring impact honestly requires the same discipline as evaluating a single output (Module 9, Level 3), scaled up — and for the same mechanistic reason: fluent, confident-feeling AI-assisted work is not automatically better work.
Activity metrics measure exposure to the mechanism, not its effect. Counting prompts run or people using the tool tells you how much generation is happening, but generation's defining property — plausible output that reads well regardless of whether it's actually better than the alternative — means volume of use has no necessary relationship to quality of outcome. This is the direct, mechanistic reason "adoption" and "impact" are different measurements that can diverge sharply: heavy use of a tool whose output isn't being verified can produce a lot of fluent, unreliable work faster.
Baselines matter because sampling variance and selection bias both distort a single after-number. An impressive AI-assisted result reported without a baseline is one sample of a generation process that inherently varies (Module 9, Level 3) — without knowing what the same task would have produced without AI assistance, or across several attempts, an isolated "after" number can't distinguish real, repeatable improvement from a favorable draw or from measuring the easiest cases first.
Exercise¶
Pick one AI-assisted task your team does regularly. Design a simple before/after measurement for it using two metrics from the table above — specify exactly what you'd measure, on how many instances, and what result would actually change your mind about whether the workflow is worth keeping.