The short answer
Compare equivalent accepted work from start to finish. Count human preparation, active supervision, review, rework, and setup alongside tool costs. Keep elapsed delivery time and quality separate. A faster first draft can still require more total work.
Define accepted work before timing it
Choose a repeatable task and a shared acceptance checklist. For product copy, that might require correct variant facts, source-backed claims, the requested format, and no missing required fields. Use the same standard for the manual and assisted versions.
Include rejected drafts, retries, and the work needed to finish them. Comparing only successful AI attempts with all manual attempts biases the result. Different quality or scope makes the comparison unresolved, even if the time arithmetic is correct.
Count active work and waiting separately
Record active person-minutes for preparation, generation supervision, review, and rework. Include one-time setup in the pilot batch; report later recurring batches separately. If several people contribute, sum their actual active time and use role-specific costs in the detailed ledger.
An unattended ten-minute generation is elapsed time, not automatically ten human labor minutes. A person reviewing another task during that wait cannot spend the same minute twice. Track delivery time with timestamps; overlapping work means it will not equal the sum of all stage durations.
| Record | Include | Keep separate |
|---|---|---|
| Manual baseline | Preparation through final accepted output | Estimated baseline versus directly timed baseline |
| Assisted labor | Setup, input preparation, active supervision, review, and fixes | Unattended waits and overlapping tasks |
| Tool cost | Incremental charges plus an explicit share of subscriptions where relevant | Already-counted costs and assumed cash savings |
| Quality | Acceptance result and rejected attempts | Speed of an unaccepted draft |
Compare an illustrative batch
The starting values model 20 accepted product descriptions. Manual work takes 18 minutes each, including review. Assisted work takes 2 minutes of preparation, 1 of active supervision, 4 of review, and 2 of rework per item, plus 60 minutes of setup for the batch.
The calculator stays in your browser. It uses one labor rate and one batch of equivalent work; it does not assess output quality, connect to your accounts, or send your amounts to analytics.
Time, review & cost
Figures below describe your assumptions. Confirm equivalent accepted work before treating a modeled advantage as a result.
- Manual human time
- 360 min
- Assisted human time
- 240 min
- Human time released
- 120 min
- Time reduction
- 33.33%
- Manual labor value
- 360
- Assisted labor value + tools
- 252
- Modeled value after tools
- 108
- Break-even review + rework per item
- 11.4 min
All cost figures use your chosen currency. Labor value is not cash saved. Waiting and elapsed delivery time are separate. Inputs and downloads stay in your browser.
Check the full arithmetic
Manual time = 20 × 18 = 360 minutes. Assisted time = 60 + 20 × (2 + 1 + 4 + 2) = 240 minutes. That releases 120 human minutes, or one-third of the manual baseline, if the work meets the same standard.
At an illustrative 60 USD per hour and 12 USD of tool cost for the batch, manual labor value is 360 USD. Assisted labor value plus tools is 252 USD. The modeled advantage is 108 USD. This is the value of capacity under the stated rate, not proof that payroll fell or revenue increased.
With eight minutes of rework per item, assisted time becomes 360 minutes and the tool cost leaves the batch 12 USD more expensive. With twelve minutes of rework, assisted time becomes 440 minutes and the modeled disadvantage is 92 USD. All three cases are fictional.
| Rework per item | Assisted human minutes | Time released | Modeled value after tools, USD |
|---|---|---|---|
| 2 min | 240 | 120 min | 108 |
| 8 min | 360 | 0 min | −12 |
| 12 min | 440 | −80 min | −92 |
Reuse the formulas with explicit assumptions
Manual minutes = item count × manual minutes per item. Assisted minutes = setup minutes + item count × (preparation + active supervision + review + rework minutes per item). Time released = manual minutes − assisted minutes.
Modeled value after tools = time released ÷ 60 × labor rate − batch tool cost. Time reduction = time released ÷ manual minutes. Negative results remain negative.
For a positive labor rate, the break-even review-plus-rework allowance per item is manual minutes per item − preparation − supervision − (setup minutes + tool cost × 60 ÷ labor rate) ÷ item count. This is an allowance within this batch, not a forecast of review effort. A negative allowance means the other modeled costs already exceed the manual baseline. At a zero labor rate, this monetary threshold is undefined.
Download the three scenarios and expected results (JSON)Measure a pilot before claiming a gain
Preselect comparable tasks, rotate or randomize the method where practical, and preserve task difficulty, workflow version, dates, and failed attempts. Have the reviewer apply the same acceptance checklist. Avoid timing the same task twice in sequence without acknowledging the learning advantage.
A 2025 METR randomized study found slower completion in one experienced-developer setting. Its February 2026 follow-up explained that selection effects and concurrent-agent time measurement made newer estimates unreliable. These are bounded research observations, not a universal verdict on today’s tools or on marketing work.
Use the ledger to report your own scoped result: accepted output, total person-minutes, tool cost, elapsed delivery time, and remaining quality issues. Describe estimated inputs as estimates. Keep the raw observations available for checking.
Download the ledger, instructions, and blank row (JSON)Source: METR: Early-2025 developer productivity study; METR: February 2026 experiment-design update
Decide what changes in the workflow
If quality is unresolved, finish the acceptance review before calling the run a success. If rework consumes the modeled gain, inspect the input quality, task boundaries, and recurring defects. A shorter generation time will not necessarily fix that bottleneck.
If the same accepted work consistently requires less total effort, identify how to use the released capacity. Keep capacity, cash cost reduction, throughput, and revenue as separate outcomes. A calculator demonstrates the assumptions; repeatable observations establish the result.
Sources & scope
Primary references checked for this edition. The notes below distinguish source-backed facts from the frameworks and examples proposed in this guide.
- Anthropic: Demystifying evals for AI agents ↗
Supports explicit success criteria, outcome checking, and complementary human review. This ledger and calculator are original editorial tools.
Checked September 13, 2026 - METR: Early-2025 developer productivity study ↗
A bounded randomized study of experienced developers; its results are not generalized to current models or commerce production.
Checked September 13, 2026 - METR: February 2026 experiment-design update ↗
Documents selection effects and difficulties measuring concurrent AI-assisted work; newer estimates were described as unreliable.
Checked September 13, 2026
AI-assisted research and drafting. Provider-specific claims link to primary sources. Frameworks are editorial proposals; worked examples are illustrative and are not employer performance results.
Editorial policy & corrections ↗Cite this guide
George Kelly. Did AI save time after review and rework? iamgeorgekelly. Updated September 18, 2026. https://www.iamgeorgekelly.com/field-guide/ai-review-rework-cost
Keep the source notes and example labels with an excerpt. For a provider requirement, follow the original documentation in Sources & scope.