← The field guide

Did AI save time after review and rework?

Compare the full cost of accepted work, including preparation, supervision, review, and corrections. Includes a local calculator and a reusable time ledger.

The short answer

Compare equivalent accepted work from start to finish. Count human preparation, active supervision, review, rework, and setup alongside tool costs. Keep elapsed delivery time and quality separate. A faster first draft can still require more total work.

Define accepted work before timing it

Choose a repeatable task and a shared acceptance checklist. For product copy, that might require correct variant facts, source-backed claims, the requested format, and no missing required fields. Use the same standard for the manual and assisted versions.

Include rejected drafts, retries, and the work needed to finish them. Comparing only successful AI attempts with all manual attempts biases the result. Different quality or scope makes the comparison unresolved, even if the time arithmetic is correct.

Source: Anthropic: Demystifying evals for AI agents

Count active work and waiting separately

Record active person-minutes for preparation, generation supervision, review, and rework. Include one-time setup in the pilot batch; report later recurring batches separately. If several people contribute, sum their actual active time and use role-specific costs in the detailed ledger.

An unattended ten-minute generation is elapsed time, not automatically ten human labor minutes. A person reviewing another task during that wait cannot spend the same minute twice. Track delivery time with timestamps; overlapping work means it will not equal the sum of all stage durations.

RecordIncludeKeep separate
Manual baselinePreparation through final accepted outputEstimated baseline versus directly timed baseline
Assisted laborSetup, input preparation, active supervision, review, and fixesUnattended waits and overlapping tasks
Tool costIncremental charges plus an explicit share of subscriptions where relevantAlready-counted costs and assumed cash savings
QualityAcceptance result and rejected attemptsSpeed of an unaccepted draft

Source: METR: February 2026 experiment-design update

Compare an illustrative batch

The starting values model 20 accepted product descriptions. Manual work takes 18 minutes each, including review. Assisted work takes 2 minutes of preparation, 1 of active supervision, 4 of review, and 2 of rework per item, plus 60 minutes of setup for the batch.

The calculator stays in your browser. It uses one labor rate and one batch of equivalent work; it does not assess output quality, connect to your accounts, or send your amounts to analytics.

Time, review & cost

Illustrative starting values
Compare the same number of equivalent accepted items.
Include preparation, review, and corrections.
Human time preparing sources and instructions.
Active human work during generation; exclude unattended waits.
Inspect the work against the shared acceptance standard.
Include retries, corrections, and final acceptance.
Include one-time setup here; do not count it again per item.
One chosen rate for both methods, in your chosen currency.
Same currency as labor; include an explicit allocation if needed.
Your assessment. This calculator cannot inspect the work. Changing inputs resets this answer.
Scenario assessmentQuality still needs review

Figures below describe your assumptions. Confirm equivalent accepted work before treating a modeled advantage as a result.

Manual human time
360 min
Assisted human time
240 min
Human time released
120 min
Time reduction
33.33%
Manual labor value
360
Assisted labor value + tools
252
Modeled value after tools
108
Break-even review + rework per item
11.4 min

All cost figures use your chosen currency. Labor value is not cash saved. Waiting and elapsed delivery time are separate. Inputs and downloads stay in your browser.

Check the full arithmetic

Manual time = 20 × 18 = 360 minutes. Assisted time = 60 + 20 × (2 + 1 + 4 + 2) = 240 minutes. That releases 120 human minutes, or one-third of the manual baseline, if the work meets the same standard.

At an illustrative 60 USD per hour and 12 USD of tool cost for the batch, manual labor value is 360 USD. Assisted labor value plus tools is 252 USD. The modeled advantage is 108 USD. This is the value of capacity under the stated rate, not proof that payroll fell or revenue increased.

With eight minutes of rework per item, assisted time becomes 360 minutes and the tool cost leaves the batch 12 USD more expensive. With twelve minutes of rework, assisted time becomes 440 minutes and the modeled disadvantage is 92 USD. All three cases are fictional.

Rework per itemAssisted human minutesTime releasedModeled value after tools, USD
2 min240120 min108
8 min3600 min−12
12 min440−80 min−92
Download the fictional time ledger (CSV)

Reuse the formulas with explicit assumptions

Manual minutes = item count × manual minutes per item. Assisted minutes = setup minutes + item count × (preparation + active supervision + review + rework minutes per item). Time released = manual minutes − assisted minutes.

Modeled value after tools = time released ÷ 60 × labor rate − batch tool cost. Time reduction = time released ÷ manual minutes. Negative results remain negative.

For a positive labor rate, the break-even review-plus-rework allowance per item is manual minutes per item − preparation − supervision − (setup minutes + tool cost × 60 ÷ labor rate) ÷ item count. This is an allowance within this batch, not a forecast of review effort. A negative allowance means the other modeled costs already exceed the manual baseline. At a zero labor rate, this monetary threshold is undefined.

Download the three scenarios and expected results (JSON)

Measure a pilot before claiming a gain

Preselect comparable tasks, rotate or randomize the method where practical, and preserve task difficulty, workflow version, dates, and failed attempts. Have the reviewer apply the same acceptance checklist. Avoid timing the same task twice in sequence without acknowledging the learning advantage.

A 2025 METR randomized study found slower completion in one experienced-developer setting. Its February 2026 follow-up explained that selection effects and concurrent-agent time measurement made newer estimates unreliable. These are bounded research observations, not a universal verdict on today’s tools or on marketing work.

Use the ledger to report your own scoped result: accepted output, total person-minutes, tool cost, elapsed delivery time, and remaining quality issues. Describe estimated inputs as estimates. Keep the raw observations available for checking.

Download the ledger, instructions, and blank row (JSON)

Source: METR: Early-2025 developer productivity study; METR: February 2026 experiment-design update

Decide what changes in the workflow

If quality is unresolved, finish the acceptance review before calling the run a success. If rework consumes the modeled gain, inspect the input quality, task boundaries, and recurring defects. A shorter generation time will not necessarily fix that bottleneck.

If the same accepted work consistently requires less total effort, identify how to use the released capacity. Keep capacity, cash cost reduction, throughput, and revenue as separate outcomes. A calculator demonstrates the assumptions; repeatable observations establish the result.

Sources & scope

Primary references checked for this edition. The notes below distinguish source-backed facts from the frameworks and examples proposed in this guide.

  1. Anthropic: Demystifying evals for AI agents ↗

    Supports explicit success criteria, outcome checking, and complementary human review. This ledger and calculator are original editorial tools.

    Checked September 13, 2026
  2. METR: Early-2025 developer productivity study ↗

    A bounded randomized study of experienced developers; its results are not generalized to current models or commerce production.

    Checked September 13, 2026
  3. METR: February 2026 experiment-design update ↗

    Documents selection effects and difficulties measuring concurrent AI-assisted work; newer estimates were described as unreliable.

    Checked September 13, 2026

AI-assisted research and drafting. Provider-specific claims link to primary sources. Frameworks are editorial proposals; worked examples are illustrative and are not employer performance results.

Editorial policy & corrections ↗

Cite this guide

George Kelly. Did AI save time after review and rework? iamgeorgekelly. Updated September 18, 2026. https://www.iamgeorgekelly.com/field-guide/ai-review-rework-cost

Keep the source notes and example labels with an excerpt. For a provider requirement, follow the original documentation in Sources & scope.

Choose your next task ↗ · Browse the source directory ↗

Keep going

How do you check AI product copy before publishing?From request to release: a supervised commerce workflow.ROAS calculator: break-even and contribution