Measure AI Productivity After Review and Rework

Z

ZharfAI Team

September 6, 202611 min read
Measure AI Productivity After Review and Rework

A support manager sees drafts appearing in seconds. The team likes the assistant, and the renewal proposal promises hundreds of hours saved. Yet senior reviewers are staying late to check policy exceptions, and reopened tickets are accumulating. Has the assistant improved productivity, or moved work somewhere the dashboard does not count?

That is an illustrative situation, not a ZharfAI customer result. It captures the decision this guide addresses: whether to expand an AI tool in a particular workflow. The useful unit is accepted work delivered with known human effort and acceptable quality—not generated text, model calls, or estimated minutes saved.

Faster generation is only one part of the job

For a support answer, work begins before drafting and ends after the answer survives the agreed follow-up period. It can include finding context, prompting, checking evidence, editing, escalation, and correcting a reopened case. If another team performs those steps, its labor still belongs in the comparison.

Define the endpoint before collecting results. A resolved ticket might mean the customer's issue is closed and not reopened within a specified window. A software change might need tests, review, merge, and a defined period without a related defect. An invoice extraction might need reconciliation against the original document. The window should fit the actual failure pattern; there is no universal seven-day rule.

Keep several clocks separate. Human labor time is the sum of people's active effort. Elapsed time includes waiting. Throughput is accepted outcomes per time period. AI can improve one while worsening another. Ten parallel generations do not represent ten simultaneous hours of human labor, and a long review queue is not captured by a fast model response.

Read productivity studies as bounded evidence

The final 2025 Generative AI at Work paper studied a staggered rollout involving 5,172 customer-support agents. It reported 15% more issues resolved per hour on average, with larger benefits for less experienced workers and small quality declines among the most skilled. This was a particular firm and workflow, not a randomized estimate for every occupation or today's assistants. QJE study.

METR's July 2025 experiment randomly allowed or disallowed early-2025 AI tools on 246 tasks undertaken by 16 experienced open-source developers. Tool access increased completion time by 19%, although developers perceived acceleration. That result challenges self-reported savings as a measurement method; it does not establish that current coding assistants slow everyone down. METR experiment.

The studies ask different questions in different settings. Averaging their percentages would not produce a useful purchasing benchmark. Treat them as evidence that local task mix, expertise, tooling, and the definition of completion matter. A model's isolated test score is another distinct measurement; our guide to benchmark exposure and contamination explains why even that score needs a clear account of what the test actually measures.

Choose a task family, not an entire job title

“Analysts using AI” is too broad for a decision. Routine document comparison, ambiguous investigation, and final approval can occupy the same person's day while imposing different verification costs. Harvard's September 2023 account of a field experiment with 758 consultants describes a jagged capability frontier: benefits depend on which tasks the system can handle. The relevant lesson is task-level evaluation, not a single adoption verdict for knowledge work. HBS research summary.

Start with a repeatable family whose outcomes are observable and whose risks permit a controlled pilot. Write down eligible work, excluded sensitive cases, required source material, acceptance criteria, and the person authorized to judge completion. Separate straightforward and difficult cases using information available before treatment, rather than labeling disappointing AI results “exceptional” afterward.

Choose one primary decision metric. For example: accepted cases per total human labor hour, with separately enforced limits on serious errors and reopen rates. Report task difficulty and quality alongside the headline. Otherwise, a team can improve the ratio simply by completing easy requests and leaving difficult ones in the queue.

Compare access fairly, including people who do not use it

The following design is ZharfAI's proposed operating method, not a protocol validated by the cited studies. Where feasible, randomly assign eligible workers or work units to the existing workflow or the AI-enabled workflow. Keep the surrounding policy, source access, and acceptance standard comparable. Record the assignment before work begins.

Choose the randomization unit around how work is shared. If colleagues continuously exchange drafts or a common reviewer changes behavior for everyone, individual assignment may mix the two conditions. Team-level assignment can reduce that mixing, but a few teams provide much less independent evidence than thousands of tickets suggest. Get statistical support for the sample-size and uncertainty calculation; do not treat every ticket from one team as an independent experiment.

The main comparison should follow original assignment, including assigned workers who rarely open the assistant. This intention-to-treat estimate answers the operational question: what happens when this group receives access under these conditions? A separate usage analysis can diagnose adoption problems, but comparing enthusiasts with nonusers does not isolate the tool's effect.

Give both groups a clear explanation of the pilot. Do not tie participation, speed, or reported savings to individual performance rankings. Collect task-level timing with an explicit purpose and retention policy, not continuous surveillance of private activity. Workers need a way to report extra checking, hidden tool use, and missing tasks without turning the study into a competition for favorable numbers.

Count the work that disappears from the report

METR's February 2026 update explains why its later experiment could not reliably establish the size of current speed gains. Participant and task selection changed, and concurrent agent use complicated time measurement. The update is a warning about the sample and the clock—not confirmation that the earlier slowdown still describes newer tools. METR methodology update.

Keep an eligibility register before assignment. For each item, retain a stable identifier, task family, difficulty information, assignment, relevant tool version, completion status, reviewer effort, correction effort, and reason for exclusion or withdrawal. Use references to protected work records where possible rather than copying sensitive content into another analytics store.

Reconcile eligible, assigned, started, completed, accepted, and followed-up counts. Unfinished work remains visible at the reporting cutoff. If the AI group abandons more hard cases, analyzing only its completed tasks can create an attractive but misleading result. Publish the backlog and missing follow-up share with the productivity estimate.

Microsoft Research describes sample-ratio mismatch as an experiment-quality warning that should be diagnosed before effects are trusted. Assignment, execution, joins, or filtering can produce the discrepancy. In a workplace pilot, compare observed allocation with the planned randomization and investigate unexpected losses; a clean allocation check alone does not prove that all outcomes are complete. Microsoft experimentation guidance.

Put review and corrections into the arithmetic

Consider a deliberately simplified example. Each group receives 100 comparable cases. All follow-up is complete, and each labor figure includes effort on unsuccessful cases. These are invented numbers for explaining the calculation, not an observed experiment.

MeasureExisting workflowAI-enabled workflow
Research and drafting labor60 hours36 hours
Review labor15 hours28 hours
Correction and reopening labor5 hours12 hours
Total human labor80 hours76 hours
Accepted cases after follow-up9590
Labor hours per accepted case0.8420.844

Drafting labor falls by 40%. Total labor falls by only 5%, while fewer cases reach acceptance. Dividing all labor by accepted cases shows essentially unchanged—and slightly worse—labor efficiency. The apparently impressive drafting improvement is not enough to justify expansion.

Do not average each worker's ratio without considering the intended estimand. For this team-level example, calculate total accepted cases divided by total counted labor, or its reciprocal, consistently across groups. Show the underlying totals so readers can inspect the denominator. The table alone cannot establish statistical significance, customer impact, or causality; those require the actual assignment design, uncertainty, and quality results.

Also inspect who absorbs the extra effort. Saving junior drafting hours while consuming scarce specialist review hours may create a capacity problem even if total labor improves. A single blended hourly rate can hide that bottleneck. Report role-specific effort and review waiting time before converting the result into money.

Protect quality without pretending that no errors means safety

Define serious errors separately from ordinary revisions. Examples include an unsupported policy commitment, a wrong payment amount, or an incorrect access instruction. The applicable workflow should determine the categories. Agree on escalation and pause conditions before the pilot, and retain existing controls; a productivity experiment does not authorize risky outputs to reach customers.

Where practical, assess samples without revealing which workflow produced them. Use the same rubric and calibrate reviewers on examples. Measure disagreement and adjudicate important differences. Otherwise, reviewers' enthusiasm or skepticism about AI can enter the quality score.

Report uncertainty around both productivity and quality. Decide in advance what minimum improvement would matter and what deterioration would be unacceptable. A small pilot with zero observed serious incidents cannot establish that a rare failure is acceptably unlikely. If the sample cannot answer the safety question, limit the claim and scope rather than converting “not detected” into “not present.”

Fix the planned analysis date and follow-up window. If repeated interim decisions are necessary, use a suitable sequential analysis plan. Stopping at the first encouraging result changes the meaning of ordinary uncertainty estimates. Safety monitoring can still operate continuously; stopping harmful exposure and declaring a productivity win are different decisions.

Saved time is not automatically cash or new output

A working paper first issued in May 2025 and revised that November studied randomized access across 66 firms and 7,137 knowledge workers. It reports two fewer email hours per week among tool-using treated workers in the experiment's second half, but no detected shift in task quantity or composition. That distinction matters: a time saving is not itself evidence of greater organizational output. NBER working paper.

Translate a credible local effect into three separate accounts: released capacity, service improvement, and actual expenditure change. Freed minutes may reduce overtime or improve responsiveness. They may also be fragmented across the day and unavailable for another task. A salary does not disappear because a drafting step takes less time.

Subtract recurring licenses, usage charges, administration, quality review, and ongoing training from any cost comparison. Show one-time implementation and learning effort separately with the period over which it is being assessed. For usage-based tools, the agent spending reservation ledger addresses a complementary issue: controlling spend before concurrent work consumes it. Budget compliance and productivity are different tests, and both matter.

Do not annualize a short pilot without stating assumptions about demand, adoption, task mix, and continued performance. A sensitivity range is more informative than a precise-looking annual saving built on an untested constant.

Make a narrow decision, then keep checking it

An expansion case needs more than a positive average. It needs credible assignment, complete enough outcomes, an improvement large enough to matter, acceptable quality, and a feasible use for the released capacity. Report separately whether the conclusion applies to novice workers, experienced workers, or particular task families. Predetermine the main subgroup comparisons and treat unexpected patterns as hypotheses to retest.

If review absorbs the gain, redesign the assistance: better source retrieval, narrower suggestions, or a different handoff may deserve another bounded experiment. If quality fails, do not compensate with a larger speed percentage. If the uncertainty spans both a worthwhile gain and a meaningful loss, report an inconclusive result and decide whether further evidence is worth its cost.

Record the model, configuration, training, and workflow tested. A material change makes an old estimate less portable. After any expansion, use post-deployment monitoring to follow accepted outcomes, reviewer workload, reopenings, and task-mix changes. Monitoring detects deterioration; it does not replace the original comparison or make every later movement causal.

The renewal meeting should end with a precise sentence: this workflow, for this population, delivered this range of accepted-work improvement under these quality limits and costs. If the evidence supports only faster drafts, say that. It is a useful finding—but it is not yet the same as a more productive team.

Source notes — reviewed September 6, 2026

#AI productivity#Workplace experiments#Human review#Enterprise AI#Measurement

Related Posts

Name one process for a discovery call

If this note maps to a real system in your organisation, start with the services page or a shipped case study.