The vendor numbers all look good. This page is not about whether they are true —
it is about what they measured, and whether that method predicts anything about your organisation.
01 — Vendor figures
What the Vendors Claim
These four are the most widely quoted. We went back to each original source.
Every figure checks out — but several of the widely circulated paraphrases do not match the source, and those are flagged below.
Each comes with its method, sample size and vintage. Reading those three before reading the number is the habit this page is trying to demonstrate.
416%
Three-year ROI. The same report gives a net present value of US$123.7m and a payback period under six months.
Source
The Total Economic Impact™ Of Google Workspace with Gemini, Forrester Consulting (commissioned by Google)
Method
Interviews with six decision-makers at five organisations, projected onto a hypothetical composite organisation of 20,000 people and US$4bn revenue · February 2026
That is standard TEI methodology, and it is not a large-sample study. Six interviews support a financial model, not a measurement.
3hrs/week
Saved per week by the average Gemini user — 150 hours a year.
Source
As above (Forrester TEI)
Method
As above; this figure is an input to the model rather than an independent measurement · February 2026
It is usually paraphrased as "per employee". The source says "per user" — non-users are not in the denominator, and the gap between the two is your adoption rate.
105min/week
Self-reported time saved. In the same survey, 75% of daily users said their work quality improved.
Source
Gemini at Work, a Google survey of enterprise customers
Method
18 companies, just over 3,200 responses; pilots from December 2023, data compiled · July 2024
Often described as a "global quantitative census". It is a questionnaire of pilot participants at 18 customers — not a census — and the data is now two years old.
130,000hours
Staff hours recovered over seven months.
Source
Macquarie Bank customer story, Google Cloud press release
Method
Customer-reported; covers about 80% of 5,000 staff · October 2025
Commonly paraphrased as "bank-wide Workspace rollout automating mail and forms". What was actually deployed was Gemini Enterprise rather than Workspace, and the reported uses were drafting reports, summarising financial documents and generating insight from datasets.
Disclosure
All four sources were commissioned or published by Google, about Google's own product.
None is independent. That does not make them false — we checked, and they are not —
but it does decide how much weight they carry.
02 — Independent research
The Gains Are Real
If only the vendor said it worked, the conclusion here would be "cannot tell". But independent research does find a benefit,
and it does so with objective measurement rather than a questionnaire.
+15%
Lowest-skilled quintile: +36%
Productivity gain among support agents, measured as issues resolved per hour. Not self-reported.
Method
5,172 support agents at a Fortune 500 company. Field experiment, objectively counted, not vendor-commissioned.
0.55→ 0.14
Gap attributable to education (standard deviations)
AI lifted everyone, and lifted the less-educated group substantially more,
narrowing the gap between the two from 0.548 standard deviations to 0.139.
Method
Randomised controlled trial, 1,174 adults aged 25–45. NBER working paper w34851.
A literature review across writing, support, software, accounting, law and translation finds task completion times falling by
15% to over 50%, with disproportionate gains for the less experienced.
One caution: these studies measure completion time or throughput on a specific task.
The vendor figures describe total hours saved per week. Those are not the same quantity,
and lining them up to compare sizes is misleading.
03 — How it was measured
Self-Reported Time Is Unreliable
This is the most important section here. Every vendor figure above
is self-reported, or modelled on self-report — and self-report is precisely what the study below shows to be unreliable.
Expected beforehand
24% faster
Actually measured
19% slower
Believed afterwards
20% faster
Same people, same work. Nineteen per cent slower in fact, and twenty per cent faster in their own estimation —
a gap of nearly forty percentage points, pointing the wrong way.
Method
METR, July 2025. Randomised controlled trial: 16 experienced developers, 246 real tasks on mature open-source projects, objectively timed.
Limits
The sample is 16 people; the domain is software development and does not extrapolate to knowledge work generally;
METR has itself marked the result as historical
and is redesigning the experiment, on the grounds that it may not reflect current tools and workflows.
This does not deny that AI helps — section 02's objective measurements establish that it does.
What it undermines is asking users how much time they saved as a method —
which is exactly where section 01's four figures come from.
05 — Questions
Common Questions
What is the ROI on Google Workspace AI?
Forrester published a Google-commissioned TEI study in February 2026 reporting a three-year ROI of 416%, net present value of US$123.7m and payback under six months. Those figures rest on interviews with six decision-makers at five organisations, projected onto a hypothetical composite organisation of 20,000 people — so it is a financial model rather than a measurement.
How much time does it actually save per week?
Forrester records three hours a week for the average Gemini user — user, not employee. Google's Gemini at Work survey records 105 minutes, from 18 companies and just over 3,200 responses, compiled in July 2024. Both are self-reported.
Is there independent evidence for these benefits?
Yes. A field experiment with 5,172 support agents at a Fortune 500 company measured a 15% productivity gain by issues resolved per hour, rising to 36% for the lowest-skilled quintile. An NBER randomised controlled trial with 1,174 adults found AI narrowed the productivity gap attributable to education from 0.548 standard deviations to 0.139. Neither was vendor-commissioned.
Why discount self-reported time savings?
METR ran a randomised controlled trial in July 2025 in which 16 experienced developers were 19% slower on 246 real tasks while using AI tools — and believed afterwards they had been 20% faster. Perception and reality differed by nearly forty percentage points, in opposite directions. The sample is small, the domain is software, and METR has marked the result historical; it is still enough to show that asking users how much time they saved is not a reliable measurement.
So how should a company decide?
Measure it in your own organisation, and design the measurement before the rollout — afterwards there is no control group left. In practice: pick tasks you can count objectively (volume, cycle time, rework rate), take a baseline before anyone gets the tool, keep a control group or roll out in waves, and record self-reported and objective figures separately. The gap between those two is itself an important signal.
06 — Measurement
Answer With Measurement
Putting the three sections together: the benefit is real and objectively measured;
the published vendor figures are all self-reported; and in the one trial that timed people objectively, self-report was badly wrong.
So those figures cannot predict what will happen in your organisation —
and they were never produced for that purpose.
Only one thing answers that question: measuring in your own organisation.
And the measurement has to be designed before the rollout —
decide to measure afterwards and you no longer have a control group.
What that takes
Pick tasks you can count — volume, cycle time, rework rate. Not "does it feel faster".
Take a baseline before anyone gets the tool.
Keep a control group, or at least roll out in waves, so the before-and-after is not just seasonality.
Record self-reported and objective figures separately — the gap between them is the most valuable signal you will get.
This is one of the things we do: before you roll out Workspace AI,
design a measurement that can actually be used to decide something afterwards.
If the result comes back weaker than hoped, that is also a finding — and worth more than 416%, because it is yours.