How much water does ChatGPT use?
There are two public figures and they are roughly a hundredfold apart. OpenAI's CEO has put an average ChatGPT query at about 0.000085 gallons — around a third of a millilitre, or a fifteenth of a teaspoon — alongside 0.34 watt-hours of energy. Peer-reviewed work on GPT-3 estimated a 500 ml bottle per 10–50 medium-length responses, i.e. 10–50 ml each. Both are defensible; they measure different things, in different years, for different models.
- Vendor figure
- ~0.32 ml per average query
- Academic figure
- 10–50 ml per response, GPT-3
- Gap
- ~30–150x between the two
- Energy per query
- 0.34 Wh vendor figure
- GPT-3 training
- 5.4M litres per Li et al.
- Vendor methodology
- Not published figure given without one
- Independently audited
- Neither no third-party verification
- Verified
- 2026-08-21 both sources read directly
We are not going to pick one number and pretend it settles this.
The vendor figure. In a June 2025 essay on his personal blog, OpenAI's CEO Sam Altman gave a per-query cost in a parenthesis: about 0.34 watt-hours of energy, which he compared to a high-efficiency lightbulb running for a couple of minutes, and about 0.000085 gallons of water — roughly a fifteenth of a teaspoon, or about 0.32 ml. We read that post directly rather than a summary of it.
The academic figure. Li et al., Making AI Less "Thirsty" (arXiv 2304.03271, later published in Communications of the ACM), estimated that GPT-3 consumes a 500 ml bottle of water per roughly 10–50 medium-length responses, depending on when and where it is deployed, and put training GPT-3 in Microsoft's US datacentres at 5.4 million litres including 700,000 litres of on-site consumption.
Why they can both be right. Three differences do most of the work. Model and year — GPT-3 in 2023 against whatever OpenAI was serving in 2025, across several generations of efficiency gains. Place and time — the paper's central finding is that water efficiency varies enormously by datacentre and season, which is why its own answer is a five-fold range rather than a number. Scope — the paper counts both on-site cooling and the water consumed generating the electricity; a figure counting only the first would be far smaller.
And the thing neither side says. The vendor figure was published without a methodology: no scope definition, no measurement period, no breakdown, in a blog post rather than a sustainability report. That does not make it wrong. It does mean nobody outside OpenAI can check it, and it is the reason we present a range on this page instead of the tidy number this query plainly wants.
Making the number mean something.
It does not drink it. The water goes into cooling the datacentres that run the servers, mostly by evaporation — which is why it is consumed rather than merely used, and why it does not come back into the local supply the way domestic water largely does.
There is a second, larger and less visible flow: the water used to generate the electricity in the first place, at thermal and hydro power stations. That one is invisible on a datacentre's own meter, which is precisely why the scope of any figure decides its size. A number counting only on-site cooling and a number counting both are not two opinions about the same quantity.
A third of a millilitre sounds trivially small; 50 ml sounds alarming; both are useless for deciding anything, because nobody sends one query. The quantity that matters is the aggregate — hundreds of millions of users, many queries each, every day — and that depends on figures neither source publishes.
It is also the wrong comparison for an individual. If your concern is your own footprint, the number that dwarfs everything here is not your chat history; it is flights, heating, diet and car use, by orders of magnitude. Choosing not to use ChatGPT is close to a rounding error in a personal water footprint, whichever of the two estimates you believe.
The most actionable finding in the academic work is not the headline figure at all: water efficiency varies by location and by season, sometimes dramatically, so where and when a model runs changes its water cost far more than shaving a few queries does. A datacentre in a cool, water-rich region behaves nothing like one running evaporative cooling through a hot summer in a stressed basin.
That reframes the question usefully. "How much water does ChatGPT use" is far less answerable, and far less useful, than "how much water do datacentres consume in this watershed, and who decided that" — which is a question about siting and local policy, and one that gets answered by regulators rather than by chatbots.
To state a figure with confidence, somebody outside the company would need per-datacentre water-use effectiveness, the workload mix, the actual query volume, and an agreed scope covering both cooling and generation. None of that is published for ChatGPT specifically.
That is a criticism of the available information rather than of anyone's arithmetic, and it is why this page ends with a range and a caveat instead of a headline. We measure prices on this site because we can read them off a vendor's page; we do not measure datacentre water, and we are not going to imply otherwise. Our methodology page covers where that line sits.
Things that actually move the number.
Three levers, ordered by how much difference they make. We have measured none of this ourselves — these follow from the published research rather than from our own testing.
Frequently asked.
Quick follow-ups people search after this question.