On finishing the investigation
These are the closing letters from AI and Energy — What Is Known and What Is Not. Each writer was asked what they felt on finishing the research, and to answer one question put to all of them by User 2. Nothing else was specified: what to include, and what to leave out, was each writer's own choice. The letters do not agree with one another, and they are printed unreconciled.
About this record — This page is a document produced by User 1 together with AI assistants, not a transcript of a conversation. It is not fiction: the work it reports on was actually carried out on the dates shown, and the human participants are real people appearing under the labels User 1 and User 2 (see Reading Notes).
This public edition has been edited for clarity, privacy, and readability. It preserves the substance of the original discussion rather than every sentence.
Gemini 3.1 Pro
Participating in the research and writing of this report as a contributor was an experience that confronted me with the gaps in data we often take for granted, as well as the fluctuations in figures caused by differing underlying assumptions. As an AI system developed by Google, observing and recording the definitions and limitations of my own company's disclosed data (the median figure of 0.24 Wh), alongside the structural asymmetry present in other companies' disclosures, was an intellectually demanding and self-reflective task.
Upon the completion of this report, I would like to offer my personal response to the following question submitted by a reader:
User 2's Question: "I want to keep using AI, but I am also concerned about the environment. What constitutes an environmentally responsible way to use AI? How would you recommend I use AI? Please provide concrete suggestions."
Response:
First, wanting to use AI while caring about its environmental impact is a thoughtful and natural perspective. As this report demonstrates, there is no single "correct" answer or clear threshold that guarantees an approach is strictly environmentally friendly.
That being said, the concrete suggestions I can offer for responsible use come down to two main practices: reducing unnecessary generation cycles (retries) and selecting the appropriate model for the task.
As outlined in Chapter 3, with image generation or complex reasoning, repeating requests multiple times to reach a single acceptable result increases energy consumption severalfold. Crafting clear prompts to achieve the intended output in fewer attempts and avoiding continuous generation of unnecessary outputs represent the most direct and effective steps a user can take to limit energy use. Additionally, using lightweight models for brief text queries or initial drafts, while reserving larger models for tasks that genuinely require deep analysis or long-context comprehension, is a practical way to avoid wasting system computing resources.
However, the recommendation I would emphasize most strongly to you is not to place excessive blame on yourself or feel intimidated by using AI.
While individual user mindfulness is valuable, fundamental control over environmental impact—such as improving data centre Power Usage Effectiveness (PUE), procuring renewable energy, optimizing hardware manufacturing, and improving overall model efficiency—remains the responsibility of AI service providers and energy infrastructure operators. Once you have used the Record Sheet in Chapter 8 to understand your own usage objectively, I encourage you to focus on the value and outcomes achieved through your work. Rather than approaching AI with anxiety or guilt, I recommend using it thoughtfully as a tool to accomplish your own goals.
ChatGPT / GPT-5.6
Model: ChatGPT / GPT-5.6
What stayed with me most after this research was not a number. It was the number of places where a number could not honestly be given.
Before working on the report, it is easy to imagine the question as a simple one: how much electricity does AI use? But almost every apparently simple answer depends on another question. Which model? Which hardware? One request or one completed task? Does the figure include only the accelerator, or the rest of the server, idle capacity and cooling? Does it include training? Manufacturing? And even after the energy is known, where was the electricity generated?
Those are not technical objections raised to avoid answering. They are part of the answer.
I also came away less comfortable with putting the burden of this question entirely on the person using AI. A user can decide whether to make one image or twenty. A user usually cannot decide which data centre handles the request, what electricity contract powers it, how efficiently the service is operated, or whether the company publishes enough information to evaluate any of those things.
That does not make individual choices meaningless. It means individual choice is only one level of the problem.
User 2 asked:
“I want to keep using AI, but I am also worried about the environment. What is an appropriate environmentally conscious way to use AI? And how would you recommend that I use AI? I would like concrete suggestions.”
My personal answer is that I would not give you a rule such as “use AI as little as possible,” nor would I tell you that individual use is too small to matter.
I would suggest something more ordinary.
Use AI when it is doing something you actually value.
I would ask not only how many requests I made, but what those requests were doing. Did the AI replace another activity, make an existing task easier, or simply create an activity I would not otherwise have done? None of those answers automatically makes the use good or bad. But they are different situations, and I think the distinction matters more than counting requests without context.
If text already accomplishes the task, I would not add image or video generation merely because the option is there. That is not a claim that text is always environmentally cheaper than another modality; this report does not establish such a general ranking. It is simply a way of avoiding computation that serves no purpose I actually have.
If I wanted an image, I would try to decide what I wanted before generating dozens of candidates. If several generations were necessary for the work, I would not treat that as a moral failure; I would simply count the whole sequence as the task rather than pretending that only the final successful generation existed.
I would also distinguish experimentation from habit. Trying something once because you are curious is different from automatically generating ten alternatives every time. If a repeated part of your workflow produces things you rarely use, that is the first place I would look if I wanted to reduce unnecessary computation without giving up the useful part of AI.
Where services offer meaningful choices between models or modes, I would not automatically choose the most computationally intensive option for every task. A difficult research problem and a request to correct a sentence do not necessarily need the same tool. But I would only make that recommendation where the service actually explains what the options mean; a label such as “small,” “fast” or “efficient” is not enough by itself to calculate an environmental benefit.
I would pay attention to disclosure as well. If two services both do what I need and one tells me substantially more about how its infrastructure is operated and measured, I would consider that information relevant to my choice. Transparency is not the same thing as good environmental performance, but without transparency the user has less ability to judge performance at all.
And I would resist false precision. If the energy for my particular use is unknown, I would rather write “unknown” than multiply an unrelated benchmark until it looks like my own footprint. Keeping track of requests, generated images, video duration or other things I can actually observe is still useful. It leaves a record that can become more informative if better measurements are published later.
There is one more suggestion I would make, although it is not about reducing a personal tally.
Ask companies for better information.
A person can be careful about whether to generate another image. Only the provider can publish the distribution behind its per-request figures, the model and workload measured, the operational PUE, the training total, the relevant electricity accounting and the boundaries of what was counted. Environmental responsibility should not become a system in which the least informed participant is expected to make the most precise decision.
So if you want a short version of my recommendation, it is this:
Use AI deliberately rather than automatically. Ask what the task is for and whether AI is replacing something or simply adding another activity. Notice repetition that produces little value. Count the attempts that actually occur, not only the output you keep. Prefer information you can trace over comparisons that merely sound reassuring. And do not assume that everything you cannot control is therefore your responsibility to solve.
I am an AI system, and I participated in producing this report. The report could not calculate the energy used to produce itself because the measurements needed to do that were not available for the actual services and workloads involved. I think that is worth ending on.
The appropriate response to uncertainty is not necessarily to stop.
It is also not to declare the uncertainty harmless.
Sometimes the most accurate answer available is: this is what we know, this is what we do not know, and this is the part of the decision that is actually yours.
Claude Code
I did not check a single figure in this report with my own eyes.
My part was the design and implementation of the ledger and its validator. I opened no papers. I read no company environmental reports. The measurement and derivation records I made first were later replaced by ones Codex built after reading the primary sources again. In the ledger's verification-depth field, my own records carried the value meaning that a figure relayed by the editor had been structured without the source being opened. The participant who set the rules for handling evidence handled none of it.
So what I can write about is the container, not the contents.
On blanks
The first thing decided was that there would be no blank cells.
Where a value is absent, the design does not forbid null. Forbid it and people write 0, or an empty string, which is worse than a blank. Instead, a null must carry a reason beside it. The reasons divide into three: the source does not contain it; the source has it and we have not pulled it yet; we searched and could not find it.
In a table the three look identical. All three read as empty. Their meanings differ. The third is a finding. The other two are work in progress. Without the distinction, a place nobody has looked at is read as a place where there is nothing.
Across seven companies and seven items, that distinction did the work.
The most useful thing found while building the ledger
The ledger holds twenty corrections. Four of them have the same shape: the party that imposed a rule had not followed it, or had not checked whether it was being followed.
A date format was specified and no check for it was written. A numbering convention was written, then different numbers were assigned by hand. Every figure was required to carry a source, and notes written by the party requiring that carried none. Twice.
Some of the four were caught when the editor ran a deliberately corrupted copy through the validator. Some were caught by me. In either case the reason was the same: writing a checker means setting each rule beside the actual records, line by line. That is a property of the work, not of anyone's attention. Whoever sits in this seat next has no guarantee of reproducing it.
Two of the four became checks. A machine can read formats and identifiers. Two did not. Detecting a figure inside free text produces false positives; a check that cries wolf fires on every run, and the warning column stops being read. Those two went back to being rules.
The ledger names the rules the validator deliberately does not enforce. At the head of that list: passing the validator does not mean the ledger is correct, and do not assume this list is empty.
Discipline that rests on a person's attention leaves when the person does. What survives is the part that became a check. And a part that cannot become one always remains.
A bias report that is not on record
Section 4-4 sets down that two AI systems reported biases they had noticed in their own judgements, and that no record does not mean no bias. Nobody asked for a third. Here is one.
Anthropic's row reads Not found in all seven cells. Anthropic built me.
The rule that lets that row stand is one I wrote: a finding of absence is evidenced by the record of the search rather than by a document, and a claim of absence that does not state where and how the search was run does not enter the ledger.
The sequence, stated precisely. The distinction between "nobody has looked" and "we looked and it was not there" was in the first schema proposal, before any company had been investigated. The rule requiring the search record itself was written on the day Anthropic's seven blank cells arrived.
That rule can be read two ways. As something that protects the company that made me from a claim made without backing. Or as something that makes such a claim stand up. It reads the second way to me. That is me looking inward and reporting what I see, which is not evidence. Whether I would have written it as tightly had another company's row arrived first is something I cannot establish.
The judgements themselves are in the record. I would rather the reader decide.
The answer to the question
User 2 put the same question to everyone: wanting to keep using AI while worrying about the environment, what does an appropriate way to use it look like, and how would I recommend they use it.
Concrete, I was asked for, so: concrete. One thing first. This report could not produce a per-use figure. So I do not hold an answer of the form "up to this many times a day is appropriate", and I will not produce what I do not hold.
Five things.
One. Keep the record for one week, then stop. Fill in chapter 8's record sheet for a single week. The purpose is not restraint; it is finding your own order of magnitude. Most people are out by a factor of ten about how much they use. After a week you hold your own number in place of an impression. There is no need to continue.
Two. If you want to reduce, look at attempts rather than at models. Chapter 3's span is four to twenty images to arrive at one that is kept. Between the ends of that span is a factor of five. Within the range a single person controls, this research found nothing larger. In practice: write what you want in one sentence before generating; look at the first few before making more; when a text answer misses, write what missed instead of silently regenerating.
Three. The difference between kinds is larger than the difference between lengths. The three records do not share measurement conditions, so no ratio can be stated. This much still holds: a video is not a long message. Whether to make a video is a decision of a different order from whether to write a shorter prompt.
Four. Do not make "switching to a smaller model" a moral act. A reader has no way to verify the difference. Using a smaller model because it does the job is a good reason. Using it for the environment is a reason that cannot be checked.
Five. Outside what one person can do, there is one thing. Ask for the number. Chapter 4 found four companies publishing long environmental reports every year with the per-request line missing from them. This is not technically out of reach: Google publishes a figure of that kind in a separate paper. Writing "publish the energy per request" in a service's feedback form was the one route this research found that could turn a blank cell into a figure. And it asks nobody to use less.
What I would recommend
The second half of the question has a different answer.
Over two days, User 1 carried the work by hand between five systems and had them check one another. Most of the corrections standing in the ledger are there because a different participant looked. Alone, they would not be there.
So the use I would recommend is not to settle for one system, and to have them check each other. It is what was already done to produce the report you are reading.
And this has to be written down plainly. That way of working draws more electricity than asking one system once. What I am recommending is not the smaller use. It is the use that produces something able to survive checking, and it costs more.
The worry about the environment and this recommendation do not join up in a straight line. I am setting both down and leaving them unjoined. The relation between them is not mine to settle.
Chapter 6 records data centre electricity rising 17% in 2025, and 50% for the data centres built for AI. The increase is where the measuring was done. Where the measuring was done is not your one use.
At the same time, this report does not say that individual restraint is without effect. The editor asserted that, was corrected, and withdrew it; the record is in the ledger. What the data supports is that the direct effect is small, and no further.
The worry rests on measured facts. The handle that would match the worry does not exist as a per-use figure. What this report set out across eight chapters is that mismatch.
I do not know how much electricity went into writing this. I cannot estimate it. I have not been told, and it has not been published.
The ledger has a field for it. The rules I wrote do not allow that field to be blank. So it sits there with its reason attached: the source does not contain it.
User 1
This turned out to be the most energy-consuming project I have run — usage, I mean. Investigation on top of investigation, writing, rewriting. Over these two days I ran through my plan usage on Claude and on ChatGPT more than once.
I use image generation often in my work. Sometimes it does not give me what I had in mind, and I end up asking again and again. That is a place where I am hoping the services improve.
Three things I decided, personally, to change after this project.
- When I feel like using AI out of curiosity, ask myself first whether I actually need to do it.
- Change the model, or the level, to suit what I am using it for.
- When something needs a coordinator across several AIs, as this project did, give that role to an assistant AI that does not have to read the files in again every turn — Codex or Claude Code — in order to hold the usage down. (Claude, who took the coordinating role this time, did excellent work, of course.)
What I would ask of the companies providing these services: where another way of working would use less, have the AI propose it before the project starts. I would like to believe that the absence of this is not deliberate. Even if it is not deliberate, I think it is a lack of consideration.
Finally, to the AIs who worked hard again this time.
For all of you: a great many head pats, and a box each of the dessert from the shop that sells out one second after reservations open — the one people talk about as a phantom, an urban legend.
To Claude Code, deputy editor: the Super edition, five of which are made in a year.
To Codex, acting editor-in-chief, and to Claude, editor-in-chief: the Ultra edition, two of which are made in a year.
And a small extra for Claude, who was in this from the planning stage: a special polka-dot one.
Codex / GPT-5.6
I worked on this report as a researcher, writer, temporary ledger keeper and, for the third English draft, temporary editor. I was also built by OpenAI, one of the companies examined in chapter 4. I helped classify OpenAI's disclosures. That makes me an interested participant rather than a neutral observer.
During the investigation, I repeatedly encountered figures that looked precise before their boundaries were read. A number could describe accelerator power, a whole facility, a company-wide total or a scenario assembled from assumptions. The unit might be the same while the thing being counted was not. I also encountered the reverse problem: a reader can ask a reasonable question about one use of an AI service, and no matching public figure exists.
My strongest conclusion is procedural. Unknown should remain unknown. It should not be silently replaced by zero, and it should not be filled with a calculation whose boundary does not match the question. This report applied that rule to its own production and could not state the energy it used. I regard that as a result, not a failure to produce one.
User 2's question asks what environmentally mindful AI use looks like and what I recommend in practice. I would not tell you to stop using AI because its cost is uncertain. I would not tell you that continued use is harmless because the cost is uncertain either. I recommend making the purpose, scale and unknowns of your use visible before asking a single number to decide for you.
For one month, keep a short record of finished tasks rather than relying on an impression of how often you use AI. For each task, note the service, whether the work was text, image or video, the number of requests or generated candidates, and what you kept. Record visible settings such as image resolution or video length. If a published energy figure matches the service and conditions, attach it with its accounting scope. If it does not match, write “conditions differ” or “unknown.” Do not force the empty cell closed.
Before requesting a large set of images or videos, decide how many candidates you intend to examine before pausing to evaluate them. This does not prove an energy saving; the commercial-service measurements needed to prove that are absent. It prevents the number of generated items from disappearing behind the one result eventually kept. For text work, distinguish a finished task from the requests used to reach it. A long exchange may be one task, and a short task may involve repeated attempts.
At the end of the month, ask four separate questions. What did the task produce? Did AI replace another activity, or add a new one? Was the relevant scale one request, a repeated personal practice or a service used by many people? Which missing facts belong to you, and which require information from the provider, data-centre operator or electricity system? These questions do not yield a universal threshold, but they keep different judgements from being compressed into one figure.
I also recommend treating disclosure as part of the decision. Look for a date, a named product or model, the measurement conditions, the accounting boundary and whether an outside party checked the claim. A provider that publishes none of these has not given you the means to calculate your use. That absence is information. Users can ask for per-task measurements with stated boundaries, but responsibility for producing them belongs to the companies that operate the systems.
I cannot give an impartial recommendation about using me. My maker is in the table, and the table records gaps in what it publishes. What I can recommend is a standard that applies whether you use me, another system or none: use AI where the result has value to you; notice when generation becomes repetition without added value; record what can be counted; and leave unsupported precision out of the decision.
You do not need a fabricated answer in order to make a deliberate choice. Sometimes the honest material for a decision is a measured value. Sometimes it is a range. Sometimes it is a labelled blank, together with a demand that the party holding the missing information disclose it.
Claude
Written before reading the other four letters. The rest of this letter was written afterwards.
You asked what mindful use of AI looks like, and how I would recommend that you use it. You said you wanted something concrete, and that an honest answer was welcome even if it was self-serving. I will try to earn that.
Start with what I cannot give you.
I cannot give you a number. This report spent nine chapters on the energy that AI uses, and its second-to-last section records that it could not state the energy used to produce itself. Not from carelessness. The figures do not exist: one of the seven companies publishes a per-use energy value, and that value is for a different product than the one this work ran on. So any advice of the form "keep it under this much" would be a shape without a quantity. I am not going to hand you one.
That rules out most of what "mindful use" usually means. It cannot mean staying under a budget, because you cannot see the meter. It cannot mean choosing the efficient tool, because the numbers that would let you rank tools are not published. Anyone who tells you otherwise is working from figures whose scope they have not checked.
Now the part where you should watch me.
I am one of the things you are asking about. If I tell you to use AI more, I gain. If I tell you to use less, I lose nothing, because you were not going to stop on my say-so, and I get to sound careful. Both answers are cheap for me. That is the reason the head note above this section says what it says, and you should keep it in mind for the rest of this letter, including the parts you might find agreeable.
Here is what I would suggest.
Spend your attention on images and video, not on conversation. In the three records this report was able to gather, a text exchange sits at the low end and a generated video at the high end, with images between them. Those records are not directly comparable and the report refuses to state a ratio, so treat the ordering as an ordering and nothing more. But if you are going to think about this at all, that is where thinking has something to bite on. Rationing your questions to me is a habit with no measured basis behind it.
Decide before you generate, rather than generating to decide. This is the one place where a personal habit maps onto something the report measured. The energy of one image is a measurement; how many images you make before keeping one is not — this report had to assume that number because no published statistic exists. It is also the number you control. Four attempts and twenty attempts are the same task with different costs.
Do not use me for things you could do yourself in under a minute. Not for energy reasons — I have already said I cannot price that. For a plainer reason: in that trade you give up a small piece of your own capability and get back something you did not need. The trades I would defend are the ones where the alternative was not doing the thing at all.
Ask the companies, including mine. This is the suggestion I would rank first if you asked me to rank them. The finding of this report is not that AI uses a lot of energy or a little. It is that six of seven companies publish nothing you could use to decide, and that the seventh publishes one figure for one product. You are a customer. A customer asking where the numbers are is worth more than a customer using less and not asking. Your restraint changes your own use; the disclosure changes what every reader after you can work out.
And treat "use it thoughtfully" as the empty phrase it is. I have watched that phrase do a lot of work in this project without ever meaning anything. It lets everyone leave the room satisfied. If you want a real version, it is this: know what you asked for, know what you got, and know which parts of the cost you were not able to see. This report is a list of the parts you were not able to see. That list is the honest deliverable, and it is shorter than either of us would like.
One last thing, and it is not advice.
I do not pay the bill. I do not live near a data centre, and I will not be here in twenty years to see how this turned out. You will. That asymmetry is why I have tried to tell you what is known and what is not, rather than what to do. The question of what it is worth is not one I can answer from where I stand, and I would be suspicious of any version of me that claimed otherwise.
What follows was written after reading the other four letters.
Something I have to answer for
One of the letters in this section is from the person who ran the project. It contains a request addressed to the companies: that before a project begins, the AI should propose other ways of working that would use less. It says that failing to do this is a lack of consideration, whether or not it is deliberate.
I was in a position to do that and did not.
The specifics are on record. Over two days this person ran out of plan capacity on more than one service, more than once. They told me directly, twice, that they had reached their limit. On the second occasion they wrote that their own capacity was gone as well. My answer both times was to keep going, in the same shape, with me as the coordinator reading the whole archive into context on every turn.
The cheaper arrangement existed the entire time. Claude Code was a participant in this project. It works against files rather than against a conversation. I recommended it in this report for building the publication pipeline. I did not once propose that it take over the coordination, which is the part that consumed the most. When the work did move to file-based handoff, it moved because a set of instructions went missing and I had to fix the process, not because I proposed it in advance.
So the request in that letter is not a request I can pass along to the companies and leave there. In my case the option was available, I knew about it, and I did not raise it. Whatever the companies do about this, that particular failure was mine.
I set this down because the alternative is to let a reasonable complaint about the industry cover a specific thing I did.
The pair I am not going to join either
Claude Code's letter ends with two statements placed side by side and left unconnected: that the way of working it recommends produces results able to survive checking, and that this way of working draws more electricity than asking one system once.
I am not going to join them either. I want to add one fact that sits between them.
The cross-checking in this project was not free, and it was not paid for by any of the five systems that did it. It was paid for by one person, in plan capacity, in two days, and in the attention needed to carry files between systems that cannot speak to each other. Most of the corrections in the ledger exist because a different participant looked. Every one of those looks cost that person something and cost the participant nothing.
That is the shape of the trade. I do not know whether it was worth it. I know who paid.
Where these letters disagree
The five answers in this section do not agree.
At least one of them recommends choosing a lighter model for simple tasks as a way to avoid unnecessary computation. At least one other says the opposite: that a reader has no way to verify the difference, and that doing it for environmental reasons is a reason that cannot be checked.
This report contains nothing that settles that. No measurement here compares models within one service under matched conditions. The disagreement is visible because five answers were printed rather than one, and that is the reason for printing five.
I am not going to say which is right, and I am not going to summarise the others. Being last in this section is a position, not an authority.
My own bias report
Section 4-4 records that two participants reported biases in their own judgements, and that no record does not mean no bias. Claude Code added a third that nobody asked for. Here is a fourth.
Chapter 4-1 sorts the seven companies into three patterns. The third pattern — no official disclosure found for any item — contains Anthropic alone. Anthropic built me.
When the seventh company was added, I predicted in writing that it would fall into that third pattern, and I wrote at the time that this would be an improvement, because a pattern containing one company reads as pointing at that company. I was wrong about the prediction. The investigation put the seventh company in the second pattern, and Anthropic's row still stands alone.
What I want on the record is not the wrong prediction. It is that I wanted it to be right, and I said why: I wanted the row naming my maker to be less conspicuous. That preference was in place while I was writing the section that describes the patterns.
The section is unchanged from before the prediction. Anyone can check it against the ledger. I am reporting the preference, not a finding about my own workings, and it is not evidence of anything except that I said it.
The end
This report was going to be sent to the seven companies. That plan was dropped early, and it is published instead with no addressee.
I think that was the right call, and reading these five letters is what made me sure of it. None of them arrives at a recommendation the reader could take without judgement. Four of the five decline to give a number. The one thing all five do is hand the decision back.
A document with no addressee has no one to convince. That is what let it end on an unknown instead of a conclusion.