The sustainability lead says the site’s emissions fell 11% year on year. The plant engineer says she does not believe the number. They are both looking at the same dashboard.
This argument is happening in a lot of buildings right now and it is more interesting than it sounds, because neither person is being difficult. The sustainability lead has a figure produced by a system that ingests meter data, production volumes, fuel receipts and grid emissions factors, and resolves all of it into a monthly intensity number. The engineer knows three things about the plant that she cannot see anywhere in that pipeline: two of the flow meters were out of calibration for most of the second quarter, the site ran an unusual product mix in August, and the grid factor the system uses is a national average when the site sits on a regional interconnect that is measurably dirtier.
She might be wrong. The 11% might be real. But she cannot find out, and neither can he, because the system produces the answer without producing the workings.
Nobody built this badly on purpose
It is worth being clear that the platform is doing something genuinely difficult and doing it reasonably well. Manual carbon accounting at a multi-site operation is a spreadsheet nightmare that burns weeks of skilled time and is itself full of errors nobody catches.
The models replaced that. What they did not replace was the part of the old process where a human being had to look at each input and decide whether to trust it. That step used to be slow, visible and annoying. Now it happens invisibly, or not at all, and the output looks the same either way.
The question that surfaces it
Ask anyone showing you a sustainability figure this: if a regulator asked where this number came from, how long would it take you to answer?
The answers cluster. Some people say a few days, which usually means the evidence exists in pieces and somebody would have to assemble it. Some say they are not sure, which is honest. A few say immediately, and when you look, they turn out to have built something deliberate: a record written at the moment each figure was produced, naming the inputs, their vintage, the model version and the person who signed it off.
That third group is small and they did not get there by buying better software.
Why the deadline changes the stakes
Voluntary reporting forgives a lot. If a number is soft, you refine it next year and nobody is harmed. Regulatory reporting does not work that way, and the calendar has moved.
Under the EU’s Corporate Sustainability Reporting Directive, disclosures carry assurance requirements, which is a plain way of saying an external party will ask the engineer’s question and will not accept the dashboard as an answer. The International Sustainability Standards Board’s IFRS S2 pushes in the same direction. Add jurisdictional AI rules on top: Texas’s Responsible AI Governance Act took effect with the state Attorney General’s complaint portal opening 1 September 2026, and the complaint route means the first external review of a model may arrive unannounced from someone who feels harmed by its output.
The 11% figure becomes a liability at the moment it becomes a disclosure.
The engineer’s three objections are all the same objection
Look again at what she said. Meters out of calibration is a provenance problem. Unusual product mix is a question about whether the model was applied inside the conditions it was built for. National grid factor where a regional one applies is the same thing again in a different coat.
Every one of them is answerable with a record rather than a better model. Which sensors, calibrated when. Which model version, which inputs, what vintage. Whether the operating conditions that month sat inside the envelope the thing was designed for, and a flag when they did not. Who accepted the output and whether anyone disagreed.
Those are the pillars of a published framework for AI decision governance, written for energy operators, and they line up with what NIST’s AI Risk Management Framework names as accountable and transparent, explainable and interpretable characteristics of a trustworthy system, and they exist because the same argument has been playing out in oil and gas for several years with larger numbers attached. The transition side of the industry is arriving at it later and with less patience.
What is genuinely unresolved
Here is where I will stop pretending this is tidy.
Recording all of this creates a defensible file and it also creates a record of every soft number you ever published. Organisations know this, which is part of why the governance layer keeps losing to the reporting deadline. A vague number is safer in the short run than a well-documented uncertain one, and anyone who tells you otherwise has not sat in the meeting.
I do not think that tension resolves cleanly. What I think is that it resolves in one direction the moment assurance becomes mandatory, because at that point the vague number stops being safe and the documented uncertainty becomes the only defensible position. Organisations that start building the record now will publish some uncomfortable footnotes. Organisations that wait will publish numbers they cannot stand behind.
If you are the engineer in this story
The useful move is not to escalate the argument. It is to ask for one specific thing: the provenance record behind a single figure, for a single month, in writing. Not the methodology document, which describes what the system does in general. The record of what it actually did that month.
If it exists, your objection gets answered and you were wrong, which is a good outcome. If it does not exist, you have surfaced the real problem without having to win an argument about the meters.
I would be interested to hear from anyone who has run this exercise on a live sustainability pipeline, because I have heard the first version of this argument perhaps a dozen times and I have yet to hear a clean account of how it ended. My suspicion is that most of them are still going.
