WEBVTT

00:00:00.000 --> 00:00:11.000
And in this session, we have four interesting talks so Without much of an explanation, I'll invite Zach to discuss the great stuff that's happening in Atlas in terms of environmental impact and sustainability for computing. I'll give you a 15 minute timer and

00:00:11.000 --> 00:00:17.000
It will buzz with a very funny sound so that you know when your time is up and then we have time for questions.

00:00:17.000 --> 00:00:19.000
So yes, whenever you're ready, go ahead.

00:00:19.000 --> 00:00:27.000
Thanks very much, Sigonia. And I will do my best to run through our Atlas Computing Sustainability Studies here.

00:00:27.000 --> 00:00:35.000
On the top here, you can see a distribution of the computing that we run worldwide. It really is a global effort.

00:00:35.000 --> 00:00:40.000
And I'll tell you about some of the reasons that that's important and interesting in the talk.

00:00:40.000 --> 00:00:53.000
So for scale, we operate something like 700,000 cores of compute worldwide, and you see peaks up to a million cores when we take over big chunks of high performance computing systems.

00:00:53.000 --> 00:01:00.000
And we have something like an exabyte of data between disk and tape at the moment.

00:01:00.000 --> 00:01:08.000
That's a combination of grid sites that we talked about earlier in the session and high performance computers as used to be called supercomputers.

00:01:08.000 --> 00:01:20.000
Cloud computing and some volunteer computing. And again, those peaks are mostly driven by those high performance computing centers where we're users, but certainly not the reason for existence of the data centers.

00:01:20.000 --> 00:01:25.000
When we look forward to the future to the HLHC, those resources are going to grow.

00:01:25.000 --> 00:01:30.000
So we know by the time the HLHC starts up in sort of 2030ish now.

00:01:30.000 --> 00:01:36.000
We're going to need a factor of three to five more compute, disk, and tape.

00:01:36.000 --> 00:01:45.000
And by the end of data taking in 2041, the end of the collaboration, we'll need another factor of two. So we're talking about an order of magnitude growth.

00:01:45.000 --> 00:01:55.000
Which means we have a nice lever arm to adjust if we can improve things, we can pull down the projected carbon cost of those resources.

00:01:55.000 --> 00:02:01.000
In general, we make resource requests 12 to 18 months in advance. So in year.

00:02:01.000 --> 00:02:05.000
Jevon's paradox applies where you make something faster and we just do more of it.

00:02:05.000 --> 00:02:16.000
Over the course of one to two to three years, there is opportunity for reduction of resources through optimization. And we have seen that where we've made a workflow faster.

00:02:16.000 --> 00:02:26.000
We've made some data smaller and therefore two years out, three years out, we've requested fewer resources and had some savings there.

00:02:26.000 --> 00:02:34.000
So obviously, if we can pull down these lever arms, then we'll reduce the carbon cost of the experiment. But there are a lot of other things that we can do as well.

00:02:34.000 --> 00:02:38.000
So one simple thing is to just build awareness in the collaboration of what's going on.

00:02:38.000 --> 00:02:44.000
Every user, when they run a job on the distributed computing system gets an email like this when it's done.

00:02:44.000 --> 00:02:53.000
And you notice down here there's an estimated carbon footprint for the task section, which gives them a link to what it means and how it's calculated and so on.

00:02:53.000 --> 00:03:09.000
This includes now an average carbon footprint averaged worldwide for the tasks that they've run. And we keep internally the information about the actual carbon footprint for that task in that location.

00:03:09.000 --> 00:03:19.000
For the grid carbon intensity there. We provide a worldwide average for a few reasons. One is there are a lot of data centers. We don't control them. We don't have all the information about all of them. We don't know where there are solar panels yet.

00:03:19.000 --> 00:03:26.000
So we're trying to get that information until we know we we shouldn't uh charge sites extra.

00:03:26.000 --> 00:03:31.000
Second thing is CPU doesn't sit idle. So if a user moves a job from Poland to Norway.

00:03:31.000 --> 00:03:38.000
We'll put a different job on the site in Poland. We won't just let the CPU sit idle. So moving jobs around doesn't actually save carbon.

00:03:38.000 --> 00:03:43.000
The third reason, and this is something that you'll hear me say a few times.

00:03:43.000 --> 00:03:49.000
If all the users take action and everyone pushes their work to Norway.

00:03:49.000 --> 00:03:55.000
The Norwegian data center will get absolutely hammered. The disks will burn out faster.

00:03:55.000 --> 00:04:05.000
The network will be hammered. And the jobs will become less efficient and we'll still be running the data centers elsewhere. So actually.

00:04:05.000 --> 00:04:12.000
Recommending that they all do that would cost us in carbon terms. It would be a negative.

00:04:12.000 --> 00:04:23.000
Solution. So there are a lot of cases like this where if you sort of take too myopic a view you can make a recommendation that's harmful.

00:04:23.000 --> 00:04:31.000
Globally. The last reason is it's a nice correlation. You want to say if somebody makes their code faster that they have a lower carbon footprint.

00:04:31.000 --> 00:04:34.000
And if you do a global average, that correlation is stronger.

00:04:34.000 --> 00:04:45.000
It's, of course, not intended to shame anybody. We just want to make them aware of what's going on. And we also do this now for production. So you can see here's a production that costs two tons of carbon.

00:04:45.000 --> 00:04:52.000
About 130 of which was failed jobs. That's pretty typical. 5%, 6% failures is about right.

00:04:52.000 --> 00:05:02.000
For our grid productions. We also have started to build awareness early on in tutorials. So there's this Atlas analysis tutorial.

00:05:02.000 --> 00:05:06.000
And we teach folks how to run a failing job and then diagnose what happened.

00:05:06.000 --> 00:05:10.000
And then to look at the carbon impact of that failing job.

00:05:10.000 --> 00:05:19.000
We find people are not familiar with how this all works. And we found cases in the wild where users have retried jobs hundreds of times.

00:05:19.000 --> 00:05:26.000
And they just keep failing around the world. Hundreds of times. So this seems like an education issue.

00:05:26.000 --> 00:05:31.000
Now we just need to make people more aware of how to fix things and what goes wrong.

00:05:31.000 --> 00:05:37.000
Then there are a bunch of things that we can do in terms of policies that impact our carbon footprint.

00:05:37.000 --> 00:05:43.000
I really like these because building a new data center, buying new hardware takes years.

00:05:43.000 --> 00:05:51.000
Takes years to replace old hardware. If you change your policy, if you change what you store, that's almost instant in terms of its carbon impact.

00:05:51.000 --> 00:05:54.000
So these are really nice things to be able to get after.

00:05:54.000 --> 00:06:04.000
Mostly historically, they've been thought of in financial terms. For things like the data carousel, which is a system by which we use tape more than disk.

00:06:04.000 --> 00:06:11.000
And move data actively on and off of tape. The idea was we'd save cash because tape is cheaper than disk.

00:06:11.000 --> 00:06:16.000
We also save carbon because in carbon terms, tape is cheaper than disc.

00:06:16.000 --> 00:06:20.000
So this is a case where there's a nice synergy. Everything's pointing in the same direction.

00:06:20.000 --> 00:06:26.000
Obviously, there are other places where they point in opposite directions and we have to justify things somehow.

00:06:26.000 --> 00:06:32.000
Another nice example is data reproduction. So if a user says, I want to save this data just in case I need it.

00:06:32.000 --> 00:06:43.000
Do I save it or do I delete it? And our most recent math says if they request it less than about 15% of the time.

00:06:43.000 --> 00:06:49.000
We should just delete it and reproduce it on demand. And the request rate is about 25% right now.

00:06:49.000 --> 00:06:59.000
Give or take. So right now we're better off keeping it That balance could change in the future such that it's actually better for us to delete the data and then reproduce it if they ask.

00:06:59.000 --> 00:07:08.000
The last obvious policy is to use what we have. So we talked earlier today about trigger farms for several experiments.

00:07:08.000 --> 00:07:15.000
Those got built. We paid the embodied carbon. We paid the carbon for the Tier Zero Center.

00:07:15.000 --> 00:07:19.000
Those are all for operation of the LHC. The LHC doesn't operate year round.

00:07:19.000 --> 00:07:28.000
So when it's not on, we need to use that compute as best we can to make sure that we amortize the embodied carbon over more compute.

00:07:28.000 --> 00:07:37.000
A few other things we're looking into are things like waste lost and unused data, right? It's okay, in my opinion, if we use carbon.

00:07:37.000 --> 00:07:48.000
For science. It's not okay if we waste carbon while doing science So this was a look at when we produce data that no one ever looks at.

00:07:48.000 --> 00:07:53.000
And you see like a year ago some data that were produced that were never looked at.

00:07:53.000 --> 00:08:00.000
Most of these things turn out to be two inclusive patterns somebody asked for all the Monte Carlo. They don't really need all the Monte Carlo.

00:08:00.000 --> 00:08:05.000
Or they ask for a big production, check a few files, find a bug.

00:08:05.000 --> 00:08:09.000
And then the rest of the production just goes ahead and they don't kill it off fast enough.

00:08:09.000 --> 00:08:13.000
So these sorts of things cost us and we're trying to get better.

00:08:13.000 --> 00:08:17.000
We're, of course, trying to improve our CPU over wall efficiency.

00:08:17.000 --> 00:08:25.000
You can see the trend over the last few years. It's upward. It's slow, but it's upward. It's hard to get higher than 90% efficiency in a lot of these cases.

00:08:25.000 --> 00:08:32.000
And there's a constant effort to sort of reduce the serial portions of many core multi-core jobs as well.

00:08:32.000 --> 00:08:45.000
We also have automated systems to take sites offline when they have problems. So there's a significant potential savings when a site that starts failing can be taken offline and we don't have lots of failing jobs there.

00:08:45.000 --> 00:08:52.000
We're also preparing to release a significant amount of our event generation output as open data.

00:08:52.000 --> 00:09:00.000
So this would be a sort of service to the community. It's good scientifically. That's great. It's hard because we have to document everything that we've produced publicly.

00:09:00.000 --> 00:09:04.000
Okay, that documentation is good for us, so maybe that's okay.

00:09:04.000 --> 00:09:08.000
We're going to release something like 10 million CPU hours of event generation.

00:09:08.000 --> 00:09:14.000
And now we have to ask a really interesting question that we'll have to come back to in about a year to see what the answer is.

00:09:14.000 --> 00:09:23.000
One is, does the community pick it up at all? Do they actually use it? A second is, does it actually reduce the amount of event generation that is run in the community?

00:09:23.000 --> 00:09:29.000
But it's not obvious that people will necessarily use ours and not generate their own or won't generate more.

00:09:29.000 --> 00:09:38.000
End users. And the third is, does it actually result in a carbon savings? Or in reality, do people use the same or even more carbon?

00:09:38.000 --> 00:09:49.000
Because they're able to run over larger numbers of events. So we'll have to see what in carbon terms this this does. It might be different from the answer to what scientifically it does.

00:09:49.000 --> 00:09:53.000
So hopefully I'll come back in a year and tell you more.

00:09:53.000 --> 00:10:03.000
We've also looked into job failures quite a bit. We found out it's very important to de-correlate issues to avoid victim blaming, right? So we have this tier zero center at CERN.

00:10:03.000 --> 00:10:09.000
It's very efficient. The failures are quite rare. That's because it runs very well controlled jobs.

00:10:09.000 --> 00:10:15.000
We have other data centers that have kind of unique resources where there are very high failure rates.

00:10:15.000 --> 00:10:24.000
And we expect fairly high failure rates. So we've tried to build this sort of expected failure rate based on worldwide failures.

00:10:24.000 --> 00:10:29.000
And then compare sites based on their observed and expected failure rates.

00:10:29.000 --> 00:10:42.000
And then look for cases where the observed rate is way different from the expectation, either above or or below. And that's where we want to dig in to see what's going on, not just because there's a high failure rate, but because there's a higher than expected.

00:10:42.000 --> 00:10:53.000
Failure rate. So we're starting to get into this. And also starting to get into automatic retry policies. So a job runs on the grid, it fails.

00:10:53.000 --> 00:10:58.000
And we have an automated system by which we say, this failed. Let's try it again somewhere else.

00:10:58.000 --> 00:11:04.000
Or mothers failed. It was probably transient. Let's just go again and it might succeed next time.

00:11:04.000 --> 00:11:12.000
And obviously, if you get that wrong, you end up retrying a bunch of jobs that are just going to fail over and over again. And you should stop.

00:11:12.000 --> 00:11:24.000
So we need to constantly re-examine these sorts of policies. We've talked a lot in this workshop about software development for sustainability. It goes without saying we should continue to improve our software.

00:11:24.000 --> 00:11:29.000
One thing that's worth pointing out is it's possible to run CPUs at lower frequency.

00:11:29.000 --> 00:11:43.000
To improve their performance per watt. So this is a nice little pot of The frequency of the CPU, against the performance per watt. So actually, if you lower the frequency of the CPU, you get better performance per watt.

00:11:43.000 --> 00:11:46.000
If you care only about power, that's clearly a good idea.

00:11:46.000 --> 00:11:50.000
If you have to deliver a certain number of hep scores, say.

00:11:50.000 --> 00:12:01.000
Then you might have to buy more CPUs in order to provide that much power. So you have to balance embodied and operational carbon in order to get that working point, right?

00:12:01.000 --> 00:12:11.000
We've also been working on checkpointing. That's the ability to stop a job right at state to disk and then read it off of disk and start exactly where you left off.

00:12:11.000 --> 00:12:16.000
To see if we can do a better job during outages, brownouts.

00:12:16.000 --> 00:12:30.000
Regular downtime, that sort of thing. In the future, we're also thinking about this sort of situation where this is a projection to 2035 of of German power. And there are periods where there's a huge amount of renewable energy.

00:12:30.000 --> 00:12:35.000
So if we can take a data center, run it when there's a ton of renewables.

00:12:35.000 --> 00:12:41.000
And then basically turn it off. When the renewables are gone at night.

00:12:41.000 --> 00:12:46.000
And then turn it back on when the load is lower than the demand.

00:12:46.000 --> 00:12:56.000
Renewables are higher than the demand. We might be able to actually overall save quite a lot of carbon and maybe even make our computing largely more neutral.

00:12:56.000 --> 00:13:06.000
I mentioned validation that we need to improve the production of buggy samples and avoid that as much as possible. That's another ongoing effort.

00:13:06.000 --> 00:13:18.000
Sustainability is not an Atlas problem or an LHC problem. So we've had a lot of success talking to other folks about what sorts of things we can do. There's a little list of folks we've collaborated with.

00:13:18.000 --> 00:13:27.000
The idea isn't to duplicate anything that they're doing. It's to do our best to share data, work.

00:13:27.000 --> 00:13:36.000
And in case conclusions are workload dependent, like that frequency set point I showed you, that's workload dependent. It's not the same for Atlas as it is for Google.

00:13:36.000 --> 00:13:39.000
We need to know what the answer is for Atlas workloads.

00:13:39.000 --> 00:13:50.000
So these sorts of things help us out a lot. Here's a little example of some of the studies that we've collaborated on. So improving the packing of jobs in the batch system.

00:13:50.000 --> 00:14:14.000
Cern had a great system called BEER, which runs compute on storage elements. They realized CPU is over-provisioned on storage, so let's run compute on it. We just introduced low memory jobs to allow better packing of production into a batch system. This is like a 15 year old problem that resulted in a lot of idle CPU that we've finally gotten a solution to.

00:14:14.000 --> 00:14:23.000
You can schedule storage background tasks at low carbon times and save a lot of carbon, it turns out. So these are things like data re-indexing and error correction.

00:14:23.000 --> 00:14:29.000
That need to happen on a big storage system and you just schedule them in a better way.

00:14:29.000 --> 00:14:39.000
Build a new data center. They amortize over five to 10 years. So almost certainly you will want to build a new data center if you can make any significant improvement in PU.

00:14:39.000 --> 00:14:46.000
Run the system hot. You don't lose any performance. The hardware lifetime isn't affected. Run the system hot.

00:14:46.000 --> 00:14:54.000
Run water cooling if you can. Reuse your waste heat if you can. These are all things that Atlas sites have been working on and trying to improve.

00:14:54.000 --> 00:15:00.000
So we're doing our best to build a more complete model of the carbon footprint for Atlas computing.

00:15:00.000 --> 00:15:04.000
You heard earlier today, that's hard. We're doing our best. Hopefully we'll get there.

00:15:04.000 --> 00:15:13.000
It is clear that limited models and limited information risk harmful recommendations, which is something we want to avoid.

00:15:13.000 --> 00:15:23.000
We're trying to make folks more aware of the environment. The goal here is to make recommendations towards a more sustainable computing model As we approach the HLHC.

00:15:23.000 --> 00:15:28.000
Since I'm out of time, I'll just mention my two favorite examples of that.

00:15:28.000 --> 00:15:35.000
Dramatic music. One is GPU usage. So we talked about GPU usage a lot earlier.

00:15:35.000 --> 00:15:40.000
There's a payoff point where I've bought a GPU and I've paid for it in embodied carbon terms.

00:15:40.000 --> 00:15:49.000
But I'm not running it enough to make it worthwhile. And rough calculations say 30 to 60% is the magic number.

00:15:49.000 --> 00:15:57.000
Maybe closer to 60% in the future. That if I'm not putting at least that average load on a GPU, I should have just bought CPU.

00:15:57.000 --> 00:16:00.000
There are a lot of other questions that we can worry about.

00:16:00.000 --> 00:16:07.000
We're also closely watching extrapolations that affect these recommendations, like the decarbonization of the power grid.

00:16:07.000 --> 00:16:14.000
Which makes embodied carbon more important than operational carbon. And changes to processors.

00:16:14.000 --> 00:16:26.000
If you want to read more about anything I just said, we just yesterday had a paper hit archive with all of these studies and more so happy reading. Thanks.

00:16:26.000 --> 00:16:31.000
Thanks, Zach, for this really amazing talk. Are there any questions?

00:16:31.000 --> 00:16:39.000
For Zach?

00:16:39.000 --> 00:16:42.000
I don't ah okay greg go ahead

00:16:42.000 --> 00:16:54.000
I didn't quite understand your comment that running the process as hot doesn't affect or degrade their lifetime.

00:16:54.000 --> 00:16:56.000
Why do you say that?

00:16:56.000 --> 00:17:05.000
So processors generally come with a range that they're happy to run in and it's wider than you might think.

00:17:05.000 --> 00:17:11.000
And indeed, surely if you run them well above that recommended range.

00:17:11.000 --> 00:17:20.000
You're going to get into problems. But what we found was several data centers. In the particular case of this study, it was Brookhaven.

00:17:20.000 --> 00:17:33.000
Tended to run something like 15 to 20 degrees below the top of the recommended range. And actually, they could just run closer to the top of the recommended range.

00:17:33.000 --> 00:17:42.000
And there's no indication we can find in any literature that that causes any problems for the lifetime.

00:17:42.000 --> 00:17:43.000
As long as you're within the range.

00:17:43.000 --> 00:17:52.000
Okay. Do these processes presumably have operating systems in them that slow the clock speed down if they get too close to the limit, right?

00:17:52.000 --> 00:18:08.000
That's what we were worried about and why we checked the performance as a function of those fan speeds in this case, which is a proxy for temperature set points and didn't find any significant issue.

00:18:08.000 --> 00:18:15.000
Thank you. Because we're on this slide, I had a question of my own. So on your point of building a new data center.

00:18:15.000 --> 00:18:16.000
Hmm.

00:18:16.000 --> 00:18:33.000
Of course that costs money and the funding agencies are not always happy to give that money. So is there some kind of discussions where Atlas is actively suggesting to the funding agencies that data centers should be refurbished.

00:18:33.000 --> 00:18:35.000
Or how does this work?

00:18:35.000 --> 00:18:46.000
What we're able to concretely say is here's some easy math you can do to see if in operational terms, operational carbon terms, this will pay off.

00:18:46.000 --> 00:18:53.000
We can also say, here's some easy math you can do to see if in power consumption terms in terms of your electricity bill.

00:18:53.000 --> 00:19:05.000
This will pay off. And in some cases, you can make those two arguments together And they'll buy it and say, okay, yeah, actually after five years, I'll be saving money. So this is worth doing.

00:19:05.000 --> 00:19:11.000
It also depends on the location. So in Norway, don't build a new data center.

00:19:11.000 --> 00:19:16.000
The power is too green. In Poland, definitely build a new data center and put solar panels on the roof.

00:19:16.000 --> 00:19:25.000
So we're trying not to make any blanket recommendations, but to give people the tools they need or a little equation that they can fill in some numbers in.

00:19:25.000 --> 00:19:31.000
To say, oh, yeah, this totally makes sense for us and we should approach our funding agency about it.

00:19:31.000 --> 00:19:34.000
Perfect. Hopefully it works out in the long run. Right. So do we have any last pressing questions for Zach?

00:19:34.000 --> 00:19:39.000
I hope so.

00:19:39.000 --> 00:19:41.000
Otherwise, we can also move the discussions to the matter most channel.

00:19:41.000 --> 00:19:46.000
If not

