WEBVTT

1
00:00:00.870 --> 00:00:18.259
Sukanya Sinha: Great stuff that's happening in Atlas in terms of environmental impact and sustainability for computing. I'll give you a 15 min timer and it will buzz with a very funny sound, so that you know when your time is up, and then we have time for questions. So yes, Zach, whenever you're ready, go ahead.

2
00:00:18.810 --> 00:00:38.859
Zach Marshall: Thanks very much, and I will do my best to run through our Atlas computing sustainability studies. Here on the top. Here you can see a distribution of the computing that we run worldwide. It really is a global effort. And I'll tell you about some of the reasons that that's important and interesting in the talk.

3
00:00:40.030 --> 00:00:52.339
Zach Marshall: So for scale, we operate something like 700,000 cores of compute worldwide. And you see peaks up to a million cores when we take over big chunks of high performance computing systems.

4
00:00:52.450 --> 00:00:58.999
Zach Marshall: and we have something like an exabyte of data between disk and tape at the moment.

5
00:00:59.740 --> 00:01:19.430
Zach Marshall: That's a combination of grid sites that we talked about earlier in the session and high performance computers, as used to be called supercomputers, cloud computing, and some volunteer computing. And again, those peaks are mostly driven by those high performance computing centers where we're users, but certainly not the reason for existence of the data centers.

6
00:01:19.690 --> 00:01:24.590
Zach Marshall: When we look forward to the future, to the Hlhc. Those resources are going to grow.

7
00:01:24.720 --> 00:01:44.530
Zach Marshall: So we know by the time the Hlhc starts up in sort of 2030 ish. Now, we're going to need a factor of 3 to 5 more compute, disk and tape. And by the end of data, taking in 2041, the end of the collaboration. We'll need another factor of 2. So we're talking about an order of magnitude growth.

8
00:01:44.740 --> 00:01:53.430
Zach Marshall: which means we have a nice lever arm to adjust. If we can improve things, we can pull down the projected carbon cost of those resources.

9
00:01:54.330 --> 00:02:04.319
Zach Marshall: In general, we make resource requests 12 to 18 months in advance. So in year, Jevon's paradox applies where you make something faster, and we just do more of it

10
00:02:05.217 --> 00:02:24.869
Zach Marshall: over the course of one to 2 to 3 years. There is opportunity for reduction of resources through optimization. And we have seen that where we've made a work flow faster, we've made some data smaller, and therefore 2 years out, 3 years out we've requested fewer resources and and had some savings there.

11
00:02:25.550 --> 00:02:33.330
Zach Marshall: So obviously, if we can pull down these lever arms, then we'll reduce the carbon cost of the experiment. But there are a lot of other things that we can do as well.

12
00:02:33.510 --> 00:02:52.269
Zach Marshall: So one simple thing is to just build awareness in the collaboration of what's going on. Every user when they run a job on the distributed computing system gets an email like this when it's done. And you notice down here there's an estimated carbon footprint for the task section which gives them a link to what it means and how it's calculated, and so on.

13
00:02:53.448 --> 00:02:59.190
Zach Marshall: This includes now an average carbon footprint averaged worldwide

14
00:02:59.570 --> 00:03:23.049
Zach Marshall: for the tasks that they've run, and we keep internally the information about the actual carbon footprint for that task in that location, for the grid carbon intensity. There we provide a worldwide average for a few reasons. One is, there are a lot of data centers. We don't control them. We don't have all the information about all of them. We don't know where there are solar panels yet. So we're trying to get that information until we know we we shouldn't

15
00:03:23.240 --> 00:03:24.800
Zach Marshall: charge sites. Extra.

16
00:03:25.270 --> 00:03:37.460
Zach Marshall: Second thing is, CPU doesn't sit idle. So if a user moves a job from Poland to Norway, we'll put a different job on the site in Poland. We won't just let the CPU sit idle. So moving jobs around doesn't actually save carbon.

17
00:03:38.120 --> 00:03:43.004
Zach Marshall: The 3rd reason and this is something that you'll hear me say a few times.

18
00:03:43.460 --> 00:03:48.360
Zach Marshall: if all the users take action and everyone pushes their work to Norway.

19
00:03:48.900 --> 00:04:10.480
Zach Marshall: The Norwegian data center will get absolutely hammered, the disks will burn out faster, the network will be hammered, and the jobs will become less efficient, and we'll still be running the data centers elsewhere. So actually recommending that they all do that would cost us in carbon terms, it would be a negative

20
00:04:11.890 --> 00:04:23.119
Zach Marshall: solution. So there are a lot of cases like this where, if you sort of take too myopic a view, you can make a recommendation that's harmful globally.

21
00:04:23.870 --> 00:04:33.290
Zach Marshall: The last reason is, it's a nice correlation. You want to say, if somebody makes their code faster, that they have a lower carbon footprint. And if you do a global average, that correlation's stronger.

22
00:04:34.090 --> 00:04:53.689
Zach Marshall: it's, of course, not intended to shame anybody. We just want to make them aware of what's going on. And we also do this now for production. So you can see, here's a production that cost 2 tons of carbon, about 130 of which was failed jobs. That's pretty typical. 5, 6% failures is about right for our grid productions.

23
00:04:54.460 --> 00:05:05.500
Zach Marshall: We also have started to build awareness early on in tutorials. So there's this Atlas analysis tutorial, and we teach folks how to run a failing job and then diagnose what happened.

24
00:05:05.730 --> 00:05:09.879
Zach Marshall: and then to look at the carbon impact of that failing job.

25
00:05:10.296 --> 00:05:29.920
Zach Marshall: We find people are not familiar with how this all works, and we found cases in the wild where users have retried jobs hundreds of times, and they just keep failing around the world hundreds of times. So this seems like an education issue. Now, we just need to make people more aware of how to fix things and what goes wrong.

26
00:05:31.230 --> 00:05:53.460
Zach Marshall: Then there are a bunch of things that we can do in terms of policies that impact our carbon footprint. I really like these because building a new data center buying new hardware takes years takes years to replace old hardware. If you change a policy, if you change what you store, that's almost instant in terms of its carbon impact. So these are really nice things to be able to get after.

27
00:05:53.800 --> 00:05:58.210
Zach Marshall: Mostly, historically, they've been thought of in financial terms

28
00:05:58.290 --> 00:06:25.020
Zach Marshall: for things like the data carousel, which is a system by which we use tape more than disk and move data actively on and off of tape. The idea was we'd save cash because tape is cheaper than disk. We also save carbon, because in carbon terms tape is cheaper than disk. So this is a case where there's a nice synergy, everything's pointing in the same direction. Obviously, there are other places where they point in opposite directions, and we have to justify things somehow.

29
00:06:25.760 --> 00:06:33.870
Zach Marshall: Another nice example is data reproduction. So if a user says, I want to save this data just in case I need it. Do I save it? Or do I delete it?

30
00:06:34.400 --> 00:06:44.750
Zach Marshall: And our recent most recent math says, if they request it less than about 15% of the time, we should just delete it and reproduce it on demand.

31
00:06:45.060 --> 00:06:52.449
Zach Marshall: And the request rate is about 25%. Right now give or take. So right now, we're better off keeping it.

32
00:06:52.660 --> 00:07:04.840
Zach Marshall: That balance could change in the future, such that it's actually better for us to delete the data and then reproduce it if they ask. The last obvious policy is to use what we have. So we talked earlier today about the

33
00:07:05.100 --> 00:07:26.240
Zach Marshall: trigger farms. For several experiments. Those got built. We paid the embodied carbon. We paid the carbon for the tier 0 center. Those are all for operation of the Lhc. The Lhc doesn't operate year round. So when it's not on, we need to use that compute as best we can to make sure that we amortize the embodied carbon over more. Compute.

34
00:07:27.660 --> 00:07:37.539
Zach Marshall: A few other things we're looking into are things like waste lost and unused data. Right? It's okay. In my opinion, if we use carbon for science.

35
00:07:37.850 --> 00:07:46.900
Zach Marshall: it's not okay. If we waste carbon while doing science right? So this was a look at when we produce data that no one ever looks at.

36
00:07:47.200 --> 00:07:51.839
Zach Marshall: And you see, like a year ago, some data that were produced that were never looked at.

37
00:07:52.210 --> 00:08:03.969
Zach Marshall: Most of these things turn out to be too inclusive patterns. Somebody asked for all the Monte Carlo. They don't really need all the Monte Carlo, or they ask for a big production, check a few files, find a bug.

38
00:08:04.190 --> 00:08:05.150
Zach Marshall: and then

39
00:08:05.700 --> 00:08:12.219
Zach Marshall: the rest of the production just goes ahead and they don't kill it off fast enough, right? So these sorts of things cost us, and we're trying to get better.

40
00:08:12.510 --> 00:08:24.360
Zach Marshall: We're, of course, trying to improve our CPU over wall efficiency. You can see the trend over the last few years. It's upward. It's slow, but it's upward. It's hard to get higher than 90% efficiency in a lot of these cases.

41
00:08:24.750 --> 00:08:44.089
Zach Marshall: and there's a constant effort to sort of reduce the serial portions of many core multi-core jobs as well. We also have automated systems to take sites offline when they have problems. So there's a significant potential savings. When a site that starts failing can be taken offline. And we don't have lots of failing jobs there.

42
00:08:44.980 --> 00:08:50.980
Zach Marshall: We're also preparing to release a significant amount of our event generation output as open data.

43
00:08:51.340 --> 00:09:02.939
Zach Marshall: So this would be a sort of service to the community. It's good scientifically, that's great. It's hard because we have to document everything that we produced publicly. Okay, that documentation is good for us. So maybe that's okay.

44
00:09:03.617 --> 00:09:13.929
Zach Marshall: We're going to release something like 10 million CPU hours of event generation. And now we have to ask a really interesting question that we'll have to come back to in about a year to see what the answer is?

45
00:09:14.020 --> 00:09:30.270
Zach Marshall: One is, does the community pick it up at all? Do they actually use it. A second is, does it actually reduce the amount of event generation that is run in the community? It's not obvious that people will necessarily use ours and not generate their own, or won't generate more and use ours?

46
00:09:30.270 --> 00:09:48.630
Zach Marshall: And the 3rd is, does it actually result in a carbon savings, or in reality do people use the same or even more carbon because they're able to run over larger numbers of events. So we'll have to see what in carbon terms this does, it might be different from the answer to what scientifically it does.

47
00:09:49.040 --> 00:09:51.549
Zach Marshall: So hopefully, I'll come back in a year and tell you more.

48
00:09:52.530 --> 00:09:55.450
Zach Marshall: We've also looked into job failures. Quite a bit.

49
00:09:55.750 --> 00:10:08.810
Zach Marshall: We found out it's it's very important to decorrelate issues to avoid victim blaming. Right? So we have this tier 0 center at Cern. It's very efficient. The failures are quite rare. That's because it runs very well controlled jobs.

50
00:10:08.860 --> 00:10:34.759
Zach Marshall: We have other data centers that have kind of unique resources where there are very high failure rates and we expect fairly high failure rates. So we've tried to build this sort of expected failure rate based on worldwide failures, and then compare sites based on their observed and expected failure rates, and then look for cases where the observed rate is way different from the expectation either above or below.

51
00:10:35.030 --> 00:10:43.079
Zach Marshall: And that's where we want to dig in to see what's going on, not just because there's a high failure rate, but because there's a higher than expected failure rate.

52
00:10:43.830 --> 00:10:52.129
Zach Marshall: So we're starting to get into this and also starting to get into automatic retry policies. So a job runs on the grid. It fails.

53
00:10:52.280 --> 00:11:03.559
Zach Marshall: and we have an automated system by which we say this failed. Let's try it again somewhere else, or well, this failed. It was probably transient. Let's just go again, and it might succeed next time.

54
00:11:03.780 --> 00:11:14.699
Zach Marshall: And obviously, if you get that wrong, you end up retrying a bunch of jobs that are just going to fail over and over again, and you should, you should stop. So we need to constantly reexamine these sorts of policies.

55
00:11:15.779 --> 00:11:23.119
Zach Marshall: We've talked a lot in this workshop about software development for sustainability. It goes without saying we should continue to improve our software.

56
00:11:23.240 --> 00:11:41.499
Zach Marshall: One thing that's worth pointing out is, it's possible to run cpus at lower frequency to improve their performance per watt. So this is a nice little plot of the frequency of the CPU. Against the performance. Per Watt. So actually, if you lower the frequency of the CPU, you get better performance. Per Watt.

57
00:11:42.370 --> 00:12:00.359
Zach Marshall: if you care only about power. That's clearly a good idea. If you have to deliver a certain number of Hep score, say, then, you might have to buy more cpus in order to provide that much power. So you have to balance embodied and operational carbon in order to get that working point right?

58
00:12:01.265 --> 00:12:10.399
Zach Marshall: We've also been working on checkpointing. That's the ability to stop a job, write it, state to disk, and then read it off of disk, and start exactly where you left off

59
00:12:10.560 --> 00:12:18.240
Zach Marshall: to see if we can do a better job during outages. Brownouts, you know, regular downtime, that sort of thing

60
00:12:18.620 --> 00:12:29.599
Zach Marshall: in the future. We're also thinking about this sort of situation where this is a projection to 2035 of of German power, and there are periods where there's a huge amount of renewable energy.

61
00:12:29.850 --> 00:12:34.759
Zach Marshall: So if we can take a data center, run it when there's a ton of renewables

62
00:12:34.990 --> 00:12:45.130
Zach Marshall: and then basically turn it off when the renewables are gone at night and then turn it back on when the load is lower than the demand. The the

63
00:12:45.320 --> 00:12:54.689
Zach Marshall: renewables are higher than the demand. We might be able to actually overall save quite a lot of carbon, and maybe even make our computing largely more neutral.

64
00:12:55.751 --> 00:12:58.718
Zach Marshall: I mentioned validation that we need to improve

65
00:12:59.420 --> 00:13:04.770
Zach Marshall: the production of buggy samples and and avoid that as much as possible. That's another ongoing effort.

66
00:13:06.100 --> 00:13:34.960
Zach Marshall: Sustainability is not an Atlas problem or an Lhc problem. So we've had a lot of success talking to other folks about what sorts of things we can do. There's a little list of folks we've collaborated with. The idea isn't to duplicate anything that they're doing. It's to do our best to share data work. And in case conclusions are workload dependent like that frequency set point I showed you, that's workload dependent. It's not the same for atlas as it is for Google.

67
00:13:35.260 --> 00:13:38.169
Zach Marshall: We need to know what the answer is for atlas workloads.

68
00:13:38.460 --> 00:13:59.849
Zach Marshall: So these sorts of things help us out a lot. Here's a little example of some of the studies that we've collaborated on so improving the packing of jobs in the batch system. Cern had a great system called beer, which runs compute on storage elements. They realized CPU is over provisioned on storage. So let's run compute on it.

69
00:14:00.030 --> 00:14:27.330
Zach Marshall: We just introduced low memory jobs to allow better packing of production into a batch system. This is like a 15 year old problem that resulted in a lot of idle CPU that we've finally gotten a solution to. You can schedule storage background tasks at low carbon times and save a lot of carbon. It turns out. So these are things like data, re-indexing and error correction that need to happen on a big storage system. And you just schedule them in a better way.

70
00:14:28.310 --> 00:14:38.209
Zach Marshall: build a new data center. They amortize over 5 to 10 years. So almost certainly you will want to build a new data center. If you can make any significant improvement in pue.

71
00:14:38.490 --> 00:14:45.040
Zach Marshall: Run the system hot. You don't lose any performance. The hardware lifetime isn't affected. Run the system hot.

72
00:14:45.160 --> 00:14:52.020
Zach Marshall: run water cooling if you can reuse your waste heat if you can. These are all things that Atlas sites have been working on and trying to improve.

73
00:14:53.410 --> 00:15:12.060
Zach Marshall: So we're doing our best to build a more complete model of the carbon footprint for Atlas computing you heard earlier today. That's hard. We're doing our best. Hopefully, we'll get there. It is clear that limited models and limited information risk harmful recommendations, which is something we want to avoid.

74
00:15:12.775 --> 00:15:22.159
Zach Marshall: We're trying to make folks more aware of the environment. The goal here is to make recommendations towards a more sustainable computing model. As we approach the Hlhc.

75
00:15:22.440 --> 00:15:28.790
Zach Marshall: Since I'm out of time, I'll just mention my 2 favorite examples of that dramatic music.

76
00:15:30.330 --> 00:15:43.460
Zach Marshall: one is Gpu usage. So we talked about Gpu usage a lot earlier. There's a payoff point where I've bought a Gpu, and I've paid for it in embodied carbon terms. But I'm not running it enough to make it worthwhile.

77
00:15:43.690 --> 00:16:10.700
Zach Marshall: and rough calculations, say 30 to 60% is the magic number. It's maybe closer to 60% in the future that if I'm not putting at least that average load on a Gpu. I should have just bought CPU. There are a lot of other questions that we can worry about. We're also closely watching extrapolations that affect these recommendations like the decarbonization of the power grid which makes embodied carbon more important than operational carbon

78
00:16:10.910 --> 00:16:23.210
Zach Marshall: and changes to processors. If you want to read more about anything. I just said we just yesterday had a paper hit archive with all of these studies, and more so happy reading. Thanks.

79
00:16:25.380 --> 00:16:31.469
Sukanya Sinha: Thanks, Zach, for this really amazing talk. Are there any questions for Zach?

80
00:16:38.400 --> 00:16:41.609
Sukanya Sinha: I don't. Okay, Greg, go ahead.

81
00:16:41.930 --> 00:16:50.739
Greg Hallewell: I didn't quite understand your comment that running the processes hot doesn't affect or degrade their lifetime.

82
00:16:51.880 --> 00:16:55.330
Greg Hallewell: Why do you say that.

83
00:16:55.710 --> 00:17:04.379
Zach Marshall: So processors generally come with a a range that they're happy to run in, and it's wider than you might think.

84
00:17:04.900 --> 00:17:19.169
Zach Marshall: And indeed, surely, if you run them well above that recommended range, you're going to get into problems. But what we found was several data centers. In the particular case of this study it was Brookhaven.

85
00:17:19.750 --> 00:17:26.199
Zach Marshall: tended to run something like 15 to 20 degrees below the

86
00:17:27.109 --> 00:17:38.219
Zach Marshall: top of the recommended range, and actually they could just run closer to the top of the recommended range, and there's no indication we can find in any literature that that causes any

87
00:17:38.340 --> 00:17:39.970
Zach Marshall: problems for the lifetime.

88
00:17:41.330 --> 00:17:43.410
Zach Marshall: As long as you're with the Ranger.

89
00:17:43.600 --> 00:17:51.669
Greg Hallewell: Do these. These processes presumably have operating systems in them that slow the clock, speed down if they get too close to the limit right.

90
00:17:52.280 --> 00:18:03.779
Zach Marshall: That's what we were worried about, and why we checked the performance as a function of those fan speeds. In this case, which is a proxy for temperature set points and didn't find any significant issue.

91
00:18:04.370 --> 00:18:05.430
Greg Hallewell: Okay. Thanks.

92
00:18:07.560 --> 00:18:25.930
Sukanya Sinha: Thank you. Because we're on this slide. I had a question of my own. So on your point of building a new data center. Of course, that costs money, and the funding agencies are not always happy to give that money, so is there some kind of discussions where Atlas is actively

93
00:18:26.210 --> 00:18:34.090
Sukanya Sinha: suggesting to the funding agencies that data centers should be refurbished. Or how does this work.

94
00:18:36.270 --> 00:18:45.210
Zach Marshall: What we're able to concretely say is, here's some easy math you can do to see if in operational terms, operational carbon terms. This will pay off.

95
00:18:45.928 --> 00:18:53.460
Zach Marshall: We can also say, here's some easy math you can do to see if in power consumption terms in terms of your electricity bill. This will pay off.

96
00:18:53.960 --> 00:19:04.248
Zach Marshall: And in some cases you can make those 2 arguments together, and they'll buy it and say, Okay, yeah, actually, after 5 years, I'll be saving money. So this is worth doing.

97
00:19:04.910 --> 00:19:29.669
Zach Marshall: it also depends on the location. So in Norway don't build a new data center. The power is too green in Poland, definitely build a new data center and put solar panels on the roof. Right? So we're trying not to make any blanket recommendations, but to give people the tools they need, or a little equation that they can fill in some numbers in to say, Oh, yeah, this totally makes sense for us, and we should approach our funding agency about it.

98
00:19:31.120 --> 00:19:42.499
Sukanya Sinha: Perfect. Hopefully, it works out in the long run. Right? So do we have any last pressing questions for Zach. Otherwise we can also move the discussions to the matter. Most channel.

99
00:19:43.030 --> 00:19:45.369
Sukanya Sinha: If not, then let's thanks. Act.

