WEBVTT

1
00:00:09.196 --> 00:00:12.860
Dwayne Spiteri: Okay, fine. So assume I'll just share my slides and

2
00:00:13.020 --> 00:00:15.399
Dwayne Spiteri: had them up. They disappeared somewhere.

3
00:00:16.030 --> 00:00:17.050
Dwayne Spiteri: Okay.

4
00:00:32.400 --> 00:00:33.589
Dwayne Spiteri: let's see, I'm doing.

5
00:00:45.411 --> 00:00:49.509
Dwayne Spiteri: Can you guys hear me? Okay, I'm sorry. Just having some trouble with my sound a bit.

6
00:00:50.910 --> 00:00:52.440
Stefania Juks: Yes, this one should be fine.

7
00:00:57.940 --> 00:00:58.720
Dwayne Spiteri: Okay.

8
00:01:16.190 --> 00:01:21.129
Dwayne Spiteri: okay, perfect. Right? Apologies for the delay. I can now share my screen.

9
00:01:25.850 --> 00:01:27.209
Dwayne Spiteri: And you guys can see that.

10
00:01:28.820 --> 00:01:29.619
Stefania Juks: Yes, that's perfect.

11
00:01:30.030 --> 00:01:36.239
Dwayne Spiteri: Perfect. Okay, then start the timer. Okay? So hi, I

12
00:01:36.570 --> 00:01:49.979
Dwayne Spiteri: I'm doing. I'm working here at Daisy for the past 6 months. And I am part of a project to try and make some sustainable or more sustainable our

13
00:01:50.210 --> 00:01:56.480
Dwayne Spiteri: designs and our futures and our work, that we submit to the data center here at Daisy.

14
00:01:57.140 --> 00:01:59.230
Dwayne Spiteri: And so

15
00:01:59.620 --> 00:02:13.910
Dwayne Spiteri: the Idf is the international data analysis facility. And this is kind of the ecosystem that uses from around the world submit scientific work to Daisy. So we have a large scientific team from Hep on campus.

16
00:02:13.940 --> 00:02:42.390
Dwayne Spiteri: and they use resources. And we have an external team from different experiments that also use them. And while in this talk of sustainability, while I work alongside the colleagues at the data center, it's important to note that sustainability efforts are going to be limited. If the wider parts of the ecosystem don't actually talk to each other and feature into the discussions and the work going on to try to make this more sustainable.

17
00:02:42.690 --> 00:02:59.089
Dwayne Spiteri: And what I'm funded by, and I'll talk about it very briefly in a moment is this research facility 2.0 project, which is also got colleagues at Cen and Alba, and some partners in industry which are working together with research labs to try to make them more sustainable.

18
00:02:59.740 --> 00:03:11.870
Dwayne Spiteri: So what does the head part fit in? So typically most of people in the community that use our resources are 2 types. There's the uses that are local to Daisy, that interact with our naf resources.

19
00:03:12.431 --> 00:03:21.909
Dwayne Spiteri: And there's the experimental facilities that will send work packages or pilots to our grid services and their submission frameworks.

20
00:03:22.140 --> 00:03:34.089
Dwayne Spiteri: I'll talk about the difference between that and mapping grid a bit later. But the idea here is that different user groups will use our resources in a different manner. And so there's no one. Size fits all solution to everybody.

21
00:03:34.400 --> 00:03:50.340
Dwayne Spiteri: and the pilot is a packet of work that is sent to us from the experiment. It runs very quickly a local job on our resources, and then it says the experiment, send the rest of the job our way, and then it takes out those resources for a fixed period of time and runs work.

22
00:03:51.310 --> 00:03:59.659
Dwayne Spiteri: So I very briefly, commented about this on an earlier slide. But the the research facility is an EU Funded project which is looking at

23
00:03:59.710 --> 00:04:26.679
Dwayne Spiteri: the ways we can make and shape our research infrastructure to make them more sustainable in the future. And the work packages are roughly such that there's 3 groups. There's the identify phase which is complete where we look at investigated technologies. And there's a prototyping phase which are currently in where we are doing lots of things to try and design and create components that be used in future facilities. And then there's implementation where we build demonstrators.

24
00:04:26.700 --> 00:04:35.950
Dwayne Spiteri: Daisy, in this sense, is the only institution in one of the partners looking at and to develop strategies for data centers. Given that research

25
00:04:36.260 --> 00:04:41.140
Dwayne Spiteri: facilities in the future, especially accelerators where all of them require data centers.

26
00:04:43.040 --> 00:04:54.659
Dwayne Spiteri: So the Daisy data center is managed as 2 separate clusters, and it's split over the main site in Hamburg and the smaller sister site in Zoyton, which is Brandenburg, just south of Berlin and

27
00:04:55.182 --> 00:05:00.027
Dwayne Spiteri: the Daisy data center is about 4 times as big as the

28
00:05:00.850 --> 00:05:08.520
Dwayne Spiteri: as the Zoyton one. And in both systems. The main average power usage of the main is our compute systems.

29
00:05:08.670 --> 00:05:20.198
Dwayne Spiteri: But this is not everything. And given that, then we're mainly based on the main site. I wanted to talk a bit more about the the data center. Specifically. So, as I said, the main

30
00:05:20.510 --> 00:05:39.350
Dwayne Spiteri: power usage is compute. But it's not all, and it's not everything. So you can see in the plot on the left that all of the main components that we've measured come out of our data center. So most of it is compute. But there's a lot in terms of the storage, some in terms of operations. So that's like the it.

31
00:05:39.780 --> 00:05:50.349
Dwayne Spiteri: The people that the machines that we need to require to keep things going, networking infrastructure, and the water cooling which comes from a different place. On site.

32
00:05:50.550 --> 00:06:17.719
Dwayne Spiteri: Now all of these are very useful, but our compute clusters is split into 3 main groups, Naf Grid and Maxwell, and, as I've mentioned before, the Naf is a series of computers which is for local Hep uses, and the grid is not necessarily directly accessed by our local group, but they are accessing those resources via the institutions that they come from. So atlas cms experiments.

33
00:06:19.170 --> 00:06:36.120
Dwayne Spiteri: So what are the site? Challenges that we have to sustainability. So I want to talk about very briefly 2 things that we are doing at Daisy that will lead into the sculpting, and of the idea of the future, of submission, of work at the site.

34
00:06:36.310 --> 00:06:51.070
Dwayne Spiteri: So at the current, one of the problems that we're having is the model of the energy provision is limiting us from saving power in the sense that because we are one of the major con power consumers in Hamburg.

35
00:06:51.190 --> 00:07:09.769
Dwayne Spiteri: Our provider has more interest in predictable user patterns. So they want us to operate within a power variance band. And if our what draw exceeds or goes lower than this, then we get penalized in a heavy cost price. So there's not necessarily an incentive for us to go as low as possible.

36
00:07:10.280 --> 00:07:13.030
Dwayne Spiteri: And so we have

37
00:07:13.860 --> 00:07:22.510
Dwayne Spiteri: but there is some incentive to try to be flexible with our load. So the 1st thing I want to talk about is some of the seasonal savings we've done as our data center.

38
00:07:23.470 --> 00:07:29.319
Dwayne Spiteri: So this the way. So in the winter, when the usage of demand is high,

39
00:07:29.730 --> 00:07:56.340
Dwayne Spiteri: from the whole of Hamburg, we were asked to see if we could shed some of the work on our nodes, and this was completed a few years ago. Typically speaking, when we talk about compute experiments, have this quota of how much resource that we have to give them and most sites in the in the Lhc computing grid over provision. So we we provide more resource than and they request from us.

40
00:07:56.440 --> 00:08:04.749
Dwayne Spiteri: So we wanted to try to ramp down to the only pledge levels over some holiday by slowly shedding work without making sure the jobs failed.

41
00:08:05.353 --> 00:08:12.459
Dwayne Spiteri: We found that this took about a week to get to these pledge levels, and it took a very long time

42
00:08:12.880 --> 00:08:14.100
Dwayne Spiteri: to do so.

43
00:08:14.503 --> 00:08:31.310
Dwayne Spiteri: And we also ran some simulation to see how long it were to drain a single node, and if the model of the future is to be expected, we want to be more flexible when it comes to workload. So of the order of half an hour to an hour, and

44
00:08:31.390 --> 00:08:50.380
Dwayne Spiteri: with the current ways of the this experiment, pilots where they require 48, 72 h of our time. The shedding of this work becomes over a long course of time, spread over a large period of days, and it becomes very difficult for us to make savings quickly.

45
00:08:50.640 --> 00:09:01.796
Dwayne Spiteri: That's what we noted. So some of the things we're trying to do. So coming up. So in the next couple of months. We're trying to do some summer savings. So when the usage on the site is going to be quite high,

46
00:09:02.230 --> 00:09:25.409
Dwayne Spiteri: we want to try to reduce the amount of power that the data center is producing and some of our experiments so as a site, we were asked to try to see if we could save some power, and so, instead of changing the amount of of work we will try to then do is set the clock frequency. So the amount of calculations the computer can do per second, we can change that. So we can reduce that.

47
00:09:25.450 --> 00:09:48.770
Dwayne Spiteri: And the work still gets done. But it happens over a longer period of time. And so we estimate, if that if we in this period of time where they're asking us to save power, we go to the minimum frequency or a lower frequency that the machines can handle. We can save about 40 kilowatts of power draw within that time.

48
00:09:49.120 --> 00:09:59.310
Dwayne Spiteri: And again, this is not. This is about trying to operate within that band. And then there was some idea about whether we could turn off the oldest machines together to save some excess power.

49
00:09:59.780 --> 00:10:13.780
Dwayne Spiteri: But this slightly interesting, because when we have this, we need to talk about accounting, because when we have a quota, we say to the experiments that we are going to provide. This amount of Hep score hours, and the Hep score is a as A

50
00:10:14.040 --> 00:10:30.489
Dwayne Spiteri: is a measure of how good. A machine is at running work, and ours is integrated over time. And so when we clock our frequency, this changes. And so at the moment we're going to report some sort of very rough estimate of a balance between this minimum and maximum and

51
00:10:30.993 --> 00:10:40.700
Dwayne Spiteri: power that we have in and out of this period that we're using. And hopefully, this should be a very rough guide to what we could do in the future.

52
00:10:43.530 --> 00:11:06.339
Dwayne Spiteri: So the other thing that we have, and that was the seasonal saving. So that's something we could do. On a shorter time scale on a longer time. Scale. The electricity price seems to be moving away from this model of cost, price and network, and moving towards something which rewards flexibility in the site. So can you adjust your energy usage on demand.

53
00:11:06.370 --> 00:11:21.819
Dwayne Spiteri: Now, as a whole site. Daisy, is not that flexible when it comes to energy usage, because the accelerators and the associated cryo require a lot of constant power being used if they've been operation.

54
00:11:22.265 --> 00:11:37.664
Dwayne Spiteri: So what we're trying to do is see how much of our load is potentially flexible. And you can see the data center here incorporates maybe about 8%, but potentially with assisted from the cryo and others associated.

55
00:11:38.376 --> 00:11:56.730
Dwayne Spiteri: power draws. It could be a bit more than that, and we need to see how flexible we are in terms of our power usage here with the data center just so potentially in the future, we could try to see if we can adapt our power loads to a signal given to us by a power company, for example.

56
00:11:57.720 --> 00:12:04.160
Dwayne Spiteri: So one of the things that I'm that so my personal work here is is the creation of a digital twin of the Daisy data center.

57
00:12:04.682 --> 00:12:20.759
Dwayne Spiteri: It was you initially created the University of Glasgow expanded by funding from R. 2 0. The idea is to simulate the data center, compute and have an estimate of how much carbon is being used in this operation.

58
00:12:21.170 --> 00:12:37.069
Dwayne Spiteri: The idea here is that we take in some data something like the carbon intensity of the local grid. The the architecture quantity and specs that we have of the machines in our data center.

59
00:12:37.360 --> 00:12:44.190
Dwayne Spiteri: Then we get some machine schematics like our and the heptical benchmark I mentioned earlier about how well the machines are doing work.

60
00:12:44.300 --> 00:13:06.860
Dwayne Spiteri: and then we can feed in different energy strategies. And then we can simulate per time step, how much carbon is used for particular given work. And then we can compare this with different energy, saving strategies to see if we can estimate for our whole cluster whether we save energy or not, without actually having to impact the cluster. And this is a low carbon way and highly

61
00:13:06.860 --> 00:13:15.399
Dwayne Spiteri: highly non, disruptive way of testing these things without actually having to to do them or buy new machines. To test these things as well.

62
00:13:15.760 --> 00:13:24.770
Dwayne Spiteri: So the currently the output in is very basic, but gives you some real and simulated time of the duration, some information about how much jobs is finished.

63
00:13:25.132 --> 00:13:54.579
Dwayne Spiteri: The total average CPU duration, and some estimates on the amount of power we use and the amount of carbon dioxide. Equivalent emissions were done for this work, and the numbers themselves should be taken with a pinch of salt. The idea there is to compare and contrast them with other things you could do. So use case one is, can we save carbon by shifting work? So similar to what we're doing in the afternoon. If we have to deliver a specific amount of

64
00:13:54.960 --> 00:13:59.719
Dwayne Spiteri: of one for specific amount of time, if we drop our

65
00:14:01.510 --> 00:14:26.330
Dwayne Spiteri: our frequency to a particular level. How much do we actually save? And this is at the simulation of Glasgow. What I've did, where I found that we could each job we could have can run less. Co. 2. But we have an overall reduction in jobs. But also, if you actually care about shifting energy away from peak times we see an overall peak time, energy reduction, which is also separately important.

66
00:14:27.594 --> 00:14:48.279
Dwayne Spiteri: There's quite another of ways. You can use this simulation. I don't want to go through more of them here. But I did want to just sort of generally talk about the overall picture in terms of how the leave of back to the beginning how the entirety of the ecosystem can help, and in terms of creating things that can be simulated, and also into best practices.

67
00:14:48.280 --> 00:15:05.089
Dwayne Spiteri: So sort of color coded them here. It's like what the users can do is different to what the excellence can do is different to what the hardware vendors can do. And we need to have an ongoing discussions with new clients like hardware vendors and utilities to potentially give us better tools

68
00:15:05.210 --> 00:15:10.177
Dwayne Spiteri: at combating the effects that we have here.

69
00:15:10.970 --> 00:15:22.850
Dwayne Spiteri: so one of the things is this pilot info functionality can be except can be done in conjunction with the external facilities. And potentially, we can come up with a solution that works for both them and us.

70
00:15:23.190 --> 00:15:31.139
Dwayne Spiteri: So the last slide and precise experience to be as environmentally conscious as possible. They need to work

71
00:15:31.170 --> 00:15:57.290
Dwayne Spiteri: together with us and with the users to create the most impactful solutions. So things like pipelines and functionality can be really the best we could do. We both need to be able to plan ahead, because this is seems to be what the future is doing. We need to be able to dynamically access resources and experiments and sites need to be able to dynamically provide and withdraw provision of resources. This is going to be the future

72
00:15:57.650 --> 00:16:19.620
Dwayne Spiteri: and energy markets. This in Germany potentially move to a model where the price is contingent on site, flexibility and money maybe really is the greatest motivator here, and we're doing lots of work a day to try to do this. But we need to reduce our runtime. The main thing that it doesn't matter what work you run. The main thing that compacts your Co. 2

73
00:16:20.589 --> 00:16:29.419
Dwayne Spiteri: status is whether or not you could reduce the amount of time you're running your simulation on compute, and that's my! Thank you so much.

74
00:16:34.250 --> 00:16:42.190
Shreyasi Acharya: Thanks a lot. It was pretty on time and a wonderful job. Please go ahead if anyone has any questions.

75
00:16:53.186 --> 00:16:53.843
Shreyasi Acharya: Okay,

76
00:16:54.690 --> 00:17:09.012
Shreyasi Acharya: so maybe just a quick question. I don't know if I missed this. But could you please, repeat once more. So you had your plans for the summer and already. Your winter savings.

77
00:17:09.690 --> 00:17:26.680
Shreyasi Acharya: with load sharing in the winters, and also with changing the frequencies of the CPU. So like quantitatively do you have. Did you already show, like, how much of the computational capacity we would be losing versus the energy that we like save.

78
00:17:28.059 --> 00:17:51.059
Dwayne Spiteri: So for this, for this one here, specifically, because upcoming. This is work that is to be done this coming summer. We don't have this. And this is one of the things we're trying to do. So. The reason why we want to do this is because a we want to have this functionality in our back pocket and be, we need to test really what happens to the provision of our resources and what that looks like. So this is going to be a 1st real test of

79
00:17:51.059 --> 00:18:09.529
Dwayne Spiteri: this outside of software simulations to see if we can do this and how much it is impacted. But it's important to note here that we always over provision. So even when we reduce the frequency, we should still be more than providing the the.

80
00:18:09.599 --> 00:18:11.999
Dwayne Spiteri: the resources, we said, we are for this time period.

81
00:18:14.500 --> 00:18:23.370
Shreyasi Acharya: Thank you. I'd like to see the results. Yeah, whenever you have them. So, Jan, we will take a quick question. Before moving to the next.

82
00:18:23.550 --> 00:18:29.427
Yann Coadou: Thank you. Just a very quick one. When you mentioned using

83
00:18:30.590 --> 00:18:37.749
Yann Coadou: frequency scaling, you are aiming it at not using the full power of the CPU when

84
00:18:38.200 --> 00:18:43.619
Yann Coadou: when you have high carbon emissions due to your power source. But

85
00:18:44.180 --> 00:18:51.920
Yann Coadou: imagine that you don't have this issue anymore, that you don't have varying carbon dependency.

86
00:18:52.270 --> 00:19:10.320
Yann Coadou: The fact that you reduce the CPU frequency means also that you run longer. So you consume more energy while running longer. Have you compared the impact of a higher frequency on a shorter time to lower frequency and longer time running.

87
00:19:11.540 --> 00:19:13.140
Dwayne Spiteri: In terms of emissions.

88
00:19:13.450 --> 00:19:38.580
Dwayne Spiteri: Yes, that's that is a that's a good question. So if we are to try to divorce the idea of the local intensity from the running. What we actually find is 2 things. If we don't actually care about the the carbon, we might actually save the actual energy. So we actually care about when the energy is consumed. So if we want to have a whole flat.

89
00:19:39.040 --> 00:19:46.889
Dwayne Spiteri: hi, Dave, of this energy usage across sort of Hamburg. We might want to operate outside of

90
00:19:47.241 --> 00:20:01.389
Dwayne Spiteri: hours, and this might not affect our computes. They might work longer, but they do actually use a roughly around the same amount of energy. So there have been studies, at least from Glasgow that showed that if you run lower but longer

91
00:20:01.390 --> 00:20:25.210
Dwayne Spiteri: you do to use a similar amount of energy, and in some cases you can actually slightly reduce the energy you use, but very minuscule. The reason, really, why you should think about scaling down. If you're not worried about intensity is to share the energy burden across 2 different time. Zones. So you can have a wider view of the

92
00:20:25.894 --> 00:20:34.995
Dwayne Spiteri: energy balance for a city or for your organization. And so when you use power is also as important as how much you use

93
00:20:35.660 --> 00:20:47.950
Dwayne Spiteri: and sharing that out can be beneficial in terms of carbon savings because of marginal costs of things like district heating irrespective of the external sources of power.

94
00:20:50.420 --> 00:20:50.880
Yann Coadou: Thanks.

95
00:20:53.090 --> 00:20:53.760
Shreyasi Acharya: Right

96
00:20:55.520 --> 00:21:01.914
Shreyasi Acharya: although we are a bit late. But I would also like to take one more question from Greg, and then we move to the next

97
00:21:03.320 --> 00:21:18.129
Greg Hallewell: Yeah, Hi, the the question is actually related to the question that Jan just asked, and this slide that you're showing right now, and that is, you need to do 2 things, probably to really benefit from this clock speed reduction.

98
00:21:18.130 --> 00:21:43.139
Greg Hallewell: You also need to probably cool the individual processor chip at the chip level, and when the clock frequency goes down, then the power dissipation goes down, and you would need a cooling system which was reactive enough to reduce the coolant flow and reduce the heat exchange. And I don't think that exists right now, because these servers are cooled at the server level rather than the individual processor, chip level.

99
00:21:44.780 --> 00:21:47.599
Greg Hallewell: Do you have any comments or thoughts on that.

100
00:21:48.180 --> 00:21:53.729
Dwayne Spiteri: That's a that's a very good point. So the cooling systems that we have

101
00:21:54.280 --> 00:22:09.590
Dwayne Spiteri: have a a flux that we we pull into. And there's a there's a steady rate of change of the amount of cooling. Now, what we actually see here is that potentially, we are saving some power because

102
00:22:09.590 --> 00:22:37.519
Dwayne Spiteri: the water is coming in. There's a cycle of of cooling and heating, and if the computer is, have a lower temperature, then the difference in water intake to water outtake is going to be slightly less, and if that is lowered, then the amount of energy required to cool the water, to go back into the intake is lowered as well. So yes, you are correct. We do potentially also need to work on a more dynamic cooling, but reducing the amount of

103
00:22:37.520 --> 00:22:54.230
Dwayne Spiteri: power that the CPU takes across the entire site will affect the the rate of water cooling applied onto there, and therefore it will reduce the amount of power, because the the water, the difference in water temperature will be less than if we're running at Max power.

104
00:22:55.670 --> 00:22:56.609
Greg Hallewell: Okay. Thanks.

105
00:22:56.850 --> 00:22:57.949
Greg Hallewell: Yep. Okay.

106
00:22:59.980 --> 00:23:06.339
Dwayne Spiteri: I am on the same matter most so please just ping me any messages that you want to to get me on that as well.

107
00:23:09.270 --> 00:23:20.260
Shreyasi Acharya: Thanks a lot, Duane. It was a wonderful yeah. Presentation as well. Very interesting discussions so we would like to move to the next.

