WEBVTT

00:00:00.000 --> 00:00:09.000
Thank you so much.

00:00:09.000 --> 00:00:16.000
I should share my own slides and I had them up. They disappeared somewhere.

00:00:16.000 --> 00:00:32.000
Okay.

00:00:32.000 --> 00:00:45.000
Let's see.

00:00:45.000 --> 00:00:51.000
Can you guys hear me okay? I'm sorry, just having some trouble with my sound a bit.

00:00:51.000 --> 00:00:58.000
Yes, this one should be fine.

00:00:58.000 --> 00:01:16.000
Okay.

00:01:16.000 --> 00:01:26.000
Okay, perfect. Right. Apologies for the delay. I can now share my screen.

00:01:26.000 --> 00:01:29.000
And you guys can see that. Perfect. Okay. Then start the timer.

00:01:29.000 --> 00:01:32.000
That's perfect.

00:01:32.000 --> 00:01:59.000
Okay, so… Hi, I… I'm in James Materi. I'm working here at Daisy for the past six months and uh i um part of a project to try and make sustainable or more sustainable um designs and our futures and our work that we submit to the data center here at Daisy.

00:01:59.000 --> 00:02:10.000
The IDAF is the Interdisciplinary data analysis facility and this is kind of the ecosystem that users from around the world submit scientific work to Daisy.

00:02:10.000 --> 00:02:21.000
So we have a large scientific team from HEP on campus And they use resources and we have an external team from different experiments that also use them.

00:02:21.000 --> 00:02:28.000
And while in this talk of sustainability, while I work alongside the colleagues at the data center.

00:02:28.000 --> 00:02:42.000
It's important to note that sustainability efforts are going to be limited if the wider parts of the ecosystem don't actually talk to each other and feature into the discussions and the work going on to try to make this more sustainable.

00:02:42.000 --> 00:03:00.000
And what I'm funded by, and I'll talk about it very briefly in a moment, is this research facility 2.0 project, which has also got colleagues at CERN and ALBA and some partners in industry which are working together with research labs to try to make them more sustainable.

00:03:00.000 --> 00:03:14.000
So where does the head pot fit in? So typically most of people in the community that use our resources are two types. There's the users that are local to Daisy that interact with our NAF resources.

00:03:14.000 --> 00:03:20.000
Basic experimental facilities that sent work packages or pilots to our grid services.

00:03:20.000 --> 00:03:34.000
And their submission frameworks. I'll talk about the difference between that and map and grid a bit later. But the idea here is that different user groups will use our resources in a different manner. And so there's no one size fits all solution to everybody.

00:03:34.000 --> 00:03:51.000
And the pilot is a packet of work that is sent to us from the experiment. It runs very quickly a local job on our resources and then it saves the experiment send the rest of the job our way. And then it takes out those resources for a fixed period of time and runs work.

00:03:51.000 --> 00:04:13.000
So I very briefly commented about this on an earlier slide, but the research facility is an EU funded project which is looking at the ways we can make and shape our research infrastructure to make them more sustainable in the future. And the work packages are roughly such as three groups. There's the identify phase, which is complete, where we look at investigated technologies

00:04:13.000 --> 00:04:26.000
And there's a prototyping phase, which are commonly in, where we are doing lots of things to try and design and create components that be used in future facilities. And then there's implementation where you build demonstrators.

00:04:26.000 --> 00:04:43.000
Daisy, in this sense, is the only institution and one of the partners looking at to develop strategies for data centers, given that research facilities in the future, especially accelerators where all of them require data centers.

00:04:43.000 --> 00:04:53.000
So the Daisy data center is managed as two separate clusters and is split over the main site in Hamburg and the smaller sister site in Zoiten, which is Brandenburg, just south of Berlin.

00:04:53.000 --> 00:05:08.000
And the data center is about four times as big as the as the Zoiten one and in both systems, the main average power usage of the main is our compute systems.

00:05:08.000 --> 00:05:17.000
But this is not everything. And given that we're mainly based in the main site, I wanted to talk a bit more about the data center specifically.

00:05:17.000 --> 00:05:36.000
As I said, the main power usage is compute, but it's not all and it's not everything. So you can see in the plot on the left that all of the main components that we've measured come out of our data center. So most of it is compute, but there's a lot in terms of the storage

00:05:36.000 --> 00:05:53.000
Some in terms of operations. So that's like the it um people that the machines that we need to require to keep things going, networking infrastructure and the water cooling which comes from a different place and on site. Now, all of these are very useful

00:05:53.000 --> 00:05:58.000
But our compute clusters are split into three main groups, NAF, grid, and Maxwell.

00:05:58.000 --> 00:06:04.000
And as I've mentioned before, the NAF is a series of computers which is for local HEP users.

00:06:04.000 --> 00:06:19.000
And the grid is not necessarily directly accessed by our local group, but they are accessing those resources via the institutions that they come from. So Atlas CMS experiments.

00:06:19.000 --> 00:06:23.000
So what are the side challenges that we have to sustainability?

00:06:23.000 --> 00:06:36.000
So I want to talk about very briefly two things. That we are doing at Daisy that will lead into the sculpting and of the idea of the future of submission of hep work at the site.

00:06:36.000 --> 00:06:45.000
So the current, one of the problems that we're having is the model of the energy provision is limiting us from saving power.

00:06:45.000 --> 00:06:51.000
In the sense that because we are one of the major power consumers in Hamburg.

00:06:51.000 --> 00:06:58.000
Our provider has more interest in predictable user patterns. So they want us to operate within a power variance band.

00:06:58.000 --> 00:07:05.000
And if our what draw exceeds or goes lower than this, then we get penalized in a heavy cost price.

00:07:05.000 --> 00:07:10.000
So it's not necessarily an incentive for us to go as low as possible.

00:07:10.000 --> 00:07:23.000
And so we have… But there is some incentive to try to be flexible with our load. So the first thing I want to talk about is some of the seasonal savings that we've done as our data center.

00:07:23.000 --> 00:07:36.000
So in the winter when the usage of the demand is high um from the whole of Hamburg, we were asked to see if we could shed some of the work on our nodes.

00:07:36.000 --> 00:07:46.000
And this was completed a few years ago. Typically speaking, when we talk about compute, experiments have this quota of how much resource that we have to give them.

00:07:46.000 --> 00:08:05.000
And most sites in the LHC computing grid over-provision. So we provide more resource than And they request from us. So we wanted to try to ramp down to the only pledge levels over some holiday by slowly shedding work without making sure the jobs failed.

00:08:05.000 --> 00:08:13.000
We found that this took about a week to get to these pledge levels and it took a very long time.

00:08:13.000 --> 00:08:19.000
To do so. And we also ran some simulation to see how long it were to drain a single node.

00:08:19.000 --> 00:08:30.000
And if the model of the future is to be expected, we want to be more flexible when it comes to workload. So of the order of half an hour to an hour.

00:08:30.000 --> 00:08:39.000
And with the current ways of the experiment pilots where they require 48, 72 hours of our time.

00:08:39.000 --> 00:08:50.000
The shedding of this work becomes over a longer course of time, spread over a large period of days and it becomes very difficult for us to make savings quickly.

00:08:50.000 --> 00:09:02.000
That's what we noted. So some of the things we'll try to do, so coming up so in the next couple of months, we're trying to do some summer sailing so when The usage on the site is going to be quite high.

00:09:02.000 --> 00:09:07.000
We want to try to reduce the amount of power that the data center is producing and some of our experiments.

00:09:07.000 --> 00:09:25.000
So as a site, we were asked to try to see if we could save some power. And so instead of changing the amount of work, we will try to then do is set the clock frequency so the amount of calculations the computer can do per second, we can change that. So we can reduce that.

00:09:25.000 --> 00:09:32.000
And the work still gets done, but it happens over a longer period of time.

00:09:32.000 --> 00:09:49.000
And so we estimate if we, in this period of time where they're asking us to say power, we go to the minimum frequency or a lower frequency that the machines can handle, we can save about 40 kilowatts of power draw within that time.

00:09:49.000 --> 00:09:53.000
And again, this is about trying to operate within that band.

00:09:53.000 --> 00:10:00.000
And then there was some idea about whether we could turn off the old machines together to save some excess power.

00:10:00.000 --> 00:10:14.000
But this is slightly interesting because when we have we need to talk about accounting because when we have a quota we say to the experiments that we are going to provide this amount of have score hours. And HEP score is a as a

00:10:14.000 --> 00:10:18.000
Is a measure of how good a machine is at running work and how is it's integrated over time.

00:10:18.000 --> 00:10:30.000
And so when we clock our frequency, this changes. And so at the moment, we're going to report some sort of very rough estimate of balance between this minimum and maximum.

00:10:30.000 --> 00:10:43.000
Power that we have. In and out of this period that we're using. And hopefully this should be a very rough guide to what we can do in the future.

00:10:43.000 --> 00:10:52.000
So the other thing that we have, that was a seasonal savings. So that's something we can do on a shorter timescale. On a longer time scale.

00:10:52.000 --> 00:11:06.000
The electricity price seems to be moving away from this model of cost price and network and moving towards something which rewards flexibility in the site. So can you adjust your energy usage on demand.

00:11:06.000 --> 00:11:22.000
Now, as a whole site, Adaisy is not that flexible when it comes to energy usage because the accelerators and the associated cryo require a lot of constant power being used if they've been in operation.

00:11:22.000 --> 00:11:32.000
So what we're trying to do is see how much of our load is potentially detectable. And you can see the data center here incorporates maybe about 8%.

00:11:32.000 --> 00:11:40.000
But potentially. With assisted from the cryo and other associated power draws, it could be a bit more than that.

00:11:40.000 --> 00:11:47.000
And we need to see how flexible we are in terms of our power usage here with the data center.

00:11:47.000 --> 00:11:58.000
Just so potentially in the future we could try to see if we can adapt our a power load to a signal given to us by a power company, for example.

00:11:58.000 --> 00:12:15.000
So one of the things that my personal work here is is the creating of a digital twin of the Daisy data center It was initially created at the University of Glasgow, expanded by private current funding from RP20. The idea is to simulate the data center compute

00:12:15.000 --> 00:12:21.000
And have an estimate of how much carbon is being used in this operation.

00:12:21.000 --> 00:12:37.000
The idea here is that we take in some data, something like the carbon intensity of the local grid um the the architecture quantity and um specs that we have of the machines in our data center.

00:12:37.000 --> 00:12:44.000
Then we get some machine schematics like power and the heptschool benchmark I mentioned earlier about how well the machines are doing work.

00:12:44.000 --> 00:12:55.000
And then we could feed them different energy strategies And then we can simulate per time step how much carbon is used for a particular given work. And then we can compare this with different energy saving strategies to see if we can

00:12:55.000 --> 00:13:02.000
Estimate for our whole cluster whether we save energy or not without actually having to impact the cluster.

00:13:02.000 --> 00:13:13.000
And this is a low carbon way and highly non-disruptive way of testing these things without actually having to do them or buy them new machines.

00:13:13.000 --> 00:13:25.000
To test these things as well. So the carbon of the output in is very basic but gives you some realist simulated time of the duration, some information about how much jobs have finished.

00:13:25.000 --> 00:13:42.000
The total average CPU duration and some estimates on the amount of power we use and the amount of carbon dioxide equivalent emissions were done for this work. And the numbers themselves should be taken with a pinch of salt The idea there is to just compare and contrast them

00:13:42.000 --> 00:13:58.000
With other things you could do. So use case one is can we save carbon by shifting work? So similar to what we're doing in the afternoon, if we have to deliver a specific amount of I've run for a specific amount of time if we

00:13:58.000 --> 00:14:07.000
Drop our our frequency to a particular level, how much do we actually save?

00:14:07.000 --> 00:14:18.000
And tested out the simulation of Glasgow, what I did, I found that we could each job we could have can run less CO2, but we have an overall reduction in jobs.

00:14:18.000 --> 00:14:27.000
But also if you actually care about shifting energy away from peak times, we see an overall peak time energy reduction, which is also separately important.

00:14:27.000 --> 00:14:48.000
There's quite another of ways you can use this simulation. I don't want to go through all of them here, but I did want to just sort of generally talk about the overall picture in terms of how the lever back to the beginning how the entirety of the ecosystem can help in terms of creating things that can be simulated and also into best practices.

00:14:48.000 --> 00:15:05.000
So I've sort of color coded them here. It's like what the users can do is different to what the experiments can do is different to what the hardware vendors can do. And we need to have um ongoing discussions with new clients like hardware vendors and utilities to potentially give us better tools

00:15:05.000 --> 00:15:10.000
At combating the effects that we have here.

00:15:10.000 --> 00:15:26.000
So one of the things is this pilot info functionality can be um except um can be done in conjunction with the experiment facilities and potentially we can come up with a solution that works for both them and us So the last slide.

00:15:26.000 --> 00:15:36.000
Precise experience to be as environmentally conscious as possible they need to work together with us and with the users to create the most impactful solutions.

00:15:36.000 --> 00:15:42.000
So things like patterns and functionality can be really the best we could do. We both need to be able to plan ahead.

00:15:42.000 --> 00:15:57.000
Because this seems to be what the future is doing. We need to be able to dynamically access resources and experiments and sites need to be able to dynamically provide and withdraw provision of resources. This is going to be the future.

00:15:57.000 --> 00:16:10.000
And energy markets in Germany potentially moved to a model where the price is contingent on site flexibility and money maybe really is the greatest motivator here. And we're doing lots of work at Daisy to try to do this.

00:16:10.000 --> 00:16:15.000
But we need to reduce our runtime. The main thing, it doesn't matter what work you run.

00:16:15.000 --> 00:16:34.000
The main thing that compacts your CO2 status is whether or not you could reduce the amount of time you're running your simulation on compute. And that's my Thank you so much.

00:16:34.000 --> 00:16:52.000
For your time and a wonderful job. Please go ahead. If anyone has any questions.

00:16:52.000 --> 00:17:17.000
Okay. So maybe just a quick question. I don't know if I missed this, but could you please repeat once more? So you had… your plans for the summer and already your winter savings With load shedding in the winters and also with changing the frequencies of the CPU.

00:17:17.000 --> 00:17:27.000
Like quantitatively, did you already show like how much of the computational capacity we would be losing versus the energy that we like save.

00:17:27.000 --> 00:17:37.000
So for this one here specifically, because upcoming, this is work that is to be done this coming summer, we don't have this. And this is one of the things we're trying to do.

00:17:37.000 --> 00:17:51.000
The reason why I want to do this is because A, we want to have this functionality in our back pocket and B, we need to test really what happens to the provision of our resources and what that looks like. So this is going to be a first real test of

00:17:51.000 --> 00:17:59.000
This outside of software simulations to see if we can do this and how much it is impacted.

00:17:59.000 --> 00:18:02.000
But it's important to note here that we always open revision.

00:18:02.000 --> 00:18:14.000
So even when we reduce the frequency, we should still be more than providing the the the resources we said we offer this time period.

00:18:14.000 --> 00:18:19.000
Thank you. I'd like to see the results whenever you have them.

00:18:19.000 --> 00:18:22.000
So Jan, we will take a quick question. Before moving to the next.

00:18:22.000 --> 00:18:43.000
Yes. Thank you. Just a very quick one. When you mention using frequency scaling you are aiming it at not using the full power of the cpu When you have high carbon emissions due to your power source.

00:18:43.000 --> 00:19:04.000
But… Imagine that you don't have this issue anymore, that you don't have varying carbon dependency the fact that you reduce the cpu frequency means also that you run longer so you consume more energy while running longer. Have you compared

00:19:04.000 --> 00:19:11.000
The impact of a higher frequency on a shorter time to lower frequency and longer time running.

00:19:11.000 --> 00:19:12.000
In terms of emissions.

00:19:12.000 --> 00:19:22.000
Yes, that's a good question. So if we are to try to divorce the idea of the local intensity from the running.

00:19:22.000 --> 00:19:29.000
What we actually find is two things. If we don't actually care about the carbon.

00:19:29.000 --> 00:19:50.000
We might actually save the actual energy. So we actually care about when the energy is consumed. So if we want to have a whole flat idea of this energy usage across Hamburg, we might want to operate outside of hours and this might

00:19:50.000 --> 00:20:01.000
Not affect our compute. They might work longer, but they do actually use roughly around the same amount of energy. So there have been studies, at least from Glasgow that showed that if you run lower but longer.

00:20:01.000 --> 00:20:08.000
You do use a similar amount of energy and in some cases you can actually slightly reduce the energy you use but very it's minuscule.

00:20:08.000 --> 00:20:29.000
The reason really why you should think about scaling down if you're not worried about intensity is to share the energy burden across two different time zones so you can have a wider view of the energy balance for a city or for your organization.

00:20:29.000 --> 00:20:43.000
And so when you use power is also as important as how much you use and sharing that out can be beneficial in terms of carbon savings because of marginal costs of things like district heating.

00:20:43.000 --> 00:20:53.000
Irrespective of the external sources of power.

00:20:53.000 --> 00:21:03.000
Right. Although we are a bit late, but I would also like to take one more question from Greg and then we move to the next

00:21:03.000 --> 00:21:10.000
Yeah, hi. The question is actually related to the question that Jan just asked and this slide that you're showing right now.

00:21:10.000 --> 00:21:18.000
And that is… you need to do two things probably to really benefit from this clock speed reduction.

00:21:18.000 --> 00:21:39.000
You also need to probably cool the individual processor chip at the chip level And when the clock frequency goes down, then the power dissipation goes down and you would need a cooling system which was reactive enough to reduce the coolant flow and reduce the heat exchange. And I don't think that exists right now because these servers are cooled

00:21:39.000 --> 00:21:43.000
At the server level rather than the individual processor chip level.

00:21:43.000 --> 00:21:48.000
Do you have any comments or thoughts on that?

00:21:48.000 --> 00:21:57.000
That's actually a very good point. So the cooling systems that we have have a flux that we pull into.

00:21:57.000 --> 00:22:03.000
And there's a steady rate of change of the amount of cooling.

00:22:03.000 --> 00:22:17.000
Now, what we actually see here is that potentially we are saving some power because the water is coming in there's a cycle of cooling and heating. And if the computer is have a lower temperature.

00:22:17.000 --> 00:22:22.000
Then the difference in water intake to water outtake is going to be slightly less.

00:22:22.000 --> 00:22:35.000
And if that is lowered, then the amount of energy required to cool the water to go back into the intake is lowered as well. So yes, you are correct. We do potentially also need to work on a more dynamic cooling.

00:22:35.000 --> 00:22:55.000
But reducing the amount of power that the CPU takes across the entire site will affect the rate of water cooling applied onto there. And therefore, it will reduce the amount of power because the water the difference in water temperature will be less than if we're running at max power.

00:22:55.000 --> 00:23:00.000
Okay, thanks. Yep.

00:23:00.000 --> 00:23:09.000
I am on the somatimos, so please just ping me any messages that you want to get me on them as well.

00:23:09.000 --> 00:23:13.000
Thanks a lot, Duane. It was a wonderful presentation as well. Very interesting.

00:23:13.000 --> 00:23:21.000
Discuss

