WEBVTT

1
00:00:18.984 --> 00:00:38.984
For the bigger picture. You're all probably aware of the worldwide computing grid, WLCG, which operates 1.4 million CPU costs worldwide And this is at about 100 sites and the University of Manchester is one of them. And this is why we study it here as a small pilot study.


2
00:00:38.984 --> 00:00:51.984
That can later then be increased to the bigger picture. And… So why do we want to do this? We don't want to propose immediate changes.


3
00:00:51.984 --> 00:01:02.984
Because these are all very preliminary studies but we have to get familiar with these kind of studies, which the field isn't that familiar with yet, especially with life cycle assessments.


4
00:01:02.984 --> 00:01:15.984
And to measure the embodied carbon. And this is because a lot of the funding bodies are putting more and more effort into this and more and more and more prioritizing sustainability.


5
00:01:15.984 --> 00:01:22.984
And also because the European strategy as well as HECUP Also, I prioritized this more and more.


6
00:01:22.984 --> 00:01:29.984
So we really have to gain some expertise on this. I'll start with the description of the project.


7
00:01:29.984 --> 00:01:33.984
I was called tackling the energetic cost of computing at the University of Manchester.


8
00:01:33.984 --> 00:01:47.984
This was funded by a small seed funding grant here at the University of Manchester as part of the Sustainable Futures group of projects and it's a one-year project that has been ongoing since August of last year.


9
00:01:47.984 --> 00:01:57.984
And we try to do two things here. We estimate the energy consumption of running code that we use in the HEP community on our local clusters.


10
00:01:57.984 --> 00:02:01.984
And we do a lifecycle assessment of the hardware and the software together.


11
00:02:01.984 --> 00:02:09.984
And this is an interdisciplinary project between the engineering department where the PI is Ben Parks.


12
00:02:09.984 --> 00:02:15.984
At the University of Manchester and here the particle physics department where I'm the co-PI.


13
00:02:15.984 --> 00:02:33.984
And Lewis will talk about the software part of this a bit later. And he's an emphasis student working The hardware that we are studying is the Neuter cluster, so the WLCG is separated into these tiers from 0 to 3. Zero is the


14
00:02:33.984 --> 00:02:49.984
One of turn to three are the smallest kind of local clusters at the individual institutes like universities And that's why we start with the tier three here at the University of Manchester. There's eight nodes with 96 processors each, which can be used for batch jobs.


15
00:02:49.984 --> 00:03:00.984
And 20 nodes with 16 processors each, which can be both used interactively and dispatch jobs. And we're studying N96 processor on Nokia.


16
00:03:00.984 --> 00:03:04.984
It also has GPUs, but this is not part of the project for now.


17
00:03:04.984 --> 00:03:14.984
And the University of Manchester also has a tier two cluster, so the larger kind of cluster part of the grid, which is actually one of the largest in the UK.


18
00:03:14.984 --> 00:03:27.984
And interesting for this project, even though we didn't study this cluster in particular, is that the hardware from this tier two cluster is being passed down to the Neutta tier three cluster whenever the Tier 2 cluster gets upgrades.


19
00:03:27.984 --> 00:03:41.984
So we basically always one generation behind. Cutting edge on the tier two.


20
00:03:41.984 --> 00:03:42.984
You can do with this.


21
00:03:42.984 --> 00:03:52.984
And the software that we are studying is Herwig. Why? Because this is one of the most widely used Moncarlo generators. And if you look at writing flopter from the compute budget of Atlas and Monte Carlo generation is a very large part of this.


22
00:03:52.984 --> 00:03:53.984
Thanks.


23
00:03:53.984 --> 00:03:58.984
Slice at the top then. And we're studying the Treyan process with this, generating it.


24
00:03:58.984 --> 00:04:08.984
And this has two stages, the integration process, which we run single threaded. It can also be run multi-threaded, but it's interesting to compare the two different compute paradigms here.


25
00:04:08.984 --> 00:04:24.984
And the event generation, which is from multi-threaded. And with this over to Andrews. So hi there Good morning, everyone. As Toiles mentioned, my name is Luis Miguel, and I've been working in Herrig here at the University of Manchester.


26
00:04:24.984 --> 00:04:37.984
So let's get straight into it. So there are many ways of profiling software, the energy of software. For our students, the most important ones were RAPL, CodeCarbon, and Prometheus.


27
00:04:37.984 --> 00:05:00.984
Well, Rubble asked the previous… presentation to us. It's an interprocessor feature that allows for almost real-time monitoring of the CPU and RAM energy consumption COPE carbon just provides a simple interface to access rubble via rubble Python library. Prometheus interfaces with power supplies so well power plugs sorry


28
00:05:00.984 --> 00:05:07.984
That allows for measurements of power supplies in the for the nose in the Noether cluster, for example.


29
00:05:07.984 --> 00:05:22.984
But there are also another tools like panda using Atlas, hip score in the near future probably green algorithms that is online calculator that can be integrated with slurry systems as well.


30
00:05:22.984 --> 00:05:37.984
So the neutered cluster consists of an array notes, the power distribution unit or PDU fits these nodes via the PSU and the PSU then distributes the energy to each component in a working node.


31
00:05:37.984 --> 00:05:43.984
So what we are actually measuring with Prometheus is the PSU of each individual node.


32
00:05:43.984 --> 00:05:47.984
We are indirectly measuring all the energy that comes into the node.


33
00:05:47.984 --> 00:05:55.984
With gold carbon and travel, we are only measuring the CPU energy directly. So keeping that in mind when we look at the results.


34
00:05:55.984 --> 00:06:07.984
Talking about results. So we benchmarked the drill down process. In this case in your screen, you're seeing two protons going to two electrons And we use an Intel Xeon Gold 50 to 20.


35
00:06:07.984 --> 00:06:25.984
That is the 96 threaded CPU here at the tier three cluster. And as Tawaii has mentioned, we differentiated between two steps, integration and generation So we found that the relationship between the number of events generated and the energy consumed


36
00:06:25.984 --> 00:06:42.984
To very good approximation linear. And we also found that the power consumption was pretty much constant during the whole simulation at least up to 10 to the power of eight events But we do expect some thermal effects to kick in at some point.


37
00:06:42.984 --> 00:06:48.984
But well, now we have this energy consumption but what do we actually do with it? How do we convert it to CO2 emissions?


38
00:06:48.984 --> 00:06:54.984
Well, that is a simple and at the same time complex question.


39
00:06:54.984 --> 00:07:03.984
So we use a scaling factor. That this scaling factor called carbon efficiency is really location dependent so If you do an experiment, for example, in the UK here.


40
00:07:03.984 --> 00:07:15.984
You will get totally different results for your CO2 emissions if you do it, for example, in the US or Germany or france In France, it's going to be much much less impactful, for example.


41
00:07:15.984 --> 00:07:26.984
So this is something to obviously keep in mind, especially when working with worldwide collaboration like like percent.


42
00:07:26.984 --> 00:07:40.984
So yeah, so now we have the CO2 emissions. We can do something with it. We can do extrapolations and see what how will happen. I don't know, for example, FCCS, LAC scales for I don't know, 10 to a power of 16 events, for example.


43
00:07:40.984 --> 00:07:46.984
And we found that we found that given that these are order of magnitude estimations.


44
00:07:46.984 --> 00:07:54.984
Around a million tons of CO2 would be emitted in FCC conditions for central mass synergy of 100 tvs.


45
00:07:54.984 --> 00:07:59.984
This is based on a almost linear model of the two protons going to two electrons.


46
00:07:59.984 --> 00:08:15.984
For example some of the other processes will have slightly different relationships but this is something to keep in mind, especially for thermal effects because those are that's how it's how looks into the data.


47
00:08:15.984 --> 00:08:22.984
And a fun fact that we found is we did the same exact procedure in a laptop.


48
00:08:22.984 --> 00:08:35.984
And we found that surprisingly, the laptop seems to be a little bit more efficient in terms of event generated per energy consumed at least up to 10 to a power of five events.


49
00:08:35.984 --> 00:08:49.984
And also, it seems to be faster. As seen in the right, up to 10 to the power of 4 events. But we have to keep in mind that these are really not many events and tend to have a power of four events at a really long number for any


50
00:08:49.984 --> 00:09:06.984
Particle physics applications, so keep that in mind. Our little caveat is that We only produce data for 10 to the power of five events in the laptop. So keep in mind that beyond that point, the model could fail and will fail because of thermal effects so as always


51
00:09:06.984 --> 00:09:14.984
We need more data, guys. Okay, thanks. And I'll talk about the lifecycle assessment, which is the second part of this project.


52
00:09:14.984 --> 00:09:22.984
And this was mostly performed by our colleagues from the engineering department and especially by Nico, who is a post of that.


53
00:09:22.984 --> 00:09:35.984
So what is a lifecycle assessment? It considers both the embodied carbon in the production of the of the hardware as well as running the software over extended amount of time.


54
00:09:35.984 --> 00:09:46.984
We studied three different test cases here at the university, which are desktops with screen laptops and especially the HPC in this case all neutral T3.


55
00:09:46.984 --> 00:09:54.984
This all follows the ESO standards. And if you want to do something like this, please consider doing this.


56
00:09:54.984 --> 00:09:58.984
Of course then can be very well compared to other lifecycle assessments.


57
00:09:58.984 --> 00:10:06.984
The software is the commonly used software for this, which is called Sima Pro and the Ecoinvent database.


58
00:10:06.984 --> 00:10:15.984
As Lou said, the energy consumption of heroin gets linear with the number of events. So we can just use the power consumption as one input.


59
00:10:15.984 --> 00:10:25.984
The other input is then the runtime. And then we also need all the different hardware components and what kind of material they are built from.


60
00:10:25.984 --> 00:10:39.984
And this is actually quite a challenging part. Because even the vendors don't necessarily know this information. Even if you buy two products with the same product ID, you don't necessarily get always the same hardware. What you get is just hardware


61
00:10:39.984 --> 00:10:45.984
That can perform up to certain specifications, but it's not clear that it was produced in the same way.


62
00:10:45.984 --> 00:10:54.984
So this is very complicated. Another thing you have to have a micro lifecycle assessment is you have to define your system boundaries very well.


63
00:10:54.984 --> 00:11:09.984
So, for example, in our case, we do not consider the end of life so the either recycling or going to waste of the components Because here at the University of Manchester, the end of life is managed by a contractor


64
00:11:09.984 --> 00:11:20.984
Which resells components if they're still usable. And according to the university, this saves about 450 tons of CO2 equivalents.


65
00:11:20.984 --> 00:11:37.984
During the latest teaching year. Here are some very preliminary results. You should really look at these more with the lens of this is what we can study rather than these are exact values. So please take me with a grain of salt.


66
00:11:37.984 --> 00:12:01.984
But one interesting thing that was found out is that the global warming impact, mostly due to CO2, emission is dominated by the electricity consumption by the HPC. So during the runtime Whereas for desktop and laptops computers, it's usually dominated by the production, so by the embodied carbon.


67
00:12:01.984 --> 00:12:14.984
And another thing that we studied is different replacement policies. So if you replace your hardware every five, seven or nine years, in this case, the HPC hardware for the Neuta cluster.


68
00:12:14.984 --> 00:12:21.984
And they're… if you replace it more often in this particular scenario, so five years.


69
00:12:21.984 --> 00:12:27.984
Your consumptions are higher than if you replace it. Every nine years.


70
00:12:27.984 --> 00:12:41.984
And… want to bring this into this bit of a bigger picture. We're part of a broader effort in HEP, as you now all know from this workshop, hopefully.


71
00:12:41.984 --> 00:12:53.984
And the WLCG, so the computing grid already organized a workshop on this and there was a dedicated session also in the latest WXCG workshop.


72
00:12:53.984 --> 00:12:56.984
So please have a look at that if you want to.


73
00:12:56.984 --> 00:13:10.984
And then… And there will also be a WLCG Sustainability Forum be set up. So if you want to contribute to this effort, please consider joining this.


74
00:13:10.984 --> 00:13:29.984
And just quickly our conclusion. So one lesson that we learned is that sustainability analysis and especially lifecycle assessment can be really hard because you're not necessarily have all the information that you need to do this. And so you have to do some estimates. You can't really get around this.


75
00:13:29.984 --> 00:13:51.984
It's very important to do this. Because funding agencies put more and more prioritize this more and more. So we have to be ready for this And even the profiling of power consumption can be quite complex because you really have to consider what you're actually measuring. Is it only CPU? Is it the whole chassis of your


76
00:13:51.984 --> 00:14:03.984
Of your server. And in terms of outcomes of this project, so we increased awareness already here at the University of Manchester, for example, with some guest presentations at the C++ course for sustainable software.


77
00:14:03.984 --> 00:14:18.984
And longer term, we plan to inform the procurement strategies here at the University of Manchester. We're already in contact with the IT services and research IT who are actually very interested in this kind of effort.


78
00:14:18.984 --> 00:14:27.984
And even longer term, we really try to integrate this into the broader context of the WNCG and Hope.


79
00:14:27.984 --> 00:14:30.984
To work together with maybe some of you on this future.


80
00:14:30.984 --> 00:14:33.984
So any thanks. Thank you.


81
00:14:33.984 --> 00:14:45.984
Thanks for the talk, Louis and Willis. It was nice to see a tag team approach. We already have a hand raised by Dwayne, so please go ahead.


82
00:14:45.984 --> 00:15:04.984
Hi, that was a very nice talk. Just a very quick question, specifically on the point you based on this last slide. So you said that you were looking to try to apply pressure on your procurement process In what ways do you see you can change the procurement process for the better?


83
00:15:04.984 --> 00:15:17.984
Yeah, so one thing, as I said, these are very preliminary studies, but one thing that one can really change here is the replacement policies.


84
00:15:17.984 --> 00:15:22.984
So it really matters if you replace your hardware every five, seven or nine years.


85
00:15:22.984 --> 00:15:29.984
This is what we found here is five years has a larger impact, but it really depends on what kind of hardware. It can be the other way around.


86
00:15:29.984 --> 00:15:40.984
Where if you buy more efficient hardware. It can actually be more beneficial to change your hardware more often so that you're always on the cutting edge of efficiency.


87
00:15:40.984 --> 00:15:52.984
So things like that. Can really be uh yeah used as an input for that.


88
00:15:52.984 --> 00:15:53.984
Thank you.


89
00:15:53.984 --> 00:15:54.984
Thanks.


90
00:15:54.984 --> 00:16:02.984
Perfect. And That brings us to the exact time slot of the session.


91
00:16:02.984 --> 00:16:09.984
I would suggest you guys please connect to the matter most where we can continue the discussions on this very exciting topic.


92
00:16:09.984 --> 00:16:14.984
And then I hand over to Jan, whose connection hopefully is fine.


93
00:16:14.984 --> 00:16:15.984
For the continuation of the session.


94
00:16:15.984 --> 00:16:16.220
Yes. In case Kenya. Yeah, sorry about this.


