WEBVTT

1
00:00:02.414 --> 00:00:18.463
Luis Villar & Tobias Fitschen: Okay, thanks. Yeah, we can go ahead. So yeah, this leads actually quite well on from the last question, because we did study embodied carbon and operational carbon. And yeah, we were doing this at our local computing cluster here at the University of Manchester.


2
00:00:19.319 --> 00:00:23.134
Luis Villar & Tobias Fitschen: For the bigger picture. Like your


3
00:00:23.514 --> 00:00:43.123
Luis Villar & Tobias Fitschen: or probably aware of the worldwide computing grid, the Wlcg. Which operates 1.4 million CPU cores worldwide, and this has had about 100 sites, and the University of Manchester is one of them, and this is why we study it here as a small pilot study that can later then be increased to the bigger picture.


4
00:00:43.234 --> 00:00:43.979
Luis Villar & Tobias Fitschen: And


5
00:00:45.834 --> 00:00:55.694
Luis Villar & Tobias Fitschen: so why do we want to do this? We don't want to propose immediate changes, because these are all very preliminary studies, but we


6
00:00:56.084 --> 00:01:25.894
Luis Villar & Tobias Fitschen: have to get familiar with these kind of studies which the field isn't that familiar with yet, especially with life cycle assessments, and to measure the embodied carbon. And this is because a lot of the funding bodies are putting more and more effort into this and more and more prioritizing sustainability, and also because the European strategy, as well as hiccup, also prioritizes more and more. So we really have to gain some expertise on this.


7
00:01:26.649 --> 00:01:29.334
Luis Villar & Tobias Fitschen: I'll start with the description of the product


8
00:01:29.634 --> 00:01:34.023
Luis Villar & Tobias Fitschen: I would call tackling the energetic cost of computing at the University of Manchester.


9
00:01:34.244 --> 00:01:47.173
Luis Villar & Tobias Fitschen: This was funded by a small seed funding grant here at the University of Manchester as part of the sustainable Futures group of projects. And it's a 1 year project that has been ongoing since August of last year.


10
00:01:47.414 --> 00:02:01.783
Luis Villar & Tobias Fitschen: and we try to do 2 things here. We estimate the energy consumption of running code that we use in the Hub community on our local clusters, and we do a lifecycle assessment of the hardware and the software together.


11
00:02:02.064 --> 00:02:15.373
Luis Villar & Tobias Fitschen: And this is inter an interdisciplinary project between the engineering department where the Pi is Ben Parks also here at the University of Manchester, and here the particle physics department where I'm the co-pi.


12
00:02:15.484 --> 00:02:22.694
Luis Villar & Tobias Fitschen: and Louis will talk about the software part of this a bit later, and he's an emphas student working


13
00:02:25.164 --> 00:02:40.594
Luis Villar & Tobias Fitschen: the hardware that we are studying is the neuter cluster. So the Wlcg is separated into these tiers from 0 to 3. 0 is the local one at certainty. 3 are the smallest kind of local clusters at the individual institutes like universities.


14
00:02:40.594 --> 00:02:59.324
Luis Villar & Tobias Fitschen: And that's why we start with the tier 3 here at the University of Manchester. There's 8 nodes with 96 processors each, which can be used for batch jobs and 20 nodes with 16 processes each, which can be both used interactively and just batch jobs. And we're studying a 96 processor node here.


15
00:02:59.474 --> 00:03:04.233
Luis Villar & Tobias Fitschen: It also has Gpus. But this is not part of the project, for now.


16
00:03:04.394 --> 00:03:31.034
Luis Villar & Tobias Fitschen: and the University of Manchester also has a tier, 2 cluster, so the like larger kind of cluster part of the grid, which is actually one of the largest in the Uk. And interesting for this project, even though we didn't study this cluster in particular is that the hardware from this tier 2 cluster is being passed down to the newer tier. 3 cluster. Whenever the tier 2 cluster gets upgrades. So we're basically always one generation behind the


17
00:03:31.254 --> 00:03:32.774
Luis Villar & Tobias Fitschen: cutting edge on the tier 2.


18
00:03:33.964 --> 00:03:41.824
Luis Villar & Tobias Fitschen: And the software that we are studying is Herwig. Why? Because this is one of the most widely used Monte Carlo generators, and if you look.


19
00:03:41.824 --> 00:03:43.044
Greg Hallewell: What you can do is this.


20
00:03:43.524 --> 00:03:50.604
Luis Villar & Tobias Fitschen: Writing plot there, and from the compute budget of Atlas and Monte Carlo. Generation is a very large part of this


21
00:03:50.854 --> 00:03:53.554
Luis Villar & Tobias Fitschen: slice at the top there, and


22
00:03:53.794 --> 00:04:12.803
Luis Villar & Tobias Fitschen: and we're studying the process with this generating it. And this has 2 stages. The integration process which we run single threaded. It can also be on multi-threaded. But it's interesting to compare the 2 different compute paradigms here, and the event generation which is run multi-threaded.


23
00:04:12.954 --> 00:04:18.214
Luis Villar & Tobias Fitschen: And with this I'll give over to Ruth. So Hi, there!


24
00:04:18.514 --> 00:04:24.844
Luis Villar & Tobias Fitschen: Good morning, everyone as toys mentioned. My name is Luis Miguel, and I've been working in her week here at the University of Manchester.


25
00:04:25.384 --> 00:04:37.314
Luis Villar & Tobias Fitschen: So let's get straight into it. So there are many ways of providing software, the energy of software. And for our students the most important ones were rappel called carbon, and Prometheus.


26
00:04:37.694 --> 00:04:53.973
Luis Villar & Tobias Fitschen: Well, Rappel, as the previous presentation told us, into processor feature that allows for almost real time monitoring of the CPU and RAM energy consumption. Cold carbony just provides a simple interface to access. Rappel via


27
00:04:54.374 --> 00:05:07.464
Luis Villar & Tobias Fitschen: Python Library. Prometheus interfaces with power, supplies, so where power plugs. Sorry that allows for measurements of power supplies in the for the nodes in the no extra cluster, for example.


28
00:05:08.364 --> 00:05:21.584
Luis Villar & Tobias Fitschen: But there are also another tools like a panda using atlas hipscore in the near future, probably, and green algorithms. That is a line calculator that can be integrated with slurb systems as well.


29
00:05:22.414 --> 00:05:27.584
Luis Villar & Tobias Fitschen: So the noder cluster consists consists of an array nodes


30
00:05:27.964 --> 00:05:37.013
Luis Villar & Tobias Fitschen: the power distribution unit or Pdu feeds these nodes via the psu and the Psu, then distributes the energy to each component in a working node.


31
00:05:38.004 --> 00:05:47.544
Luis Villar & Tobias Fitschen: So what we are actually measuring with Prometheus is the psu of each individual node. So we are indirectly measuring all the energy that comes into the node


32
00:05:48.134 --> 00:05:55.593
Luis Villar & Tobias Fitschen: with cold carbon and drople. We are only only measuring the CPU energy directly. So keeping that in mind when we look at the results


33
00:05:56.264 --> 00:05:57.704
Luis Villar & Tobias Fitschen: talking about results.


34
00:05:57.824 --> 00:06:18.064
Luis Villar & Tobias Fitschen: So we benchmarked the drill down process. In this case in your screen, you're seeing 2 protons going to 2 electrons, and we use an intel Xion gold, 50 to 20, that is, the 96 threaded CPU. Here at the tier, 3 nodal cluster, and as Tobias mentioned, we differentiate between 2 steps, integration and generation.


35
00:06:18.874 --> 00:06:28.154
Luis Villar & Tobias Fitschen: So we found that the relationship between the number of events generated and the energy consumed was to very good approximation linear.


36
00:06:28.454 --> 00:06:40.013
Luis Villar & Tobias Fitschen: and we also found that the power consumption was pretty much constant during the whole simulation, at least up to 10 to the power of 8 events. But we do expect some thermal effects to kick in at some point


37
00:06:40.794 --> 00:06:41.924
Luis Villar & Tobias Fitschen: a.


38
00:06:42.764 --> 00:06:50.183
Luis Villar & Tobias Fitschen: But well, now we have this energy consumption. But what do we actually do with it? How do we convert it to Co. 2 emissions? Well.


39
00:06:50.614 --> 00:07:10.983
Luis Villar & Tobias Fitschen: that is a simple and at the same time complex question. So we use a scaling factor that this scaling factor called carbon efficiency is really location dependent. So if you do an experiment, for example, in the Uk. Here, you will get total different results for your Co. 2 emissions if you do it, for example, in the Us. Or Germany or France


40
00:07:11.134 --> 00:07:20.324
Luis Villar & Tobias Fitschen: in France, is going to be much, much less impactful, for example. So this is something to obviously keep in mind, especially if you, when working with


41
00:07:20.454 --> 00:07:24.314
Luis Villar & Tobias Fitschen: worldwide collaboration, like like of stern.


42
00:07:26.914 --> 00:07:40.024
Luis Villar & Tobias Fitschen: So, yeah, so now we have the Co. 2 emissions. We can do something with it. We can do extrapolations and see what how will happen? I don't know. For example, scales for a I don't know. 10 to a power of 16 events, for example.


43
00:07:40.564 --> 00:07:54.324
Luis Villar & Tobias Fitschen: and we found that, given that these are order of magnitude estimations around a million tons of Co. 2 would be emitted in Fcc conditions for central mass energy of 100 Tvs.


44
00:07:54.754 --> 00:08:02.204
Luis Villar & Tobias Fitschen: This is based on a almost linear model of the P. 2 protons going to 2 electrons. But, for example, some of the other


45
00:08:02.734 --> 00:08:05.723
Luis Villar & Tobias Fitschen: processes will have slightly different


46
00:08:05.834 --> 00:08:15.653
Luis Villar & Tobias Fitschen: relationships. But this is something to keep in mind, especially for the thermal effects, because those are. That's how it looks into the data.


47
00:08:16.134 --> 00:08:22.403
Luis Villar & Tobias Fitschen: And a fun fact that we found is we did the same exact procedure in a laptop.


48
00:08:23.214 --> 00:08:35.454
Luis Villar & Tobias Fitschen: and we found that surprisingly, the laptop seems to be a little bit more efficient in terms of in terms of a event generated per energy consumed at least up to 10, to the power of 5 events.


49
00:08:36.304 --> 00:08:39.484
Luis Villar & Tobias Fitschen: and also, it seems to be faster.


50
00:08:39.674 --> 00:08:52.524
Luis Villar & Tobias Fitschen: as seen in the right up to 10, to the power of 4 events. But we have to keep in mind that these are really not many events, and 10 to the power of 4 events are a really long number for any particle physics application. So keep that in mind.


51
00:08:52.724 --> 00:08:54.604
Luis Villar & Tobias Fitschen: Our little caveat is that


52
00:08:55.564 --> 00:09:06.443
Luis Villar & Tobias Fitschen: we only produce data for 10 to the power of 5 events in the laptop. So keep in mind that beyond that point the model could fail and will fail because of thermal effects. So as always.


53
00:09:06.654 --> 00:09:08.274
Luis Villar & Tobias Fitschen: we need more data guys.


54
00:09:09.604 --> 00:09:22.424
Luis Villar & Tobias Fitschen: Okay, thanks. And now I'll talk about the lifecycle assessment, which is the second part of this project, and this was mostly performed by the by, our colleagues from the engineering department, and especially by Nico, who is a post of that


55
00:09:23.523 --> 00:09:30.824
Luis Villar & Tobias Fitschen: so what is a lifecycle assessment? It considers both the embodied carbon and the production of the


56
00:09:31.044 --> 00:09:46.894
Luis Villar & Tobias Fitschen: of the hardware, as well as running the software over extended amount of time. We studied 3 different test cases here at the University, which are desktops with screen laptops, and especially the Hpc. In this case, or New 33.


57
00:09:47.054 --> 00:09:59.144
Luis Villar & Tobias Fitschen: This all follows the Eso standards, and if you want to do something like this, please consider doing this because then it can be very well compared to other lifecycle assessments.


58
00:09:59.254 --> 00:10:06.334
Luis Villar & Tobias Fitschen: The software is the commonly used software for this, which is called Cma pro and the Eco. Invent database.


59
00:10:07.274 --> 00:10:18.453
Luis Villar & Tobias Fitschen: As Louis said, the energy consumption of heir gets linear with the number of events. And so we can just use the power consumption as one, input the other input is in the runtime.


60
00:10:18.494 --> 00:10:47.943
Luis Villar & Tobias Fitschen: And then we also need all the different hardware components and what kind of material they are built from. And this is actually quite a challenging part, because even the vendors don't necessarily know this information. Even if you buy 2 products of the same product, Id, you don't necessarily get always the same hardware. What you get is just hardware that can perform up to certain specifications, but it's not clear that it was produced in the same way. So this is very complicated.


61
00:10:48.044 --> 00:11:04.124
Luis Villar & Tobias Fitschen: Another thing. You have to have a micro lifecycle assessment is you have to define your system boundaries? Very well. So, for example, in our case we do not consider the end of life, so the either recycling or going to waste of the components.


62
00:11:04.124 --> 00:11:20.654
Luis Villar & Tobias Fitschen: because here at the University of Manchester the end of life is managed by a contractor which resells components if they're still usable. And according to the university, this saves about 450 tons of Co. 2 equivalents


63
00:11:20.654 --> 00:11:30.283
Luis Villar & Tobias Fitschen: during the latest teaching. Yeah, here are some very preliminary results. You should really look at these more


64
00:11:30.314 --> 00:11:46.894
Luis Villar & Tobias Fitschen: with the lens of this is what we can study rather than these are exact values. So please take it with a grain of salt. But one interesting thing that was found out is that the global warming impact, mostly due to Co. 2 emission


65
00:11:47.341 --> 00:12:01.213
Luis Villar & Tobias Fitschen: is dominated by the electricity consumption by the Fpc. So during the runtime, whereas for desktop and laptops computers, it's usually dominated by the production. So by the embodied carbon


66
00:12:01.674 --> 00:12:02.409
Luis Villar & Tobias Fitschen: and


67
00:12:03.394 --> 00:12:14.333
Luis Villar & Tobias Fitschen: another thing that we studied is different replacement policies. So if you replace your hardware every 5, 7, or 9 years. In this case the Hpc. Hardware for the newer cluster.


68
00:12:14.484 --> 00:12:25.113
Luis Villar & Tobias Fitschen: and if you replace it more often in this particular scenario, so 5 years, your consumptions are higher than if you replace it


69
00:12:25.424 --> 00:12:28.984
Luis Villar & Tobias Fitschen: every 9 years and


70
00:12:30.244 --> 00:12:56.613
Luis Villar & Tobias Fitschen: want to bring this into a bit of a bigger picture. We're part of a broader effort in Hep, as you now all know, from this workshop hopefully, and the Wxg. So the computing grid already organized a workshop on this, and there was a dedicated session also in the latest Wxg workshop. So please have a look at that if you want to.


71
00:12:56.964 --> 00:13:10.113
Luis Villar & Tobias Fitschen: and then there will also be a Wcg. Sustainability forum be set up. So if you want to contribute, contribute to this effort, please consider joining this.


72
00:13:11.244 --> 00:13:32.233
Luis Villar & Tobias Fitschen: and that's quickly our conclusion. So one lesson that we learned is that sustainability analysis, and especially lifecycle assessment, can be really hard because you not necessarily have all the information that you need to do this. And so you have to do some estimates. You can't really get around this. It's very important to do this.


73
00:13:32.374 --> 00:13:36.314
Luis Villar & Tobias Fitschen: because funding agencies put more and more


74
00:13:37.324 --> 00:13:41.134
Luis Villar & Tobias Fitschen: prioritize this more and more. So we have to be ready for this.


75
00:13:41.404 --> 00:13:53.043
Luis Villar & Tobias Fitschen: And even the profiling of power consumption can be quite complex. Because you really have to consider what you're actually measuring, is it only, CPU? Is it the whole chassis of your of your server?


76
00:13:53.324 --> 00:14:04.103
Luis Villar & Tobias Fitschen: And in terms of outcomes of this project? So we increased awareness already here at the University of Manchester, for example, with some guest presentations at the C. Plus plus course for sustainable software


77
00:14:04.344 --> 00:14:19.094
Luis Villar & Tobias Fitschen: and longer term, we plan to inform the procurement strategies here at the University of Manchester, and we're already in contact with the It services and research it, who are actually very interested in this kind of effort


78
00:14:19.274 --> 00:14:33.324
Luis Villar & Tobias Fitschen: and even longer term. We really try to integrate this into the broader context of the Wcg and hope to work together with maybe some of you on this in future. So many thanks. Thank you.


79
00:14:33.714 --> 00:14:42.464
Sukanya Sinha: Thanks for the talk, Louis, and it was nice to see a tag team approach. We already have a hand raised by Duane, so please go ahead.


80
00:14:46.364 --> 00:15:04.073
Dwayne Spiteri: Hi! That's very nice talk. Just a very quick question, or specifically on a point you raised in this last slide. So you said that you were looking to try to in apply pressure on your procurement process. In what ways do you see? You can change the procurement process for the better.


81
00:15:05.134 --> 00:15:12.770
Luis Villar & Tobias Fitschen: Yeah. So one thing as I said, these are very preliminary studies. But one thing that one can really


82
00:15:13.304 --> 00:15:40.993
Luis Villar & Tobias Fitschen: change. Here is the replacement policies. So it really matters if you replace your hardware every 5, 7, or 9 years. This is what we found here is 5 years has a larger impact. But it really depends on what kind of hardware it can be the other way around. Where, if you buy more efficient hardware, it can actually be more beneficial to change your hardware more often, so that you're always on the cutting edge of efficiency.


83
00:15:41.597 --> 00:15:43.130
Luis Villar & Tobias Fitschen: So things like that,


84
00:15:43.934 --> 00:15:48.473
Luis Villar & Tobias Fitschen: can really be yeah, used as an input for that.


85
00:15:52.933 --> 00:15:53.603
Dwayne Spiteri: Thank you.


86
00:15:54.683 --> 00:15:56.373
Sukanya Sinha: Perfect, and


87
00:15:56.783 --> 00:16:16.220
Sukanya Sinha: that brings us to the exact time slot of the session. So I would suggest to you guys, please connect to the matter most where we can continue the discussions on this very exciting topic, and then I hand over to Jan, whose connection hopefully is fine for the continuation of the session.


