WEBVTT

1
00:00:00.000 --> 00:00:01.250
Volodymyr Svintozelskyi: Who paid me well off.


2
00:00:01.250 --> 00:00:03.169
Yann Coadou: Yes, perfect. You can go ahead.


3
00:00:03.770 --> 00:00:04.829
Volodymyr Svintozelskyi: Yes, thanks a lot.


4
00:00:04.960 --> 00:00:18.229
Volodymyr Svintozelskyi: So 1st of all, good morning to everyone my name is, and I'm going to talk about. I'm going to talk. All these people. I specified on this slide about sustainability studies of big data processing in real time for high energy physics.


5
00:00:18.380 --> 00:00:36.079
Volodymyr Svintozelskyi: The customer outline of my talk. So basically, it includes, 1st of all, the high level project at Valencia. I'm gonna say what it is, what is the objective and etc. And then I will move towards the motivation and strategy of the current studies of the power consumption studies in the pocket.


6
00:00:36.270 --> 00:00:40.020
Volodymyr Svintozelskyi: and then I will try to show the typical partner Assumption.


7
00:00:40.230 --> 00:00:43.320
Volodymyr Svintozelskyi: which you might have on a typical server.


8
00:00:43.790 --> 00:00:53.459
Volodymyr Svintozelskyi: And then we'll try to make some studies with different hardware, different utilization level, and also some accelerators.


9
00:00:53.650 --> 00:01:03.430
Volodymyr Svintozelskyi: So let's start with the 1st thing. So this is the description of the project. This is the transpassal project in between Atlas and Aloc that we have at Valencia.


10
00:01:04.103 --> 00:01:13.499
Volodymyr Svintozelskyi: So we have all the information about that on this slide. And the aim of the project is basically the benchmarking continued hardware and developing fast and high patient algorithms


11
00:01:13.650 --> 00:01:15.729
Volodymyr Svintozelskyi: with a reduced power consumption.


12
00:01:16.730 --> 00:01:21.110
Volodymyr Svintozelskyi: So some of the activities that specified on this slide.


13
00:01:21.580 --> 00:01:38.410
Volodymyr Svintozelskyi: so basically it includes development of different outputs, some faster processing with stock, photo, recycle, reconstruction, and etc. However, what is important for this talk is basically the measurement of software and hardware power consumption for high energy physics.


14
00:01:39.420 --> 00:01:54.020
Volodymyr Svintozelskyi: So this is the hardware that we have within this project, and with which we use to make all our studies. So basically, this is just a typical server. And the important thing is that you have to different Gpu cards.


15
00:01:54.833 --> 00:02:07.830
Volodymyr Svintozelskyi: Basically, one is a 5,000 and another is a 6,000. So a 5,000 is used to also be used for Hcv. In the system.


16
00:02:08.750 --> 00:02:23.490
Volodymyr Svintozelskyi: And then yep, we have some photos of that over there. And the last feature all set up is basically the Apc meter at rock video, which literally allows us measuring the overall power consumption of entire server.


17
00:02:23.680 --> 00:02:29.960
Volodymyr Svintozelskyi: because it's literally plugged in between the power socket and the server itself.


18
00:02:31.820 --> 00:02:50.850
Volodymyr Svintozelskyi: All right. So basically, if you take a look at some kind of existing studies, or some kind of extrapolations for the power consumption for the upcoming years, we will see that quite a drastic increase in the auto centers. Power consumption is expected in your, in your futures near future.


19
00:02:50.990 --> 00:03:03.779
Volodymyr Svintozelskyi: And the significant share of that is basically because of the it equipment. So the service itself. So they can conserve up to 50 or 60% of the power consumption for the typical data center. According to this


20
00:03:04.456 --> 00:03:16.690
Volodymyr Svintozelskyi: studies. And after that we have some green systems and lighting, etcetera, or the second. Most important is the green system, 65 to 45%.


21
00:03:18.960 --> 00:03:20.530
Volodymyr Svintozelskyi: Alright. So let's


22
00:03:20.940 --> 00:03:43.630
Volodymyr Svintozelskyi: define the tools that you can use in order to measure the our consumption of the service. So our scientists decided to use the test application, which is, which is Alan. So this is the software we use for high Level trigger one at Lhcp. So the reason why you choose this sorter is simply because I mean, we are the members of Lhcp collaboration. And it's much, much easier for us


23
00:03:44.090 --> 00:03:47.579
Volodymyr Svintozelskyi: to use the existing software. We use basically every day.


24
00:03:47.950 --> 00:03:57.650
Volodymyr Svintozelskyi: So this is the reason. So basically, there are some features of this sort of 1st of all, it can run on different various architectures.


25
00:03:57.850 --> 00:04:00.539
Volodymyr Svintozelskyi: It includes both CPU and Gpu.


26
00:04:01.020 --> 00:04:10.020
Volodymyr Svintozelskyi: It has a model of design that allows the execution of sequences sequences of the algorithm. So basically, the software is split it into the set of the algorithm.


27
00:04:10.380 --> 00:04:19.020
Volodymyr Svintozelskyi: Each one is, I mean developed to perform some specific task, and then we glue them together to form the sequence.


28
00:04:19.190 --> 00:04:24.289
Volodymyr Svintozelskyi: So we have this, this flexibility, let's say.


29
00:04:24.740 --> 00:04:40.869
Volodymyr Svintozelskyi: And the total number of the algorithm is order of 250. So also in the real conditions during the real data taken, we use order of 500 Nvidia Gpu video cards, 5,000 video cards.


30
00:04:41.640 --> 00:04:48.650
Volodymyr Svintozelskyi: and the last but not least, that's in the current status. They are done with the Lhcp simulation, someplace.


31
00:04:49.370 --> 00:04:51.940
Volodymyr Svintozelskyi: So what would be the strategy?


32
00:04:52.170 --> 00:04:53.369
Volodymyr Svintozelskyi: All the standards?


33
00:04:53.830 --> 00:05:04.549
Volodymyr Svintozelskyi: So basically the power consumption itself. It can be studied in with different tools, with different approaches. 1st of all, we can use some specific specific dedicated hard one.


34
00:05:04.560 --> 00:05:31.520
Volodymyr Svintozelskyi: So, for example, it can be the metal power distribution unit that gives us the overall power consumption of the entire setup. But it might be interesting to try to look at the individual components. So, for example, how much power is used only by the Gpu, or how much power is used only by the CPU. So all that stuff? So in order to answer these questions, we might use the device drivers. So, for example, in the new data case, we can use the Nvidia Dcgm


35
00:05:31.750 --> 00:05:38.259
Volodymyr Svintozelskyi: system that literally allows us to extract our consumption, consumed only by that piece of the hardware.


36
00:05:38.460 --> 00:05:46.950
Volodymyr Svintozelskyi: Well, the CPU is a bit more trickier about the CPU, the modern CPU. They have the system which is called the performance counters.


37
00:05:47.050 --> 00:06:02.810
Volodymyr Svintozelskyi: and by utilization of that counters we can estimate the approximate power consumption of the CPU course itself. So I believe many of you know that CPU package consists of many, many different calls basically can split the power consumption even per course.


38
00:06:04.055 --> 00:06:14.560
Volodymyr Svintozelskyi: So the goal of the current project is, basically try to answer this question, can we reduce the power consumption by optimizing the software and or harder


39
00:06:14.960 --> 00:06:15.960
Volodymyr Svintozelskyi: oops. Sorry?


40
00:06:16.918 --> 00:06:36.940
Volodymyr Svintozelskyi: So let's start with the 1st approach. Let's just try to measure the power consumption just dedicated hardware dedicated video. And we use the our processing for this stuff. And we process like a 300 million events.


41
00:06:37.260 --> 00:06:38.950
Volodymyr Svintozelskyi: not automatically start us.


42
00:06:39.150 --> 00:06:43.629
Volodymyr Svintozelskyi: And then if you upload the powerless star, you will see this kind of the curve


43
00:06:44.081 --> 00:06:52.949
Volodymyr Svintozelskyi: so you see that we have initial rise, and then almost instantly, so I don't know if you can see my character. But basically around


44
00:06:53.560 --> 00:06:58.790
Volodymyr Svintozelskyi: 200 seconds, we'll see another part, another rise to the bar consumption


45
00:06:59.340 --> 00:07:16.040
Volodymyr Svintozelskyi: that which will subsequently transform into the big big plateau. So this is how the overall behavior looks like. So we are wondering, like, what is the causes of this initial rise in the power consumption. And then like reduction.


46
00:07:16.300 --> 00:07:26.660
Volodymyr Svintozelskyi: because I mean from the nake point of view, it should be a plateau around something. The system is stable. So basically, the rest of the curve is more or less expected. But what is the reason of the plateau?


47
00:07:27.190 --> 00:07:33.050
Volodymyr Svintozelskyi: So for that, we tried to basically measure the partnership different companies.


48
00:07:33.450 --> 00:07:47.830
Volodymyr Svintozelskyi: So the 1st thing the 1st system we might use is called a Cpi interface. So this interface allows us with the modern motherboards to measure the power consumption using some onboard sensors directly on the motherboard.


49
00:07:48.070 --> 00:07:56.059
Volodymyr Svintozelskyi: And if we'll try to use this system, this basically device driver, what is important, that this is. The driver doesn't require any additional hardware.


50
00:07:56.200 --> 00:07:59.809
Volodymyr Svintozelskyi: If you try to measure the product consumption with the system, you'll see


51
00:08:00.270 --> 00:08:21.570
Volodymyr Svintozelskyi: quite a good match in between our machine with the additional hardware. And with this system, which basically means that in order to study the total bar consumption of the typical server. We don't really necessarily need some additional hardware. You might just need to write the existing tools already existing software.


52
00:08:21.890 --> 00:08:23.350
Volodymyr Svintozelskyi: It was in the or so.


53
00:08:24.909 --> 00:08:25.530
Volodymyr Svintozelskyi: Done.


54
00:08:25.630 --> 00:08:33.670
Volodymyr Svintozelskyi: We try to look also at the power consumption of the CPU's and the memory units so to do that


55
00:08:34.059 --> 00:08:41.199
Volodymyr Svintozelskyi: we use the CPU performance counters. So basically, there is a system which is called apple.


56
00:08:41.570 --> 00:08:49.669
Volodymyr Svintozelskyi: And this system allows us to extract our consumption of each individual component which is specified on the little picture on the left hand side.


57
00:08:50.748 --> 00:08:55.170
Volodymyr Svintozelskyi: So we have these studies. We decided to plot the 2 CPU's.


58
00:08:55.460 --> 00:09:03.105
Volodymyr Svintozelskyi: because on that side already have physically 2 cpus, 2 CPU. Calls. I'm sorry. And then, 2 different RAM boards.


59
00:09:03.680 --> 00:09:04.910
Volodymyr Svintozelskyi: skip soliciting.


60
00:09:05.320 --> 00:09:22.600
Volodymyr Svintozelskyi: And what we see is that the significant increase in the power consumption was only absurd for the one of the cpus. So basically, we were running cloud of sort, we're using only that CPU, so we like restricted execution only to that CPU, and that explains why the second CPU was idle during the whole process.


61
00:09:23.020 --> 00:09:25.043
Volodymyr Svintozelskyi: And then also,


62
00:09:25.760 --> 00:09:35.639
Volodymyr Svintozelskyi: I would say. Important thing is that we observed that the consumption caused by the RAM memory is negligible. With respect to all the other components.


63
00:09:35.960 --> 00:09:41.509
Volodymyr Svintozelskyi: however, still the CPU behavior, CPU power, consumption behavior is a plateau.


64
00:09:41.920 --> 00:09:47.750
Volodymyr Svintozelskyi: So it doesn't really explain why we have the speaking behavior in the beginning. So we continue with our studies.


65
00:09:48.030 --> 00:09:59.410
Volodymyr Svintozelskyi: and they switch to the Gpu. So again, now, servers have 2 different gpus, and we use this Vm system now consumption. And


66
00:09:59.670 --> 00:10:03.469
Volodymyr Svintozelskyi: again you see the portal. And again, the similar thing


67
00:10:03.800 --> 00:10:07.220
Volodymyr Svintozelskyi: support execution. We used only one of these views to restrict.


68
00:10:07.651 --> 00:10:13.840
Volodymyr Svintozelskyi: I mean, we basically stick that the execution 21 of these issues, and that's why the second one is idle


69
00:10:14.570 --> 00:10:15.660
Volodymyr Svintozelskyi: or still.


70
00:10:16.110 --> 00:10:19.649
Volodymyr Svintozelskyi: No, it doesn't explain to us why we have that rice in the beginning.


71
00:10:21.040 --> 00:10:27.449
Volodymyr Svintozelskyi: So you try to look at the overall system on this picture. I'm on the bottom of my slide.


72
00:10:27.980 --> 00:10:29.670
Volodymyr Svintozelskyi: So this


73
00:10:29.990 --> 00:10:42.880
Volodymyr Svintozelskyi: like bright color, I highlighted all the components that we basically measured already. So the only part which was kind of missing and which could potentially explain this kind of behavior was the fans of the cooling system.


74
00:10:43.270 --> 00:10:46.200
Volodymyr Svintozelskyi: So that's why we decided to basically


75
00:10:47.490 --> 00:10:58.660
Volodymyr Svintozelskyi: think how we can measure that. And in the end we ended up with attaching additional piece of the hardware, which is basically a small little device which allows us to measure the wind speed.


76
00:10:58.860 --> 00:11:16.649
Volodymyr Svintozelskyi: So now we have the funds fans in our system. So once the fans are spring up, they introduce much stronger wind, and then we can measure this wind by this with this device, and in this case we might be able to correlate like the increase in the cooling speed increase in the cooling


77
00:11:16.910 --> 00:11:22.700
Volodymyr Svintozelskyi: power consumption to the rise in the overall power consumption. So that's exactly what we did.


78
00:11:23.170 --> 00:11:51.790
Volodymyr Svintozelskyi: So on the right hand side you can see the dependences of the temperature or the CPU temperatures temperature with the time CPU and Gpu temperature in blue and orange and in black. That is the basically wind speed measured by that device. And you can see that this rise in the power. Consumption is exactly at the same point of time as the rise in overall power consumption. So now we are quite confident that this


79
00:11:52.300 --> 00:12:00.239
Volodymyr Svintozelskyi: initialize was literally because of the cooling system, because of the way how the cooling works on the surface.


80
00:12:01.626 --> 00:12:10.709
Volodymyr Svintozelskyi: So that's it. So at that point we understood the behavior of our consumption car, and then we tried to play around a little bit with that.


81
00:12:11.010 --> 00:12:18.769
Volodymyr Svintozelskyi: So the 1st thing we tried to do is we try to use different hardware to run our system and measure the part. Consumption is different hardware.


82
00:12:19.070 --> 00:12:23.619
Volodymyr Svintozelskyi: So this is used for that CPU and 2 Gpus, and you can see the results


83
00:12:24.060 --> 00:12:43.629
Volodymyr Svintozelskyi: on this plot. So basically, just to summarize these things. We observed that the faster the hardware is, the smaller overall power. Consumption will be, even though the instant instantaneous power consumption might be higher, because it's more power hardware, or still, if you integrate that.


84
00:12:43.630 --> 00:13:04.119
Volodymyr Svintozelskyi: it will be actually on the lower level, and it's can be easily seen. For the CPU case. The CPU has quite a low, instantaneous power consumption CPU only execution. But it takes very long time, and that's why, in the end, integrated energy consumption is much, much higher than the Gpu case.


85
00:13:04.550 --> 00:13:18.409
Volodymyr Svintozelskyi: So, apart from that, we also try to measure the different levels of the Gpu utilization. So the online software allows us to run the application with the threads.


86
00:13:18.810 --> 00:13:33.299
Volodymyr Svintozelskyi: And then each CPU thread is basically mapped to a cuda stream. To the some, let's say thread analog. But for the CPU. And then, if you change this number of sets and measure the part consumption to each of them we might


87
00:13:33.460 --> 00:13:38.200
Volodymyr Svintozelskyi: estimate what would be the effect of the utilization. So this is exactly what is shown on this slide.


88
00:13:38.580 --> 00:13:55.959
Volodymyr Svintozelskyi: and then for the low utilization for the blue carve you can see that the execution time is much, much larger, and then, if you try to use more threads into our hardware


89
00:13:56.120 --> 00:14:25.059
Volodymyr Svintozelskyi: actually supports. This might lead to the fact that you also will be higher in the power consumption. This is related to the fact that we have a complicated system or a scheduling. Because I mean all the additional threats that we create. They are not actually executed, but they are stored somewhere in the memory. And then the CPU does the thing which is called the contact switch. When it switches to one thread from one side to another thread. So it's not really multi-threading. But you are using this overhead.


90
00:14:26.310 --> 00:14:30.629
Volodymyr Svintozelskyi: For the for the let's say, scheduled


91
00:14:31.785 --> 00:14:43.079
Volodymyr Svintozelskyi: so, yeah, so this is for the Gpu sync. And then the last part of my talk is basically the PGA users should have PGA for the small power power consumption. So there I'm going status


92
00:14:43.400 --> 00:14:49.380
Volodymyr Svintozelskyi: which I am to offload some execution from the Gpu site from our Shorty one


93
00:14:49.965 --> 00:14:55.884
Volodymyr Svintozelskyi: to the dedicated hardware dedicated. So this is basically mostly about the tracking.


94
00:14:56.510 --> 00:15:09.529
Volodymyr Svintozelskyi: And what we see in that case is that we, I mean by uploading some tasks from from the Gpu to the dedicated hardware, we might see the decrease in the power consumption in the Gpu power consumption


95
00:15:10.030 --> 00:15:11.769
Volodymyr Svintozelskyi: which is shown on this slide.


96
00:15:11.950 --> 00:15:26.159
Volodymyr Svintozelskyi: So basically, we saw that distribution support increases by 30% and the same by the same order of the monitor, we have the decrease in the power consumption. However, the power, the studies with the overall consumption of apga plus Gpu, is still ongoing.


97
00:15:26.850 --> 00:15:28.780
Volodymyr Svintozelskyi: All right to summarize my talk.


98
00:15:29.224 --> 00:15:38.649
Volodymyr Svintozelskyi: So the 1st of all is that the what's pretty one should be an important metric for the high energy physics, especially when it comes to the development of the future software.


99
00:15:39.120 --> 00:15:43.089
Volodymyr Svintozelskyi: That sort of optimizations might potentially.


100
00:15:43.430 --> 00:15:48.960
Sukanya Sinha: Up on the 15 min. Actually, so it would be great to wrap up. Thank you.


101
00:15:52.750 --> 00:15:54.680
Volodymyr Svintozelskyi: Sorry. Could you please say again.


102
00:15:55.340 --> 00:16:00.950
Sukanya Sinha: your timer has been up for a while now, so it would be great to wrap up.


103
00:16:02.480 --> 00:16:05.639
Volodymyr Svintozelskyi: Oh, okay, I mean, just literally the last slide.


104
00:16:05.810 --> 00:16:19.940
Volodymyr Svintozelskyi: And I see a little timer in the zoom which says like minus 1 min, which might be kind of correlated to the fact that I'm getting on the last slide. But anyways, so okay, whatever. So this is the last slide. You can read the text yourself, and obtains your attention.


105
00:16:26.280 --> 00:16:35.489
Sukanya Sinha: Right? So thank you so much for this really interesting talk. Do we have any questions for Vladimir?


106
00:16:41.220 --> 00:16:43.330
Sukanya Sinha: Yes, Zach, please go ahead.


107
00:16:44.040 --> 00:16:49.199
Zach Marshall: Thanks, Lyra, this is an interesting toss. I I wanted to ask you about the Fpga usage.


108
00:16:50.043 --> 00:17:00.820
Zach Marshall: Because porting a lot of workflows to fpgas is usually pretty tough so you can get, you know, an algorithm or 2 with some work. But you're not gonna run like the whole reconstruction something like that.


109
00:17:01.823 --> 00:17:07.160
Zach Marshall: Have you looked into the balance of embodied carbon


110
00:17:07.369 --> 00:17:15.430
Zach Marshall: against operational carbon when it comes to those Fpgas to see if you're really doing something good for the environment. If you buy.


