WEBVTT

1
00:00:02.065 --> 00:00:16.065
Yes, so…


2
00:00:16.065 --> 00:00:17.065
Yes, perfect.


3
00:00:17.065 --> 00:00:45.065
Okay, they should be full screen now. Okay, thank you for the opportunity to speak. So I will discuss the EPN pharma so is the online processing farm for the ALICE experiment. I'll just jump in because the slides are a bit dense. So preliminary context so you will know that in grant three um even the um


4
00:00:45.065 --> 00:01:14.065
Let luminosity performance of the LHC accelerator has increased So they can ship an interaction rate of 50 kilohertz uh for uh sustained times like one or two hours. And to cope with this increased rate, the ALS detector underwent a major upgrade during the LS2.


5
00:01:14.065 --> 00:01:41.065
The major changes were the TPC, which now is a GMTPC and the pixel detector which is the 13 gigapixel very very granular map silicon detector and also we abandoned the trigger without in favor of continuous readout that implies implied the new computing model both in software and hardware


6
00:01:41.065 --> 00:01:50.065
And extensive views of GPUs. And GPUs are used in Alice for producing physics directly online.


7
00:01:50.065 --> 00:02:17.065
And also to calibrate ador of course otherwise there would be no real physics So that's, I think, at CERN is a unique feature. We have 350 nodes with a total of 2,800 gpus And this is the operational of the data taking the experiment cannot function without these GPUs processing online.


8
00:02:17.065 --> 00:02:40.065
And in round three, we expect to get uh luminosity of roughly six inverse nanobarn and the overall budget will be 13 nanobar inverse nanobaron for run t plus run 4.


9
00:02:40.065 --> 00:03:10.065
Okay, for the event processing basically we have now continuous readout so we take data in snapshots time frames uh this is not like free running. So free running you you we have a strobe that drives the electronics but the strobe takes this snapshot. We can set the length of the snapshot now it's 32 LHC orbits which is almost three milliseconds


10
00:03:10.065 --> 00:03:31.065
And the APN farms, the GPU farm performs tracking, calibration, reconstruction and compression synchronously so while data is being taken. Most of the GPU compute is, of course, for the TPC tracking and we have some CPU only nodes for for calibration


11
00:03:31.065 --> 00:03:43.065
And we also use efficient methods to compress this data because there is no way we can save this data in raw form anywhere.


12
00:03:43.065 --> 00:03:56.065
Basically, one characteristic is also that this the ali software is the same for synchronous and asynchronous. So for online and offline.


13
00:03:56.065 --> 00:04:02.065
So basically the same software is run in the two different scenarios.


14
00:04:02.065 --> 00:04:22.065
So this is a rather dense slide that I will just try to summarize it a bit. So we have a huge data stream coming from the detectors that is read out via FPGA cards. So we have basically 3.5 terabyte per seconds before zero suppression that goes into the


15
00:04:22.065 --> 00:04:41.065
Without nodes and here is zero suppressed basically this is mostly tpc And out of here we go into the GPU farm with a stream of 700 gigabytes per second. I'm talking about lead lead of course which is the most relevant case.


16
00:04:41.065 --> 00:04:53.065
In the EPN pharma, the events are built they are not real events. These are, as I said, just time windows that are superposed and matched.


17
00:04:53.065 --> 00:05:23.065
And all this is done on the GPU cards. So there's basically no no cpus involved except for calibration And then we have online compressions and then we store everything in um in the EOS instance at 780. So you have here the physical appearance of the of the APN pharma which i will describe later


18
00:05:23.065 --> 00:05:36.065
And I want to talk just a second of the data transport. So this happens from the readout nodes into the GPU nodes via InfiniBand and RDMA.


19
00:05:36.065 --> 00:06:01.065
So basically, the data transferred directly from the memory of the without nodes into the memory of the APN farm and they're accessed directly from from the the gpus and this is quite done quite efficiently on the apn system with because the shared memory is allocated dynamically


20
00:06:01.065 --> 00:06:18.065
And with this strategy, we were able to digest this 700 gigabyte per second but we tested the system and it's able to digest up to 1.2 terabyte per second. So from the readout into the GPU file.


21
00:06:18.065 --> 00:06:33.065
And this, of course, is more than enough to to sustain the rates that we need. In 2023, during that lead, we were able to store up to four petabytes per day in the sanity instance.


22
00:06:33.065 --> 00:06:52.065
So one word on the on the hardware so we are using mi imd md um gpus so we use mi50 and mi 100 so here i i am giving some comparisons for the ma50 gpus which is our let's say reference node


23
00:06:52.065 --> 00:07:12.065
So one GPU can replace 80 cpu cores of room generation in synchronous where most of the load is TPC and 55 cores in asynchronous where most of the load is It's still TPC but


24
00:07:12.065 --> 00:07:33.065
Not as much as in a line because most of the tracking was done in a line, of course. You can see here a node. Each node has two lumen domains so the memory is, there is a privileged access from from one cpu to the


25
00:07:33.065 --> 00:08:01.065
To the corresponding GPUs, even though of course the two nodes could the two gpus could could swap data but we try to limit that And this is the breakdown in terms of pedophl of the farm. So we have 280 nodes with my 50 gpus and 70 nodes with the MI100 GPUs. We upgraded the farm a bit during run three because we found


26
00:08:01.065 --> 00:08:17.065
That we needed more compute. We use mostly vector operation, so not tensor, not at the moment. We will use tensor cores in RAM4 for TPC clustering but this is not yet there at the moment. So most of the performance that we


27
00:08:17.065 --> 00:08:38.065
A need is a vector fp32. Operation and here you can have some numbers about the aspects of the FAR, but I think this is less relevant. The important thing is that If we would only use CPUs, we would need more than 2064


28
00:08:38.065 --> 00:09:04.065
Core servers instead of 350 so you can see from this that the advantage in terms of power is substantial So this is a bit technical slide so it's the how the synchronous processing works so as i told you, we have 3.5 terabytes per second going into the epn farms and in the APN farms we do online calibrations and basically the whole


29
00:09:04.065 --> 00:09:17.065
Dpc processing. I won't go into the details. I can answer questions later in case. And then we can use the farm also for offline reconstruction.


30
00:09:17.065 --> 00:09:25.065
And here we do the reprocessing and the full calibration and the full reconstruction and multiple calibration passes. And of course.


31
00:09:25.065 --> 00:09:44.065
Here we have also the production of analysis objects that go on disk in the grid and are accessed by the analyzers and then all the compressed time frames are stored in the on tape.


32
00:09:44.065 --> 00:10:14.065
One word on the compression. So we use this range asymmetrical numeral system, which is an entropy based coding and it's much more efficient at much more efficient than any library for for compressing like GDP for instance and it's even more efficient than than offman


33
00:10:15.065 --> 00:10:32.065
Okay, now I come to the physical infrastructure. So we have this farm hosted in these containers. So you can see here we have four containers. So we actually use only three of them one is a spare. We have a dual feed of 2.5 megawatt.


34
00:10:32.065 --> 00:10:57.065
And we have a reverse osmosis water plant to purify the water because this air coolers that you see on the top, you see four units for container they use adiabatic cooling which is much more efficient than pure mechanical cooling and the servers are of course air cooled so when the air uh the outside air is not cold enough to


35
00:10:57.065 --> 00:11:20.065
To provide enough ETH exchange we spray water purity purified water on the heat exchangers so that by the evaporation of the water we get the additional cooling capacity. And this is quite efficient. So you can see here an example when the pumps start to


36
00:11:20.065 --> 00:11:36.065
To work so you have of course a delta when the delta is exceeded then the pump brings the temperature down and and then it stops and it goes up again and so on so you have these oscillations which are natural.


37
00:11:36.065 --> 00:12:06.065
And basically the water is recuperated as much as possible. But of course, we have some water consumption also due to the fact that we need to flush the containers for bacteria and Legionella risks and so on. So we cannot keep always the same water running. So this is a first prototype of the monitoring of the environmental parameters from the farm. I show here the two most loaded containers


38
00:12:09.065 --> 00:12:32.065
So we have the energy that is consumed, the water efficiency so which is given by the water volume the ratio between water volume for over it energy and the power usage effectiveness which is basically how much this is not an efficiency effectiveness. So it tells you


39
00:12:32.065 --> 00:12:41.065
For some cooling that you use, how well you are exploiting the cooling to to cool your IT equipment.


40
00:12:41.065 --> 00:12:54.065
So to give you an example, a non-adiopathic pure mechanical cooling like the one of Cernity, I think as a POE of 1.6, 1.7. And we are around 1.0 something.


41
00:12:54.065 --> 00:13:07.065
So it's rather efficient. And I'm also calculating here the equivalent CO2 emitted by this by the IT load.


42
00:13:07.065 --> 00:13:24.065
Of course we use french energy which is less is much greener than than in other countries so this is It's not our merit, but okay it's what it is. So their reverse osmosis plant does not use any chemicals just ordinary salt


43
00:13:24.065 --> 00:13:31.065
So it does not have any real impact on the environment.


44
00:13:31.065 --> 00:13:56.065
Okay, here you see in 2024 how the farms under the heaviest load of lead lead at 50 kilohertz You can see how the farm behaves in terms of power consumption so in the on the left you see the buffer utilization on the on the gpu farm


45
00:13:56.065 --> 00:14:26.065
And you see also the input rate of time frames in the farm and the output rate and you see the luminosity is basically reflected in the data rate. And if you look at the power consumption, you can see also that the luminosity is clearly shown in the power consumption. So when the machine is


46
00:14:26.065 --> 00:14:47.065
Providing a high rate and the data taking is ongoing you you have the data rate that is following nicely the luminosity you see some dips because here the data taking was stopped so computers are not working and then it restarts and then it restops there were some operations in the control room and you can see the the poe


47
00:14:47.065 --> 00:15:01.065
So at the moment, we are always running lead in the in the winter because it's in november uh so we basically haven't used adiabatic cooling for that we have been using it for prod, proton.


48
00:15:01.065 --> 00:15:17.065
Which is not as demanding. We will use adiabatic cooling next year because the area neuron will be in summer so we will exploit the this this feature as much as we can next year.


49
00:15:17.065 --> 00:15:45.065
Okay, so just to, yeah, this last slide. So Alice has been always using GPUs for computing Since 2010. And now we have this pharma which i showed you It's very efficient in terms of compute resources and of infrastructure resources. Here, I just want to show you this figure. So if we would have to


50
00:15:45.065 --> 00:16:09.065
Use the same number of CPU only nodes of ROM generation, we will have this this number here. So these are 30 let's say 32 core servers so it's okay it's twice as much as what i mentioned before And we will have


51
00:16:09.065 --> 00:16:27.065
Twice as much of energy consumption. So we will keep forum four to two we stick to this strategy of course switching to new generation GPUs and uh yes and we will have some also higher the rates to cope with.


52
00:16:27.065 --> 00:16:39.065
And here I just mentioned some papers that we published. So if you want to know more about this setup, you can you can have a look at these papers. Thank you.


53
00:16:39.065 --> 00:16:48.065
Thank you very much, Federico. Or this very global take on your GPU firm.


54
00:16:48.065 --> 00:17:12.065
Is there a question?


55
00:17:12.065 --> 00:17:13.065
Yes.


56
00:17:13.065 --> 00:17:17.065
So I have a quick one. Actually, you mentioned how much more you would need with CPUs. In terms of investment because money is always eating us. Is it also worthwhile.


57
00:17:17.065 --> 00:17:40.065
Oh, yes, of course. If you would buy because the server cost is uh is not negligible i mean the the the physical server cost is not negligible so that's why we all choose a service that cannot as the maximum number possible


58
00:17:40.065 --> 00:17:52.065
Gpus. So I don't have a figure, but I think the farm would have cost up to two or three times more.


59
00:17:52.065 --> 00:18:20.065
Because it's… You need to buy much more motherboards and power supplies and and everything Because you cannot stick all that stream processors into into even a four-way CPU motherboard so they and also you have to think that the power consumption of the GPUs is less because they run at a slower speed.


60
00:18:20.065 --> 00:18:41.065
They have much more cores, but they run at slower speeds I think the impact on the cost compute and power consumption and resource consumption is is it's really non-negligible it's factors so


61
00:18:41.065 --> 00:18:45.065
Okay. Thank you.


62
00:18:45.065 --> 00:18:49.065
Yes, I see a hand. Please, Domenico.


63
00:18:49.065 --> 00:18:53.065
Hi. Yes, thank you. Can you hear me? Can you hear me?


64
00:18:53.065 --> 00:18:58.065
You are very faint. Yes.


65
00:18:58.065 --> 00:19:10.065
Okay, yes, I wanted to comment on that aspect. Looking at this time line that you show from 2010.


66
00:19:10.065 --> 00:19:32.065
So may I get your confirmation that essentially the investment in bringing the software on GPUs made the difference because as you mentioned, so the GPU, fewer GPUs can cope with respect to a large amount of CPUs, but this is thanks to the fact that


67
00:19:32.065 --> 00:19:33.065
Yes.


68
00:19:33.065 --> 00:19:41.065
The part of your processing can run on gpu So the software, the effort of migrating software from CPUs to GPUs has been rewarding in that case


69
00:19:41.065 --> 00:20:04.065
Yes. Thanks for the question. So here in this slide uh so the effort was For us, maybe not so huge as it could appear because we have been always doing this so since 2010 we in the so-called hlt high level trigger


70
00:20:04.065 --> 00:20:29.065
We were already able to do reconstruction and event filtering online. At the time, it was not used to filter the data because people, you know, it was people it was very bold to throw data away online without no possibility to get them back. So it was the only use for compression but this experience in run one and run two


71
00:20:29.065 --> 00:20:52.065
With the high level trigger, basically we add the framework so when we implemented this for run three we were already advanced so of course there was a lot of work but yes the fact that we can actually we could benefit to have GPU grid sites because then we could


72
00:20:52.065 --> 00:21:14.065
We could run also the asynchronous reconstruction much more efficiently because our grid sites i mean the one far from certain are only CPU. So they are much slower Then the APN alone is like 17% of the grid asynchronous the grid offline power


73
00:21:14.065 --> 00:21:26.065
So it's major and plus uh it would be also an additional effort could be into porting the Monte Carlo.


74
00:21:26.065 --> 00:21:56.065
To GPUs, which we don't have at the moment, but at least the digitization because that it's you know experiment software could give another 20% of compute margin being ported on GPU. Porting giant is maybe out of our strength but uh yes so yeah there is a work on the software which we have been doing since 2010 so we are


75
00:21:56.065 --> 00:22:02.065
Let's say we were in a good shape But it's very rewarding.


76
00:22:02.065 --> 00:22:13.065
Yeah, thanks, Federico. Sorry we're a bit late so if it's okay with you, you can maybe ask your question on Matamost.


77
00:22:13.065 --> 00:22:14.065
And we will move to the last talk of this session.


78
00:22:14.065 --> 00:22:18.065
Sure, no problem.


79
00:22:18.065 --> 00:22:21.065
Eloya. Can you share your slides?


80
00:22:21.065 --> 00:22:24.065
Yes, can you read?


81
00:22:24.065 --> 00:22:26.065
Yes, very well.


82
00:22:26.065 --> 00:22:29.065
Okay.


83
00:22:29.065 --> 00:22:36.065
Federico needs to stop sharing his screen, otherwise I cannot. Okay.


84
00:22:36.065 --> 00:22:43.025
Thank you.



