WEBVTT

1
00:00:01.685 --> 00:00:03.525
Federico Ronchetti: Yes. So


2
00:00:12.475 --> 00:00:15.134
Federico Ronchetti: okay, they should be full screen. Now.


3
00:00:15.605 --> 00:00:16.754
Yann Coadou: Yes, perfect.


4
00:00:16.755 --> 00:00:31.315
Federico Ronchetti: Okay, thank you for the opportunity to speak. So I will discuss the Epn. Farm at Cern. So is the online processing farm for the Alice experiment.


5
00:00:31.645 --> 00:00:51.194
Federico Ronchetti: I'll just jump in because the slides are a bit dense. So preliminary context. So you all know that in round 3 even the lead, lead. Luminosity performance of the Lhc accelerator has increased


6
00:00:51.535 --> 00:00:52.915
Federico Ronchetti: so they can


7
00:00:53.205 --> 00:01:17.404
Federico Ronchetti: ship an interaction rate of 50 gigahertz for sustained times like one or 2 h and to cope with this increased rate the Alice detector underwent a major upgrade during the Ls. 2. The major changes were the


8
00:01:17.405 --> 00:01:25.145
Federico Ronchetti: PC. Which now is a Gem, Tpc. And the Pixel Detector, which is a 13 gigapixel.


9
00:01:25.155 --> 00:01:50.405
Federico Ronchetti: very, very granular map silicon detector, and also we abandoned the trigger readout in favor of continuous readout that implies implied a new computing model both in software and hardware and extensive use of Gpus and Gpus are used in Alice for producing physics directly online


10
00:01:50.405 --> 00:01:55.324
Federico Ronchetti: and also to calibrate the vector of course. Otherwise there would be no, no real physics.


11
00:01:55.445 --> 00:02:22.424
Federico Ronchetti: So that's that's I think in in certain is a unique, unique feature. We have 350 nodes with the total of 2,800 gpus. And this is the operational core of the data taking the the experiment cannot function without these Gpu's processing online. And we are we in Run 3, we expect to get


12
00:02:22.835 --> 00:02:40.434
Federico Ronchetti: a luminosity of roughly 6 6 inverse Nanobarn and the the overall budget will be 13 plus run 4.


13
00:02:41.255 --> 00:03:10.014
Federico Ronchetti: Okay for the event processing. Basically, we have now continuously out. So we take data in snapshots, time frames. This is not like free running, so free running. We have a strobe that drives the electronics, but the strobe takes this snapshot we can set the length of the snapshot. Now it's 32 Lhc. Orbits, which is almost 3 ms


14
00:03:10.015 --> 00:03:18.665
Federico Ronchetti: and the Apn. Farms. The Gpu farm performs tracking, calibration, reconstruction, and compression synchronous synchronously. So


15
00:03:18.715 --> 00:03:42.255
Federico Ronchetti: while data is being taken, most of the Gpu compute is, of course, for the Tpc. Tracking, and we have some CPU only nodes for for calibration, and we also use efficient methods to to compress this data because there is no way we cannot. We can save this data. in the indoor form anywhere.


16
00:03:43.525 --> 00:03:58.615
Federico Ronchetti: The basically one characteristic is also that this software is the same for synchronous and asynchronous. So for online and and offline. So basically, the same software is run


17
00:03:59.125 --> 00:04:26.344
Federico Ronchetti: in the 2 different scenarios. So this is a rather dense slide. I will just try to summarize it a bit. So we have a huge data stream coming from the detectors that is read out via Fpga cards. So we have basically 3.5 TB per seconds before 0 suppression that goes into the readout nodes. And here is 0 suppressed. Basically, this is


18
00:04:26.345 --> 00:04:40.785
Federico Ronchetti: mostly mostly Dpc, and and out of here we go into the Gpu farm with a stream of 700 GB per second. I'm talking about lead, of course, which is the most relevant case


19
00:04:41.235 --> 00:04:52.514
Federico Ronchetti: in the Epn farm the events are built, built. They are not really events. These are, as I said, just time windows that are superposed and matched.


20
00:04:52.685 --> 00:04:59.435
Federico Ronchetti: and all this is done on the Gpu


21
00:04:59.605 --> 00:05:10.444
Federico Ronchetti: on the Gpu cards. So there's basically no no cpus involved except for calibration. And then we have online compressions. And then we store everything in


22
00:05:10.745 --> 00:05:39.845
Federico Ronchetti: in the Eos instance at 700. So you have here, the physical appearance of the of the ATM farmer, which I will describe later. And I want to talk just a second of the data transport. So this happens from the without nodes into the Gpu nodes via infiniband and Rdma. So basically, the data transferred directly


23
00:05:39.845 --> 00:06:01.064
Federico Ronchetti: from the memory of the without notes into the memory of the Apn farm. And they're accessed directly from from the Gpus, and this is done quite efficiently on the Apn system with, because the shared memory, it's allocated dynamically.


24
00:06:01.085 --> 00:06:13.194
Federico Ronchetti: And in this, with this strategy we we were able to digest this 700 GB per second. But we tested the system, and it's able to digest up to 1.2


25
00:06:13.265 --> 00:06:17.365
Federico Ronchetti: terabyte per second. So from from the result into the Gpu file.


26
00:06:17.555 --> 00:06:20.425
Federico Ronchetti: and this, of course, is more than enough to


27
00:06:20.965 --> 00:06:45.814
Federico Ronchetti: to sustain the rates that we need in 2023. During that lead we were able to store up to 4 petabyte per day in the 70 instance. So one word on the on the hardware. So we are using. Mi imd. Md. Gpus. So we use mi. 50, and mi. 100. So here I am giving some


28
00:06:45.815 --> 00:06:52.254
Federico Ronchetti: comparisons for the mi 50 gpus, which is our, let's say, reference, node.


29
00:06:52.255 --> 00:07:11.605
Federico Ronchetti: So 1 1 Gpu can can replace 80 CPU cores of ROM. Generation in synchronous where most of the load is Tpc. And 55 cores in Asynchronous, where most of the load is still Tpc, but


30
00:07:12.252 --> 00:07:40.424
Federico Ronchetti: not not as much as in online, because most of the the tracking was done in online. Of course you can see here a node. Each node has 2 Newman domains, so the memory is a privileged access from from one CPU to the to the corresponding gpus, even though, of course, the 2 nodes could, the 2 Gpus could could swap data.


31
00:07:40.425 --> 00:08:02.174
Federico Ronchetti: But we try to limit that. And this is the the breakdown in terms of petaflop of the farm. So we have 280 nodes with my 50 gpus and 70 nodes with my 100 gpus. We. We upgraded the farm a bit during entry because we found that we needed more compute.


32
00:08:02.524 --> 00:08:21.064
Federico Ronchetti: We use mostly vector, operation. So not tensor, not not at the moment. We will. We will use tensor cores in RAM 4 for Tpc clustering. But this is not not yet there at the moment. So most of the performance that we need is is a vector Fp 32


33
00:08:21.205 --> 00:08:31.754
Federico Ronchetti: operation. And here you can have some some numbers about the the aspects of the far, but I think this is less relevant. The important thing is that


34
00:08:32.153 --> 00:08:56.065
Federico Ronchetti: if we would only use cpus, we would need more than 2064 core servers instead of 350. So you can see from this that the advantage in terms of power is is substantial. So this is a bit technical slide. So it's the how the synchronous processing works. So, as I told you, we have


35
00:08:56.065 --> 00:09:06.064
Federico Ronchetti: 3.5 TB per second going into the Apn farms. And in the Apm farms, we do online calibrations. And basically the whole Tpc processing.


36
00:09:06.065 --> 00:09:29.974
Federico Ronchetti: I won't go into the details. I can answer questions later in case, and then we can use the farm also for offline reconstruction. And here we do the reprocessing and the full calibration and the full reconstruction and multiple calibration passes. And of course, here we have also the production of


37
00:09:29.995 --> 00:09:42.175
Federico Ronchetti: analysis objects that go on disk in the grid and are accessed by the analyzers, and then all the compressed time frames are stored in the


38
00:09:42.255 --> 00:09:43.195
Federico Ronchetti: on tape


39
00:09:44.833 --> 00:10:13.745
Federico Ronchetti: one word on on the compression. So we use this range asymmetric numeral system, which is an entropy based coding. And it's much, much more efficient that much more efficient than any any library for for compressing like, for instance. And it's even more efficient than than Afman.


40
00:10:15.332 --> 00:10:36.285
Federico Ronchetti: Okay. Now I come to the physical infrastructure. So we have this farm hosted in these containers. So you can see here we have 4 containers. So we actually use only 3 of them. One is a spare. We have a dual feed of 2.5 megawatt, and we have a reverse osmosis water plant to


41
00:10:36.285 --> 00:11:01.044
Federico Ronchetti: purify the water, because these air coolers that you see on the top you see 4 units for container. They use adiabatic cooling, which is much more efficient than pure mechanical cooling, and the servers are, of course, air cooled, so when the air, the outside air, is not cold enough to to provide enough heat exchange, we


42
00:11:01.045 --> 00:11:19.465
Federico Ronchetti: spray water, purified water on the heat exchanger, so that by the evaporation of the water we get the additional cooling capacity. And this is quite efficient. So you can see here an example. When the pumps start to


43
00:11:19.465 --> 00:11:36.104
Federico Ronchetti: to work. So you have, of course, a delta. When the delta is exceeded, then the pump brings the temperature down, and and then it stops, and it goes up again, and so on. So you have these oscillations, which are natural


44
00:11:36.105 --> 00:11:50.585
Federico Ronchetti: and and basically the water is recuperated as much as possible. But of course we have some water consumption also, due to the fact that we need to flush the container.


45
00:11:50.685 --> 00:12:09.434
Federico Ronchetti: the containers for bacteria and legionella risks, and so on, so we cannot keep always the same water running. So this is a 1st prototype of the monitoring of the environmental parameters from the farm we have I show here the 2 most loaded containers.


46
00:12:09.435 --> 00:12:31.935
Federico Ronchetti: So we have the energy that is consumed, the water efficiency. So which is given by the water volume, the ratio between water volume for over it energy and the power usage effectiveness, which is basically how much this is not an efficiency effectiveness. So it's it tells you.


47
00:12:31.935 --> 00:12:39.225
Federico Ronchetti: for some cooling that you use, how well you are exploiting the cooling to


48
00:12:39.225 --> 00:13:06.404
Federico Ronchetti: cool your it. Equipment. So, to give you an example, a non Adiabatic, pure, mechanical cooling, like the one of 70, I think, has a pue of 1.6 1.7, and we are around 1.0 something. So it's it's rather efficient. And I also, I'm also calculating here the equivalent Co. 2 emitted by this by the it load.


49
00:13:06.625 --> 00:13:31.174
Federico Ronchetti: Of course, we use French energy, which is less greener than than in other countries. So this is it's not our merit, but okay, it's what it is. So their reverse osmosis plant that does not use any chemicals, just ordinary salt. So it's not. It does not have any real impact on the environment.


50
00:13:31.649 --> 00:13:46.054
Federico Ronchetti: Okay, here, you see in 2024 how the farms under the heaviest load of lead lead, that 50 kilohertz. You can see how the farm behaves in terms of


51
00:13:46.365 --> 00:13:55.925
Federico Ronchetti: power consumption. So in the on the left, you see the buffer utilization on the on the Gpu farm.


52
00:13:55.925 --> 00:14:20.885
Federico Ronchetti: and and you see the also the input rate of time frames in the in, the, in the farm and the output rate. And you see, the luminosity is basically reflected in the data rate. And if you look at the power consumption, you can see also that the luminosity


53
00:14:20.885 --> 00:14:46.965
Federico Ronchetti: is clearly shown in the power consumption. So when the machine is providing a high rate. And the data taking is ongoing you, you have the data rate that is following nicely the luminosity. You see some dips, because here the data taking was stopped. So computers are not working, and then it restarts, and then it stops. There were some operations in the control room, and you can see the the Poe.


54
00:14:47.185 --> 00:15:03.664
Federico Ronchetti: So at the moment, we have been always running lead lead in the in the in the winter, because it's in November. So we basically haven't used Adiabatic cooling for that. We have been using it for product which is not as demanding.


55
00:15:03.845 --> 00:15:16.484
Federico Ronchetti: We will use Adiabatic cooling next year, because the neuron will be in summer. So we will exploit this feature as much as we can next year.


56
00:15:17.645 --> 00:15:32.494
Federico Ronchetti: Okay, so just to, yeah, I'm this last slide. So Alice has been always using Gpus for computing since 2,010. And now we have this farmer which I showed you.


57
00:15:32.495 --> 00:15:56.764
Federico Ronchetti: It's very efficient in terms of compute resources and of infrastructure resources. Here, I just want to show you this this figure. So if we would have to use the same number of CPU only nodes of ROM generation. We will have this this number here. So it's these are 30.


58
00:15:56.765 --> 00:16:05.774
Federico Ronchetti: Let's say 32 core servers. So it's okay. It's twice as much as what I mentioned before.


59
00:16:07.006 --> 00:16:13.825
Federico Ronchetti: And we will have twice as much of energy consumption. So we will keep forum, for to


60
00:16:15.355 --> 00:16:34.104
Federico Ronchetti: we stick for to this strategy, of course, switching to new generation Gpus. And yes, and we will have some also higher data rates to to cope with. And here I just mentioned some papers that we publish. So if you want to know more about this this setup, you can.


61
00:16:34.415 --> 00:16:37.555
Federico Ronchetti: You can have a look at these these papers. Thank you.


62
00:16:39.455 --> 00:16:43.795
Yann Coadou: Thank you very much, Federico. But it's very


63
00:16:43.905 --> 00:16:46.995
Yann Coadou: global. Take on the on your Gpu farm.


64
00:16:47.485 --> 00:16:49.475
Yann Coadou: Is there a question


65
00:16:57.735 --> 00:16:58.445
Yann Coadou: so


66
00:16:59.435 --> 00:17:14.814
Yann Coadou: I have a quick one. Actually, you mentioned how much more you would need with cpus in terms of investment, because money is always eating us? Is it also


67
00:17:15.455 --> 00:17:16.655
Yann Coadou: worthwhile.


68
00:17:17.045 --> 00:17:19.784
Federico Ronchetti: Oh, yes, of course this


69
00:17:19.965 --> 00:17:29.604
Federico Ronchetti: the if you would, if you would buy, because the server cost is is not negligible. I mean the the


70
00:17:29.895 --> 00:17:40.874
Federico Ronchetti: the physical server cost is not negligible. So that's why we all choose a service that can host as the maximum number possible of gpus.


71
00:17:41.425 --> 00:17:42.315
Federico Ronchetti: so


72
00:17:42.715 --> 00:17:51.845
Federico Ronchetti: it I don't have a figure, but I think the firm would have would have cost up to 2 or 3 times more.


73
00:17:52.075 --> 00:17:54.055
Federico Ronchetti: because it's a


74
00:17:54.728 --> 00:18:06.234
Federico Ronchetti: you have to have you. You need to buy much more motherboards and power supplies and and everything because you you cannot stick all that stream processors into into


75
00:18:06.405 --> 00:18:10.274
Federico Ronchetti: even a four-way CPU motherboard.


76
00:18:10.325 --> 00:18:39.494
Federico Ronchetti: So they. And also you have to think that the power consumption of the gpus is less because they run at a slower speed. They have much more cores, but they run at slower speed. So there is a, I think the the impact on the cost compute and power, consumption and resource. Consumption is, it's it's really non-negligible. It's factors. So.


77
00:18:40.595 --> 00:18:48.975
Yann Coadou: Okay, and 2, and I oh, yes, I see a hand, please, Domenico.


78
00:18:49.305 --> 00:18:53.395
Domenico Giordano: Hi, yes, thank you. Comment on this call.


79
00:18:53.395 --> 00:18:55.165
Yann Coadou: You are very faint.


80
00:18:55.165 --> 00:19:14.251
Domenico Giordano: Can you hear me? Yes. Can you hear me? Okay, yes. I wanted to comment on that aspect, looking at this timeline that you show from 2,010. So may I get your confirmation that essentially the investment in


81
00:19:14.785 --> 00:19:27.505
Domenico Giordano: bringing the software on Gpus made the difference, because, as you mentioned, so the Gpu fewer gpus can cope with respect to a large amount of cpus. But this is thanks to the fact that


82
00:19:27.625 --> 00:19:31.905
Domenico Giordano: the part of your processing can run on. Gpu.


83
00:19:31.905 --> 00:19:32.255
Federico Ronchetti: Yes.


84
00:19:32.255 --> 00:19:40.854
Domenico Giordano: The the software. The effort of migrating software from cpus to Gpu's has been rewarding. In that case.


85
00:19:41.045 --> 00:20:03.604
Federico Ronchetti: Yes, yes, thanks for the question. So here in this slide, so the effort was for us, maybe not so huge as it could appear because we have been always doing this. So since 2010, we in the so-called Hlt high level trigger.


86
00:20:03.605 --> 00:20:16.514
Federico Ronchetti: we were already able to do reconstruction and event filtering online at the time it was not used to filter the data because people you know, it was


87
00:20:16.555 --> 00:20:36.224
Federico Ronchetti: people. It was very bold to throw data away online without, you know, no possibility to get them back. So it was only used for compression. But this experience in Run one and run 2 with the high level trigger it. Basically, we had the framework. So when we implemented this for Run 3,


88
00:20:37.215 --> 00:20:51.484
Federico Ronchetti: we were already advanced. So of course, there was a lot of work. But yes, the fact that we can actually we could benefit to have Gpu grid sites, because then we could


89
00:20:52.210 --> 00:21:15.494
Federico Ronchetti: we could run also the asynchronous reconstruction much more efficiently because our grid sites, I mean the one far from Cern are only CPU, so they are much slower than than the Apn alone is like 17% of the grid asynchronous, the grid offline power of of Alice.


90
00:21:15.845 --> 00:21:25.935
Federico Ronchetti: So it's it's major and plus. It would be also an additional effort could be into porting the Monte Carlo


91
00:21:26.025 --> 00:21:55.415
Federico Ronchetti: to Gpu's, which we don't have at the moment, but at least the digitization, because that it's, you know, experiment software could give another 20% of compute margin being ported on Gpu. Porting giant is maybe out of our strength. But yes. So yeah, there is a work on the software which we have been doing since 2010. So we are.


92
00:21:55.725 --> 00:22:00.015
Federico Ronchetti: Let's say we were in a good shape. But it's very rewarding.


93
00:22:02.625 --> 00:22:08.745
Yann Coadou: Okay, thanks, Federico. So yes, sorry. We are a bit late. So


94
00:22:08.895 --> 00:22:16.744
Yann Coadou: if it's okay with you, you can maybe ask your question on mattermost, and we will move to the last talk of this session.


95
00:22:18.430 --> 00:22:21.424
Yann Coadou: Elia, can you share your slides?


96
00:22:21.425 --> 00:22:23.204
Ilaria Vai: Yes. Can you hear me?


97
00:22:24.105 --> 00:22:25.254
Yann Coadou: Yes, very well.


98
00:22:25.925 --> 00:22:26.469
Ilaria Vai: Okay.


99
00:22:31.563 --> 00:22:36.985
Ilaria Vai: Federico needs to stop sharing his screen. Otherwise I cannot. Okay, thank you.



