Arista Networks, Inc. (ANET) Earnings Call Transcript & Summary

October 30, 2024

New York Stock Exchange US Information Technology special 65 min

Earnings Call Speaker Segments

Neha Sakhalkar

executive
#1

Hello, everyone. We are excited to have you join us today. I am Neha Sakhalkar, part of the Arista Technical Services team doing a quick walk-through of the TAC webinar series. Upasana, can you move over to the next slide, please? Thank you. We kickstarted the TAC webinar series back in 2022 to provide a live forum for our customers to come and interact with TAC, while we share a few techniques for troubleshooting some common issues. Since its inception, we have covered various technical services -- technical scenarios like switch health checks, latency, packet drop issues seen across multiple environments like VXLAN, EVPN, multicast, et cetera. For today's session, we have 2 of our senior tech leads, Upasana and Vignesh who will share their insights and expertise in troubleshooting congestion and packet drop-related issues seen in AI networks. Plus we also have a few other tech leads and seasoned engineers as panelists to address any questions raised during the session. Please feel free to use the Q&A section to interact with us. Now a download link for the topology diagram was e-mailed to our registered participants last week. If you don't have it yet, don't worry. We would just recommend downloading it using the link shared on a chat now. Upasana, can we move over to the next slide? All of our TAC webinars are recorded, and those recordings with slide deck are uploaded to a community page. If you are interested in trying out these scenarios in your lab, we do also upload sample configuration file for each webinar on the same community page. Feel free to utilize that. Now in addition to our publicly available TAC webinar, we also offer customized training. This program was launched last year based on the growing interest in personalized training. It is a fully specialized paid offering led by our TAC experts. While the scenarios presented in TAC webinar sessions are broad in scope, our custom training is designed to align with your specific platform, feature, and network design requirements. It includes an in-depth analysis of recent TAC cases associated with your account, where we examine common network issues observed in your environment and provide essential recommendations to mitigate these issues moving forward. These trainings have proven to be highly beneficial for the network ops team, empowering them to effectively manage any network-related challenges that may occur in the future. So far, we have received excellent feedback on both the quality as well as the content of these sessions. If you are interested, feel free to reach out to us at tac-custom-training@arista.com. Upasana, can we move to the next slide, please? Okay. So before we move on with the troubleshooting session, let me take a few minutes to go over the networking requirements for AI. If you have any experience with AI training models, you are likely aware of how essential it is to achieve a low job completion time. Besides having high-performing GPUs, your network should be optimized for high bandwidth, low latency and advanced hashing capabilities to prevent bottlenecks that can slow down your job completion time. This can negatively impact your application performance and user experience. Therefore, in today's webinar, we will explore such common challenges related to job completion time and go over how you can benefit from Arista's advanced feature set to identify and resolve this issue. Now without further ado, let me pass on the baton to Upasana and get rolling. Go ahead, Upasana.

Upasana Dangi

executive
#2

Thanks, Neha. Hello, everyone. Thank you for joining us today. Just a quick introduction. My name is Upasana, and I'm a Senior Technical Lead with Arista Tech. And as Neha mentioned, today, Vignesh and I will continue on into round 2 of how to troubleshoot various issues in AI networks and how we can debug these problems through CLI or CloudVision. But before we get into our first problem statement for the day, let's go over some general information about the setup. So we have some 7060X5s that we're using in our leaf and spine layer. They are all streaming telemetry information to our CloudVision as-a-service instance or CVaaS. Each leaf connects to a set of GPUs, and the communication between the racks is through a 4-way CMP that comes from the spines. We do have dynamic load balancing configured to make sure that there's even traffic distribution. Our AI-related flows are using RoCE V2 and they are mapped to traffic class 3. Now within traffic class 3, I am making use of ECN and PFC to help with congestion management for my AI workloads. In the event there is congestion, for better visibility, I've also enabled LANZ as a feature on the switches, and this data is also being streamed to CloudVision. So with that, let's jump into our problem statement. Our customer is observing poor performance for their AI workloads, which is causing some jobs to fail, workloads that are running slow, and they're just trying to figure out what's causing this. So before we begin, we have to keep in mind that while our setup may be small, AI networks can be fairly scaled up. This can make troubleshooting extremely challenging, which is why we, especially in TAC, tend to rely on collecting as much information as possible before we begin the debugging process. So first up, I always like to get some more information about the timeline. And in our case, it looks like our alarm bell started around October 27 around 12:00 p.m. PST. So this gives us our initial time frame. Now as a follow-up, I also want to preemptively know if there was any network activity around this time because that could also be a potential trigger. But in our case, the network was untouched. That being said, right now, we don't have much information from the GPU side. Now next, I want to understand if this has been happening since day 1 because that could also be a fundamental issue in our config or in our design as opposed to a targeted section of the network that's contributing to the problem. Now in our case, while our users are unsure about the level of the impact, they do mention that this was not a day 1 issue. Now finally, and while this step is not always easy, when possible, it does really give us a leg up if we can narrow down a set of GPU racks that are common points. Now in our case, we did determine that Rack 1 and 2 GPUs were involved for all of the workloads. Now this way, I do get a smaller and a more tangible set of devices to focus on because given that Rack 1 and 2 GPUs might be involved, for our initial look, our traffic can traverse AI leaf switches 1 and 2 along with their spines. So with this information in hand, let's do some investigating. Now as a start to avoid checking all of these devices manually for signs of an issue, I'm first going to jump into our CVaaS instance to get a bird's eye view of the network health on the devices we are focused on. This can help us rule out the usual suspects that cause the symptoms that the customer is observing. And this is also a useful step in general to help sanitize your network health if there wasn't much information to go on, where you could actually narrow down to a section of a network. So with that, let's log into our CloudVision instance. I have already recorded the GUI and the CLI part of the debugging to save us some time, so I'm just going to walk you through that right now. All right. So we're in our CVaaS instance now. And in order to get a more compact view of the network, I'm actually going to start with a custom dashboard that I've built, which helps me keep track of some of the common issues that can affect my AI workloads. This way, I don't have to scan the full events page for all of my devices, and I can also get a quick and easy summary of what my network looked like on October 27. Now this is a very customizable events-based dashboard where a user can pick and choose different kinds of CloudVision events that are critical to their network and then collate them in 1 space for all their network devices. This way, it sort of becomes a one-stop shop for overall network health. Now as we focus on the tiles we have here, if you look at the bars that you can see, these indicate the time range and they're also color coded. In the event you see that they're green, it indicates that for the event that was specified in the view, there was no activity during the time frame and everything has been stable for that view. Now other colors can indicate that an event was triggered and the color itself indicates the severity of the underlying event. But let's come back to my dashboard. The default query I can use here can just be a straightforward tag of devices star, which looks up all of my user-defined tiles for switches that are streaming in this network. But for our scenario, since we did decide to focus on force devices, I'm going to specifically filter for them so our data becomes easier to read. That being said, if you are debugging a scenario where the data points were just not enough for you to narrow down to a section of a network, you could always still just go back and use the default query that runs on all the devices to cast that wider net for potential triggers. All right. So I have my filters on, as you can see and I'm also tracking this for October 27. So let's dig deeper into these tiles here. So looking at our first check for the day, we want to see if there was any conflict change on our devices, the ones that you see here in question. And it looks like there wasn't anything done on 27th that could have triggered our issue, and this is also in line with what the customer has mentioned. So let's move on to see other statistics that could have contributed to poor performance of the AI workloads. One such thing could be BGP neighborship flapping. We could also be incurring some layer 1 errors like XCS or symbol errors that can cause packet drops or just your basic interface flaps that can also result in these issues. However, in our case, none of these are a problem at the moment because they've been pretty stable during the day because from here, it looks like everything looks green. So let's keep scrolling down, and it looks like we do see some queue congestion events on AI Leaf 1 that seem to be coinciding with the timeline from when we first started to see the problem. Because if you see here, these events also seem to have started on October 27 around noon PST. Now let's click on one of the events that got flagged. We see that these are a few threshold events. These events are actually built from the LANZ streaming data that is coming directly from the switch. And as discussed in previous webinars, the LANZ feature, when enabled, helps us identify periods of congestion when we have bursty traffic. It is most useful for the times when the burst is for a short duration, which usually tends to get missed when we're pulling on just interface counters. But again, back to our event over here. This particular event in CloudVision indicates that during the time it was fired, our device, which in this case is AI Leaf 1, was experiencing some congestion for traffic that was intended to egress one of its interfaces, Ethernet 11/5. And during this time, the buffering exceeded the queue threshold that was defined within CVP. So let's spot check some of the other instances of this event. And it looks like a lot of these events right now at least are for AI Leaf 1, Ethernet 11/5, which from my diagram connects to GPU 12. All right, let's just go 1 more level down here. And from this graph, it kind of looks like the problem is just not isolated to excessive buffering. Coinciding with this above event, we also see evidence that our switch has been actively discarding traffic at the time of the problem. And if you open up one of these discard events, it looks like, again, it's the same interface, Ethernet 11/5 on AI Leaf 1. So let's open up one of these discard events and poke around. Now I've opened up 1 for us. And first step, let's just focus on this top left corner here, where we can see that during this time frame, our link, which was Ethernet 11/5, it was running pretty hot, which sort of explains the congestion and the Tx drop rates that you see below. And speaking of congestion, let's also take a look at the LANZ graph on this right-hand side here. This graph gives us an idea of how many buffer segments that you can see in the queue are actually being used during congestion on Ethernet 11/5. And for context, each of the segments here translates to about 254 bytes, and we'll actually be talking more on this later today as well so hold on to this thought. Looking more in the same graph, we can also see there are reports of LANZ queue drops, which again is another way for us to see if traffic was being discarded during a congestion event or was there just plain buffering. Now in our case, the discards that we see here are possible during a congestion event if our traffic hits the maximum hog limit while it was buffering. And this could explain the poor performance that the customer has been seeing. However, it is weird because I would have not even expected us to reach the point of dropping traffic because we do have ECN and PFC that help manage the congestion and slow down the traffic from the server side. So it looks like now I have a solid starting point but also some questions. And so as next steps, I'm going to move my investigation to AI Leaf 1 to see what's going on. So let's go in there. But before we actually jump into our CLI, let's also do a quick recap of what we know and what have we found out so far. Now we know that our issue started somewhere on October 27 around noon. We narrowed down our scope of debugging to AI Leaf 1 and 2 along with the spine layer from the GPU racks that were isolated. We went on to CloudVision and also cleared some basic health checks across the devices. And from these checks, we found evidence of congestion and drops on one of them, in our case, AI Leaf 1, that also aligned with our issue timeline and could also explain perhaps the problems that are faced by the customer. So as next steps, I want to jump on the CLI for AI Leaf 1 to understand the cause of congestion and also specifically why my ECN and PFC configurations are not helping here. So with that, let's actually move on from here and jump on the CLI for AI Leaf 1. All right, so this is our session for AI Leaf 1. And first step, let's focus on, are we seeing any congestion? We can do this by running show queue-monitor length, which gives us a summary of the previous LANZ events, and it's kind of similar to the graph that we saw in CloudVision earlier. But right now, we don't actually have active congestion because our last event sort of ended about 41 minutes ago. The time, by the way, over here is in UTC. But from this historic data, we can identify that all of our congestion events so far have been isolated to traffic class 1 on Ethernet 11/5. The traffic class is what's indicated in the brackets. So before we look any further, since we also saw discards in CloudVision, I'm going to do a quick check on the interface counters to see where these discards are. So on Ethernet 11/5, let's look at queue detail. And here, I just want to see if the congestion-related drops that we saw are all on TC1 only or are they spread across. And sure enough, all of our drops are on TC1. Again, this is still odd for a couple of reasons. One, these drops have started around our issue timeline so the timeline aligns for us. But that being said, as I mentioned earlier, my AI-related workflows are on TC3, which -- for which we have traffic but we don't really have drops, right? And two, while our drops in TC1 do not correlate right now to our AI flows, it's still an interesting data point because the amount of traffic and subsequent drops is not really expected in this area of the customer's network. So with this, before we continue looking for other data points, since the timeline is right in our issue window and the level of traffic in TC1 is unexpected, I want to check if the drops we see are for traffic that is truly unrelated to my AI workloads or if there's something else going on here that needs to be understood. And ideally, I'd want to sample the traffic during an active congestion window, and I also want to sample traffic that's being queued instead of everything that's going out of Ethernet 11/5. This way, I don't have to shift through all the data to figure out what traffic is facing congestion. So in order to do this, I'm going to make use of the LANZ feature in EOS. This feature allows us to run a packet capture on a copy of the traffic that is being queued. And since TC1 is showing congestion, I can be relatively certain that at least part of the traffic that I capture using this feature is going to give me the flows that I'm looking for. So I have reconfigured LANZ here, LANZ mirroring, show on section mirror. And as you can see, it has been configured, it's enabled and I'm sending a copy of my traffic to CPU instead of a destination collector because this way I can handle everything on the switch directly. The copy that is going to CPU will be rate-limited by COP so it should be relatively safe to use. Now in order to capture these packets, we are going to need to run the capture itself on a special virtual port that gets created for this feature. And you can see this port if you just run show monitor session as we're doing now. And here, we can see for LANZ mirroring, we have created a new internal port called LANZ. So let's start our capture here and just do a quick tcpdump, verbose it, run it on this new port, LANZ. Okay, so this looks to be empty right now. And it looks like we're running into this classic challenge where currently, my device is not going through congestion because there's a good chance there's no active job running. So now I have 2 options. I can either log into the switch at the exact time of the next congestion and run this capture, or I can find a way to run it in the background and collect a sample during my next discard event. So we can actually do this by using one of my favorite tools, which is the tcpdump session feature. This feature is going to allow me to create a capture session that runs in the background, which can be log-rotated to ensure we don't consume too much space. And it's kind of perfect for catching intermittent events like this one. Again, saving us some time, I have it configured on the box so assurance section tcpdump. As you can see, I have a tcpdump session defined by the name of Webinar, and I'm capturing on queue monitor instead of a physical port because we're doing this for the LANZ mirroring feature. I'm saving my captures into the flash drive as Webinar Captures. And to keep this space from growing, I'm just going to create about 5 files, you can create as many as you want. And I'm capping it at about 50 MB, which is going to keep getting log-rotated. All right. So now we can just wait for the next event to actually analyze this data. Again, fast forwarding in time to save us some more time here. Let's look at the next event that took place. Our session did collect the files from the next discard event and are saved in mount flash as Webinar Captures. Now you can pull this out to view locally on your machine as a PCAP. You can also view it on the switch directly using the tcpdump read Linux utility so I'm just going to do that here directly. So a quick tcpdump-r, for read, giving the file path for my capture. I'll just pick the last one, Webinar 4, verbose it, do see all of the headers. And then for now, let's just kept to 5 packets. So looking at the sample, we can confirm that the drops are directly related to my AI workloads because the traffic that we see that's going through congestion and likely getting dropped later are some of my AI flows. And I can confirm this because this is all RoCE V2 traffic, as you can see here. And looking at the source, it seems that source is GPU 21, which is 10.2.5.2. Also, at the same time, my packet capture is giving me a hint as to why I'm looking at congestion at TC1 and not TC3. That's because all of this traffic is coming in with a DSCP of 0, which is not being modified by the switches because we don't really have anything explicitly configured. But the switch will classify all of this as part of TC1. Now given that our PFC and ECN configurations for AI workloads were specifically configured for TC3, which is where I expect my AI flows to come in, this would explain why none of my congestion control mechanisms have helped prevent the drops by slowing things down from the center side. And from this data, I can conclude that whenever my AI jobs have included GPU 21, this has had the potential to create a huge spike in TC1, which is being left uncontrolled because we did not intend for ECN and PFC to be configured for this traffic class. This means that during high traffic volume, it can cause congestion on egress ports and can result in the discards that we are seeing. And during these events, AI workloads could experience poor performance that is being reported by the customer. So what's our solution here? Our solution is going to be to resolve this at the GPU end by updating the QoS marking there to ensure that traffic ultimately gets classified in traffic class 3. Now that being said, we also want to make sure that GPU 21 was the only GPU that was misconfigured. However, given the scale of AI networks, reviewing every GPU config manually is going to be a super painful process. So to avoid this, in an indirect way, I'm just actually just going to go back to my CloudVision instance and look up my traffic flows that are being streamed using sFlow data. I can look for the traffic flows from the time when the data was impacted, so that's October 27, and chase down any RoCE V2 traffic that is incoming with a DSCP of 0. So with that, let's move our troubleshooting back to CloudVision. Okay. So on our instance, let's do a quick sanity check of all of my traffic flows from October 27. Here, I am filtering traffic across all of my devices for any data that was incoming with a class of service of 0, as you can see here. And at the same time, also coming to a destination port of 4791, which is what my AI workloads would be using. And if I filter this data on the traffic that spans through my network, it looks like everything that's coming with a DSCP of 0 is also incoming with the source IP of 10.2.5.2, which is, again, GPU 21. If I keep scrolling through this data, I don't see any other device sourcing this traffic with this traffic class, which confirms that our problem was likely only isolated to this GPU so we should be good with what we found. We have also now been able to confirm with the customer that there have been no further issues with the workloads after the QoS mapping on GPU 21 was addressed. Okay. So with that, we are done with our debugging part. So let's go back to our slides to do a quick sum-up of our issue. So to sum up, our customer was observing poor performance for the AI workloads that started around October 27. As we gathered data, we suspected that Rack 1 and 2 GPUs might be involved, which narrowed down our initial scope to Leaf 1 and 2 and their spines. We started troubleshooting first from CloudVision, where we observed that there were periodic discards on -- and congestion events on AI Leaf 1 on Ethernet 11/5. And these events also aligned with the timeline of when our problems started. But we were surprised that ECN and PFC configurations were not helping here. So to investigate more, we moved on to AI Leaf 1 for troubleshooting. On this leaf, we saw that the drops that we were observing during the congestion periods were all on traffic class 1. And this is odd because we really didn't expect any traffic or that much traffic in this queue. So using a background tcpdump session and LANZ mirroring, we sampled one of these instances of congestion and of discards. This was mostly to determine why we were getting so much traffic in TC1 and if there was any way that they were affecting my AI workflows. So from the capture, the traffic that was experiencing congestion turned out to be RoCE V2 traffic, which was my AI-related flows received from GPU 21. They were incoming with an incorrect DSCP value of 0 from this GPU. And by default, our switches in the network will map this to traffic class 1. This was preventing flow control and ECN to help manage my congestion at the -- taking it back to the sender side since we had all of these configurations for TC3, which is where we expected our AI flows. Now while GPU 21 traffic might not always create a problem, if there was a high enough spike in traffic from this GPU, it could experience unchecked congestion in TC1, causing drops and affecting AI workloads as a whole. And while we only looked at a few switches today, this could have also translated to drops on other devices and leaves and spines in the network and just based -- depending on where GPU 21 was sending traffic. After we narrowed this down, we were still sort of unsure if GPU 21 was the only 1 misbehaving. So for completion, we went back to CloudVision and looked at historical flows to determine if there was any other AI-related traffic incoming with a DSCP of 0 in our timeline. And from here, we confirmed that only GPU 21 was the problem. Now our customer was able to confirm that this GPU was recently included in the AI workloads and that this configuration was missed. They have confirmed that the network performance is now back to normal since the QoS marking on GPU 21 has been addressed. As a rule of thumb, always a good step to ensure that whatever QoS marking you're using on your GPUs matches the configuration for ECN and PFC that you have on the switches. Now I hope this was helpful. And with that, I'm going to hand things over to Vignesh, who is going to be presenting to you the second scenario that we have for today. Thank you.

Vignesh Jothinarayanan

executive
#3

Thanks, Upasana, for the great presentation. Let me go ahead and share my screen. I hope you all can see my screen and also hear me clearly.

Neha Sakhalkar

executive
#4

Yes, we can hear you.

Vignesh Jothinarayanan

executive
#5

Okay. Hello, everyone. My name is Vignesh Jothinarayanan, and I'm the speaker for another interesting scenario we have for today that is how to troubleshoot the ECN-related issues in the AI networking fabric. Let's discuss the problem description for this scenario. In this scenario, customer was able to see a problem where they were not seeing CNP packets being received on their GPU interfaces, part of their tracker statistics. So they were expecting CNP packets to be received during those congestion events seen on those switches. And they were not seeing this so they have involved Arista team to troubleshoot this one. We all know that CNP packets or the congestion notification packet are the way to notify congestion happening between the sender and the receiver GPU in the RoCE v2 flows. And customer was not able to see these packets, and they were able to see the PFC pass frames being seen on those GPU interfaces time to time whenever they were seeing some congestion events. So they were not able to notice any visible impact seen on the performance of the fabric. However, they were worried with the fact that the ECN not working but PFC engaging early can pass many flows going over the same traffic class part of a port. As per their design, customer wanted to have ECN as the first line for congestion management and PFC to be a failsafe mechanism to help with any excessive congestion seen time to time. The reason they wanted to have this design is because ECN helps with generating congestion notification packet from the GPUs for specific flows for which the traffic needs to be reduced, the rate needs to be reduced, instead of passing on an entire traffic class for a particular port should there be any congestion event, which the PFC does. And in order to provide us with more data point, customer also provided this fact that they were able to consistently reproduce this issue whenever they were running a training job simulation test. Okay, let's also quickly discuss about the topology diagram we are dealing here. So this topology you see on the screen is the leaf and spine A networking fabric built using Arista 7060 series devices. And we are using L3 ports to connect the leaves and the spines, and it is a non-VXLAN design. And all the GPUs are connected either to a Leaf 1 or Leaf 2 via the L3 ports. And we are using BGP as a routing protocol to share the routes between the leaves and the spines. And all these ports in this topology are configured with QoS profile in order to help prioritize the RoCE v2 flows and also to enable congestion management, things like ECN and PFC on this fabric. Okay, since we have discussed about the topology diagram, let's quickly discuss how ECN-based congestion management work in AI networking fabric. Let's say, we have a sender GPU and a receiver GPU connected via a switch. Initially, the sender GPU is sending RoCE v2 flows towards the receiver GPU. And in this flow, if you notice the IP header on the ECN bits, you will be able to see ECN-capable transport bit set on these packets. And these packets go towards the receiver and the flow should be fine. Let's say after some time, the switch interface facing towards the receiver GPU is experiencing some congestion, due to which the switch will now start to mark this ECN bits part of this IP header as congestion experienced like you see on the screen. And this packet now goes to receiver GPU. On reaction to this, receiver GPU will generate a special type of packet called congestion notification packet to help notify that there is a congestion going on in between the sender and the receiver GPU. And this packet, once it reaches to the sender GPU, the sender GPU will now reduce the rate in which it is sending the traffic to the receiver GPU for those particular flows. This is how explicit congestion notification or the ECN-based congestion management works in the -- works for the RoCE V2 flows. Let's also quickly discuss how frequently the switch will be marking those traffic as congestion experience on its ECN bits should there be any congestion event. So I'm going to use this diagram here, which I put together. So let's say there is some congestion events happening on a particular port for a traffic class on a switch. And there will be some buffer usage. Let's say, the buffer usage has crossed a threshold value, which we have configured on the switches for this ECN marking. And if it's crossing that minimum threshold value, the switch will now start to mark the traffic as congestion experienced on its ECN bits. Let's say if the queue buffer usage for handling those congestion event is halfway through hitting the max threshold configuration on the switch, then 50% of the packets will be getting marked by the switches as congestion experienced. Let's say, if the queue buffer usage exceeded the max threshold configuration on the switches for those congested ports for the ECN, now 100% of the traffic will get marked. So between the minimum threshold and the maximum threshold, we see the marking will happen in a probabilistic way, part of this slope graph, which you see here. Okay, this is how the switch will decide how frequently it's going to mark the traffic should there be any congestion event. Now let's discuss the process for troubleshooting for these kind of issues. But before that, I wanted to quickly highlight some of the points we have discussed so far so that we can help use these things in the troubleshooting. So customer was not able to see the CNP packets or the Congestion Notification Packets in the network, but they were able to see the PFC pass frames, which was helping with handling the congestion time to time. And from the discussion we had now, so CNP packets are sent only when the ECN congestion experience marked packets from the switches reaches the receiver GPU. And customer wants to see the ECN engaged before the PFC in order to deal with the congestion on their fabric. And they were also able to reproduce this problem whenever they were running a training job simulation test. So how to troubleshoot these kind of issues? In order to troubleshoot this, I put together a process, which can work fairly well for these kind of issues. So that is the first step is to understand the congestion points in the network, so that could be many switches, many ports in the network. So we need to understand what are all the switches, which are experiencing the congestion time to time and what are the ports which are seeing this. And after determining that, we need to go and check whether those switches that are actually marking those traffic as congestion experienced in its ECN bits or not. So that should be checked first. Second, we need to also determine during those congestion events, how much of the buffer was actually used by those ports? And getting this value will help us compare it against the ECN settings or the configuration we have put on the switch. That is, we know that the ECN works based on the minimum threshold and the maximum threshold configuration. So we need to make sure whether the buffer usage during those congestion event is actually hitting those thresholds in order for the switch to start marking the traffic or not. Getting all these data points will help us understand these issues better and also to deploy the fixes needed to fix the larger problems for us, okay? Let's dive into each and every step now. Okay, so the first step is to understand the congestion points in the network, right? And also to determine whether the switches are actually marking the traffic or not during those congestion events, right? So in order to figure this out, there are a couple of ways. The first way is to leverage CloudVision because all these switches which we are dealing with today are streaming that telemetry data to the CloudVision. So we can quickly check that out from there to understand the congestion points and also the congested interfaces. And it is also the recommended way should there be a large number of devices, which we are involving part of our topology. But since we are dealing with very less number of devices in our topology -- in the webinar topology and also to show how to do similar things on the CLI directly, we're going to log in into the EOS CLI for the switches and check for these things, okay? So for that, let's go ahead and log in into the AI Leaf 1 to start with, okay? On the screen, you're seeing the terminal for the AI Leaf 1. In order to understand the congestion points in the network, we can leverage the Latency Analyzer feature to check for these things, right? So for that, I'm going to execute the command show-queue monitor length. Okay. On the switch, if you notice this output here, the congestion seems to be happening for the port -- sorry, ET 11/5 for its traffic class 3 time to time. And if you see the last congestion event, it started around 1 minute and 35 seconds ago and ended around 1 minute and 34 seconds ago. So if you see, the congestion event itself is a short-lived congestion event. It was lasting for close to 1 second. And if you scroll through this output here, I'm able to see only 1 port, which is congested time to time on the switch. So let's also scroll through for more lines to see whether we are seeing any other ports. So far, we were able to see only one port that is ET 11/5 for this traffic class 3 on AI Leaf 1. Let's check what's the case on AI Leaf 2. Show-queue monitor length, I'm executing the same thing here. On the AI Leaf 2, I'm able to see the port ET 11/5 for its traffic class 3, which is congested at the moment. Okay. And if you also check for the last event, so it happened 2 minutes and 32 seconds ago and ended around 2 minutes and 31 seconds ago. So even here, the congestion was lasted for close to 1 second and it was a very short-lived congestion event. And I'm not able to see any other port other than this so far, and that's the only port we are seeing congestion on this particular switch. Okay, let's also quickly check what's the case on AI Spine 1 and Spine 2. On AI Spine 1, I'm executing the same command here, show-queue monitor length. And on AI Spine 1, I'm able to see the congestion happening for the port ET 1/1 for its traffic class 3. And while checking the last congestion event, it happened 3 minutes and 19 seconds ago, as you can see on the screen, and ended around 3 minutes and 18 seconds ago. Even here, the congestion was a very short-lived congestion event, okay? And it is happening for 1 particular port that is ET 1/1 for a traffic class 3 time to time on this AI Spine 1. Let's also quickly check what's the case on AI Spine 2, which is our last device in this topology for the similar data points. On the AI Spine 2, we see the congestion happening for the port ET 1/1 for its traffic class 3, okay? And the last congestion event happened here was around 23 hours ago. So there is no active or ongoing congestion event. It happened 23 hours ago and -- yes, in a quick time frame that is close to 600 milliseconds, if you see the timestamp here, right? So that's about the ports, which are congested on this fabric. Now it's time for us to check, during those congestion events, did the switches mark the traffic as congestion experienced or not? Before checking that, I wanted to quickly validate whether the configuration needed for those ECN marking to happen is present or not. For that, on the AI Leaf 1, I'm going to check the running config. We have put together a QoS profile through which we are configuring the ECN-related things. So the QoS profile on our topology is AI Scheduler, and under the traffic class 3, we have put together the ECN-related configuration. As you can see, we have that present. And also the QoS profile is applied to all the relevant ports part of this topology. Okay, now it's time for us to check whether the switch is marking this traffic or not as congestion experienced during those events. There are a couple of ways to actually check this, right? One way is to perform a packet capture on those congested port for the outgoing traffic to see whether the IP header has those ECN bits that as congestion experienced, which is done by the switch. That is 1 way. Or another way is we can leverage the ECN counter feature, which is available on this AI Leaf 1, leaf and the spine switches we have here to check for all the ECN congestion experienced packets from the switches, okay? So we're going to leverage the ECN counter feature here for our troubleshooting. So let's go ahead and check whether the configuration is present on this device. So on the AI Leaf 1, under all the ports which are relevant in this topology, we are -- we have the traffic class 3 line put together. And under that, we are configuring the random detect ECN count, so which is the way to enable the ECN counter feature on this switch. Okay. Since we have all these things present and in order to save time, I've done similar configuration-related checks on Leaf 1, 2, Spine 1 and Spine 2, and we were able to see the ECN configurations under the QoS profile are present and the QoS profile is applied to all the relevant ports, and also the ECN counter feature is enabled on all of them, okay? Now let's leverage this ECN counter feature to check whether the switches are marking the traffic or not during those congestion event, okay? So for that, I'm going to leverage this alias which I have created. So let's check the alias I've created here. The alias I've created here is named as ECN-counters. So what this is going to do is after, whenever I'm going to execute the command ECN-counters, this is going to execute this large command for us that is show QoS interfaces ECN counters queue, which will show the ECN counters relevant to the ports. And we are using the grep filter to also filter all the ports on which the traffic class 3 where the ECN is enabled and also it will show us whether the counter is incrementing or not, okay? So since we have created the alias, let's go ahead and use it here. So I'm going to use ECN-counters to check for the switch-marked traffic during those congestion events. Okay. So while checking this counter for the congested port that is ET 11/5 for its traffic class 3, we don't see any counter increments, and the counter does not at all increment for any other ports as well. So it clearly tells us that even though there was some congestion event happening time to time on this AI Leaf 1, it is still not marking the traffic as congestion experienced in its ECN bits for the traffic. So let's go ahead and perform similar checks on the other devices as well. So I will do the same thing on AI Leaf 2. So I put together the same alias everywhere so that I can just go ahead and use the ECN counters. So even on AI Leaf 2 for its congested port that is 11/5 for this traffic class 3, we are not seeing any increments for the ECN-marked traffic. Okay, so let's go ahead and check the same thing on the AI Spine 1 and Spine 2. On the Spine 1, while checking the counters, even here, we don't see any increments happening for those congested ports. Okay, similarly, let's perform the check on the final device, ECN counters on the AI Spine 2. Okay, even on the AI Spine 2, we don't see any increments happening even though we are seeing some congestion events time to time. So this clearly tells us that time to time, we are experiencing the congestion on this fabric on some of its port but it's still not ECN marking the packets as congestion experienced. Okay. So before we move forward with the next step in our troubleshooting, I wanted to quickly validate 1 data point, which customer provided that is during the same time, PFC seems to be working and was helping with congestion management. So let's quickly validate that. For that, I'm going to execute show-priority flow control counters command on the AI Leaf 1. And we can see that ET 7/1 port, we are seeing some Tx pass frames sent all the way over to the Spine 1. And let's quickly check what's the case on AI Leaf 2 for the same output. Okay, on the AI Leaf 2, we were able to see Rx pass frames, PFC pass frames received on ET 11/5 port, which is the port which was facing the AI Leaf 1. And then the ET 12/1 is sending the pass frame downstream, okay? So it looks like during those congestion, even the AI Leaf 1 was generating the pass frame and it was creating this back pressure all the way until the sender to reduce the congestion, okay? So that tells us that even though ECN was not working, PFC was working to help with this congestion. Okay. So now move forward -- let's move forward with the next step in our checks that is, since the ECN seems to be not working, now it's time for us to validate why ECN is not working. So we wanted to understand how much of the buffer was actually used during those congestion events. Because with the buffer usage trends, we should be able to compare it against the ECN marking thresholds configured on those switches, okay? So let's go ahead and check for this thing on the switches. For that, I'm going to go back to AI Leaf 1 and perform these buffer checks. So which output will be showing us this data point here? Like how much of the buffer is used during those congestion event? As we discussed earlier, Latency Analyzer feature is one of the very useful commands to use, and that is going to show -- that is going to provide us this data point as well for us to use, okay? So I'm going to use that command here, the show-queue monitor length to check for that. So if you focus on this particular column here, that is queue length in segments, so that is -- that will be providing us that data point. That is, let's say, a port is getting congested. So this column will show us how many segments of buffer was actually used during those time or the congestion event. And each segment is of size 254 bytes. So if you do a math of multiplying the number of segment buffer used during those congestion events multiplied with 254 bytes, you will get the max buffer usage seen time to time, okay? So let's scroll through on this output to see what is the maximum buffer used by the port ET 11/5 for traffic class 3 around the congestion events. Okay, so far, we see around 2,500 is the max. We are seeing various values time to time due to the congestion events, which was seen on this switch. But the max was close to 2,500 as of now, which is an approximate value of close to 2,500 bytes -- 2,500 segments, okay? So let's do the math here. So 2,500 segments approximately multiplied into 254 bytes, which is the segment size. We are getting the max buffer usage by this port was around 635 kilobytes. So that was a max buffer used by this port time to time, okay, during those congestion event. And also, in order to save time, I've also performed similar checks on all the 4 devices in this fabric and determined that the max buffer usage was not crossing the 635 kilobytes time to time. And it was varying between 400 to 635 kilobytes approximately time to time. So that is the max buffer usage seen on these devices, okay? Now it's time for us to check, what is the threshold configuration on the devices for the ECN marking, okay? For that, I'm going to go back again to AI Leaf 1 and perform the configuration check. So we know that the ECN configuration is under the QoS profile so let's quickly check that out. On the AI Leaf 1 under the QoS profile for the traffic class 3, I'm checking the ECN threshold here. And if you see the minimum threshold is configured to be 1,000 kilobytes and maximum threshold is configured to be 1,500 kilobytes. So that means the max buffer usage seen by this fabric on its port time to time was close to 635 kilobytes. Since the minimum threshold is set to 1,000 kilobytes, it was not hitting that minimum threshold for the switch to actually start marking those traffic as congestion experienced should there be any congestion event, okay? So that is the reason the switches were not marking it and it was not helping the receiver GPU to generate the CNP packets, okay, during those congestion event. So in order to visualize this problem better, I have put together a diagram, which we can leverage here. So let's quickly check that out. So if you see this diagram here, in the bottom bar, so that kind of represents the buffer usage seen on this port time to time. So if you see during various congestion, even the buffer usage was varying and the max was close to 600 kilobytes on this fabric, but it was not hitting the minimum threshold that is 1,000 kilobytes for the switch to actually start marking those traffic as congestion experienced. So because of this, the switches were not marking those traffic and the CNP packets were not getting generated, okay? So now let's discuss the solution we have in place for this one, okay? The solution is to change the threshold settings to kind of suit our requirement here. So please note that there is no 1 value which can suit for every fabric. So it depends upon the buffer usage seen time to time on your fabric and also the requirement of how quickly the ECN needs to engage and start to mark the traffic, okay? So here, we have selected the value of minimum threshold as 256 kilobytes and maximum threshold as 512 kilobytes because it suits our fabric. And also, we need to note 1 more thing that is we should not set the minimum threshold value to a very less value because it can make the switch to mark more traffic, and it can cause more number of CNP packets generated in the network and which can pretty much bring down the performance of the fabric as well. So we need to be mindful in setting these thresholds, which kind of suits for your requirement and also the performance, the things to be kept in mind while doing these changes, okay? So let's go ahead and deploy this configuration on the switches we are dealing today. On the AI Leaf 1, I'm going ahead and configuring this. For the traffic class 3, I put together this configuration. Okay, let's go ahead and do the same on the AI Leaf 2 as well. Okay, I'll do the same on the AI Spine 1 and Spine 2. Because I was able to see pretty much similar buffer usage seen on these ports time to time on all these 4 devices, so I'm putting the same threshold settings everywhere. Okay. So since we have configured the changes now, it's time for us to validate whether it helped resolve the problem we were seeing, okay? So let's go back to AI Leaf 1 and use the ECN counters which we were using earlier. So ECN counters. Okay, now if I check the port ET 11/5 for traffic class 3, which was experiencing the congestion time to time, it has started to mark the traffic as expected. Okay, that's a good sign. Okay. Let's go back -- let's go and check the AI Leaf 2. On the AI Leaf 2 as well for the port ET 11/5 for its traffic class 3, we see the marking happening now. Okay, that's good. So let's quickly validate AI Spine 1 and Spine 2 as well. Even on AI Spine 1, the ET 1/1 port for its traffic class 3 now started to mark the traffic as congestion experienced. Okay. On the AI Spine 2, let's do the check. Okay. On the AI Spine 2, we still don't see any increments on these counters. Okay, we have performed the changes even on AI Spine 2 but still we don't see any increments. Okay, so let's quickly check whether there is any ongoing congestion on AI Spine 2 at the moment, okay? For that, let's check the queue monitor length. Okay, so that tells us why. On the queue monitor length, if you see the last congestion event was happened 1 day ago. So there is no active congestion event happening on the AI Spine 2 because of that, even though we changed the settings, we still did not see any congestion event for the switch to actually start marking the traffic as congestion experienced, okay? But on the same time, AI Spine 1 was marking the traffic, right? So let's quickly check whether it has seen some -- whether it has seen any congestion events recently. Okay, so on AI Spine 1, it was marking the traffic as congestion experienced because it saw 1 congestion event 2 minutes ago, and it was crossing the threshold as per our math here. And so that's the reason on AI Spine 1, we started to see the marking happening, okay? So does other switches as well. Okay, so let's quickly discuss the troubleshooting summary here. Initially, customer did not see the CNP packets flowing through their network even though they were seeing some congestion events. However, they were seeing the PFC pass frames helping with the congestion time to time, and they involved Arista team to troubleshoot this. To start with, we kind of determined all the congestion points in the network using the Latency Analyzer feature from the CLI. And also, we checked the ECN counters to see whether during those time, the switches marked the traffic or not as congestion experienced. And we were able to see the switches did not mark the traffic as expected. And later, we determined what was the max buffer usage trends seen on the switches time to time during those congestion events. And also, we it against the ECN conflicts and also the settings, okay? So if you see this diagram here, the buffer usage we found was around 600 kilobytes. That's the max, which was seen time to time on these ports on this fabric. And it was not hitting the minimum threshold configuration that is 1,000 kilobytes, which was set on the QoS profile on these switches. So because of this reason, the switches were not ECN marking the traffic as congestion experienced, okay? And later, we changed the ECN settings to suit the requirement of this fabric. And with that, we were able to see the issue resolved, that the switch now started to mark the traffic as congestion experienced whenever they were seeing the congestion events, okay? So that's the end of this troubleshooting scenario. So we have put together all the reference links for the tools, which we have been using in both of these scenarios here. So we'll be sharing this slide with you all so that you can check this out and also try to see whether you can use this as part of your day-to-day troubleshooting and dealing with the fabrics, okay? So with that, I will hand it over to Neha back again for the closing note. Thanks, everyone.

Neha Sakhalkar

executive
#6

Thank you, Vignesh -- actually, thank you, both Upasana and Vignesh. That was an incredibly detailed presentation. We hope that the session was beneficial to all of our attendees as well. Today's recording will be posted on our community page shortly. And in addition to the webinar recordings, our community page also includes a wide variety of articles on some other technologies written by our in-house experts. Feel free to leverage our knowledge base for gaining additional insights on our platforms and feature sets. Now with the growing interest in AI networking, we are excited to announce that we will be back next year after a short holiday season break with another AI-focused session. This session is currently scheduled on 29th Jan at 11:00 a.m. Eastern. We would love to hear your feedback. If you found today's session valuable and would love -- and if you want to see more such topics from Arista TAC, please share your feedback on the link shared in our chat right now. So now I'll keep the session open for a few minutes in case there are any pending questions. Our panelists will be available to answer them live or via Q&A.

This call discussed

For developers and AI pipelines

Programmatic access to Arista Networks, Inc. earnings transcripts and 248,000+ others is available through the EarningsCalls.dev REST API. Plans from $24.99/month — full transcripts, speaker segments, full-text search, and the recently-added /api/v1/transcripts/recent polling endpoint for ETL pipelines.