Lab 13 — Capstone: Troubleshooting Challenge¶
Companion to the NCP-AIN Certification Guide
This free lab is part of the hands-on companion to NCP-AIN Certification Guide by Vakeesan Thevarajah (Cloudfoxy Ltd). The book explains the theory, design choices and hardware behaviour behind every step.
Lab at a glance¶
| Book chapters | Chapter 4 (NVUE) · 5 (RoCE QoS) · 7 (BGP-EVPN) · 11–13 (InfiniBand, SM HA, PKeys) · 17 (IB diagnostic tools) · 22 (Ansible) · 23 (exam strategy) |
| Exam objectives | Domain 5 Troubleshooting (5.4, 5.5) plus the objectives behind each fault: 2.1, 2.2, 2.3, 3.1, 3.2, 3.4, 6.1, 6.2 |
| Time | 3–4 hours for all ten scenarios (about 15–20 minutes each), or run them in two sessions |
| Air resources used | The whole helix-b-air simulation (about 26 vCPU while running). A favourite checkpoint taken before you start |
| Prerequisite labs | Labs 0–5, Lab 6 (checkpoints), Lab 8 (ibsim set up in ~/ibsim on hxa-ibsim), Lab 12 (Ansible project). Labs 7, 9 and 11 help |
Objectives¶
By the end of this lab you will be able to:
- Work a fault from a vague ticket to a verified fix using a layer-by-layer method, not guesswork.
- Diagnose Ethernet faults on the Helix-B twin: a BGP unnumbered neighbour that won't establish, an MTU mismatch, a VNI mismatch, a missing anycast gateway, a disabled lossless RoCE profile, and a wrong VLAN pushed by Ansible.
- Diagnose InfiniBand faults in ibsim: a missing standby Subnet Manager and SM failover, a missing partition member, rising symbol errors on an uplink, and a failed link.
- Prove every fix with the same commands that showed the fault.
- Write troubleshooting notes that a colleague (or an examiner) can follow.
Background¶
This lab turns everything you built in Labs 0–12 into a practice ground. An instructor, a study partner or a script breaks one thing at a time. You get only a short ticket, written the way a user would report it. Your job is to find the fault, fix it, and prove the fix. The exam's troubleshooting domain (20%) tests exactly this skill in multiple-choice form: given symptoms and command output, pick the cause and the next step.
There are ten scenarios: six on the Helix-B Ethernet twin and four on the Helix-A InfiniBand simulator. Each scenario lists how the fault is injected (instructor only), what the student sees, three hints that give away progressively more, and the exam objective it practises. Full solutions are in a separate section so you don't read them by accident.
Use the same method every time. Start from the symptom and the scope (one host? one rail? one tenant? whole fabric?). Then work up the layers: physical link, then L2 (VLAN, MTU), then the underlay control plane (BGP), then the overlay (EVPN, VNIs, VRFs, gateways), then QoS and policy, and last the automation that put the configuration there. On InfiniBand the order is link state, then SM and LIDs, then partitions, then counters and routes. At every step, compare the broken device with a healthy peer: in a rail-optimised fabric there is always a twin that should look the same.
Warning
Cumulus VX has no Spectrum ASIC, and ibsim has no real links. Some faults don't produce the traffic symptoms they would in production (for example, lossy RoCE on VX doesn't drop anything, and ibsim symbol errors only rise when someone sets them). In those scenarios you diagnose from configuration and counters, which is also what most exam questions show you.


The figure shows only the links involved in the faults. Every leaf connects to both spines (swp31 to hxb-spine01, swp32 to hxb-spine02).
Step-by-step¶
Task 1 — Capture a baseline and take a checkpoint¶
You can only recognise "broken" if you have recorded "working". Capture a baseline from the healthy fabric before any fault is injected.
-
On
oob-mgmt-server, create a working folder and capture the Ethernet baseline:ubuntu@oob-mgmt-server:~$ mkdir -p ~/capstone/baseline && cd ~/capstone/baseline ubuntu@oob-mgmt-server:~/capstone/baseline$ for sw in hxb-spine01 hxb-spine02 hxb-leaf-r1 hxb-leaf-r2 hxb-leaf-r3 hxb-leaf-r4; do ssh cumulus@$sw "sudo vtysh -c 'show bgp summary'; sudo vtysh -c 'show evpn vni'; nv show qos roce; nv config show" > $sw.txt 2>&1 done ubuntu@oob-mgmt-server:~/capstone/baseline$ lsExpected output (illustrative):
-
Record the host reachability matrix. From
hxb-gpu01, every AURORA address should answer:ubuntu@hxb-gpu01:~$ for ip in 172.16.10.1 172.16.10.102 172.16.11.1 172.16.11.102; do ping -c 2 -W 1 $ip >/dev/null && echo "$ip ok" || echo "$ip FAIL"; doneExpected output (illustrative): four lines ending in
ok. Repeat fromhxb-gpu03for the BOREALIS addresses (172.17.10.1, 172.17.10.104, 172.17.11.1, 172.17.11.104). Check that a cross-tenant ping (for exampleping 172.17.10.103fromhxb-gpu01) fails, as Lab 5 intended. -
Capture the InfiniBand baseline on
hxa-ibsim. Start ibsim and OpenSM the way Lab 8 does, then:ubuntu@hxa-ibsim:~$ cd ~/ibsim ubuntu@hxa-ibsim:~/ibsim$ ibsim-run sminfo > base-sminfo.txt ubuntu@hxa-ibsim:~/ibsim$ ibsim-run saquery -s > base-sm-ports.txt ubuntu@hxa-ibsim:~/ibsim$ ibsim-run ibnetdiscover > base-topo.txt ubuntu@hxa-ibsim:~/ibsim$ ibsim-run iblinkinfo --switches-only -l > base-links.txt ubuntu@hxa-ibsim:~/ibsim$ ibsim-run ibqueryerrors > base-errors.txt ubuntu@hxa-ibsim:~/ibsim$ cp partitions.conf base-partitions.confsaquery -slists the ports that are SM-capable. With Lab 8's HA set-up it should show both SM hosts, so you know the master's and the standby's LIDs. -
Save everything and take a checkpoint. On every switch,
nv config diff applied startupshould be empty (Lab 6, Task 1). Then in Air use Stop Simulation and Store Checkpoint, mark the checkpoint as a favourite, and name itpre-capstone. Start the simulation from it.Warning
An Air checkpoint saves disks, not running processes. After starting from a checkpoint, restart ibsim and the OpenSM instances as in Lab 8 before you do any InfiniBand scenario.
Checkpoint:
- Baseline files exist in
~/capstone/baselineand~/ibsim/base-*. - A favourite checkpoint called
pre-capstoneexists.
Task 2 — Install the fault injector (instructor or study partner)¶
The student should not read this task. If you work alone, ask a friend to install the script, or install it without reading the case block and use random mode (Task 3).
-
On
oob-mgmt-server, create~/capstone/inject.sh:#!/bin/bash # Lab 13 fault injector. Usage: inject.sh F1..F10 | F7b | random set -euo pipefail SVI_AF=ipv4 # use "ip" if Lab 5 used the pre-5.15 SVI syntax ANSIBLE_DIR=~/helix-ansible # your Lab 12 Ansible project SIM=hxa-ibsim VICTIM_GUID=0x0000000000000000 # port GUID of one PRODUCTION member (F8), from partitions.conf nv_on() { local sw=$1; shift; ssh -o BatchMode=yes cumulus@"$sw" "$* && nv config apply -y" >/dev/null; } sim_cmd() { ssh -o BatchMode=yes ubuntu@"$SIM" "echo '$1' > ~/ibsim/simcmd"; } id=${1:?usage: inject.sh F1..F10|F7b|random} if [ "$id" = random ]; then id=F$(( (RANDOM % 10) + 1 )) echo "$id $(date)" | sudo tee -a /root/capstone-answers >/dev/null echo "A fault has been injected. Good luck." fi case $id in F1) nv_on hxb-spine02 "nv set vrf default router bgp neighbor swp3 remote-as 65200" ;; F2) nv_on hxb-spine01 "nv set interface swp2 link mtu 1500" ;; F3) nv_on hxb-leaf-r4 "nv unset bridge domain br_default vlan 111 vni 10111 && nv set bridge domain br_default vlan 111 vni 10101" ;; F4) nv_on hxb-leaf-r3 "nv unset interface vlan110 $SVI_AF vrr" ;; F5) nv_on hxb-leaf-r2 "nv set qos roce mode lossy" ;; F6) cd "$ANSIBLE_DIR" cat > playbooks/access-ports.yml <<'EOF' - name: Tidy leaf access ports hosts: hxb-leaf-r2 gather_facts: false tasks: - name: Set access VLAN on swp2 nvidia.nvue.api: operation: set force: true wait: 15 data: interface: swp2: bridge: domain: br_default: access: 111 EOF git add playbooks/access-ports.yml && git commit -qm "Tidy access ports" ansible-playbook playbooks/access-ports.yml >/dev/null ;; F7) ssh ubuntu@"$SIM" "pkill -f '[o]pensm.*-p 12'" ;; # standby SM disappears F7b) ssh ubuntu@"$SIM" "pkill -f '[o]pensm.*-p 14'" ;; # later: master fails too F8) ssh ubuntu@"$SIM" "cd ~/ibsim && sed -i.bak 's/${VICTIM_GUID}[^,;]*, *//; s/, *${VICTIM_GUID}[^,;]*//' partitions.conf && pkill -HUP -f '[o]pensm.*-p 14'" ;; F9) ( for n in 400 1500 4200 9800; do sim_cmd "PerformanceSet \"hxa-su1-leaf-r1\"[33] PortCounters.SymbolErrorCounter=$n"; sleep 120 done ) >/dev/null 2>&1 & ;; F10) sim_cmd 'Unlink "hxa-su2-leaf-r2"[34]' ;; *) echo "unknown fault $id"; exit 1 ;; esac -
Make it executable and check it parses:
Notes on the script (check on your set-up):
nv_onassumes the passwordless SSH fromoob-mgmt-serverto the switches that Lab 0 and Lab 12 set up. Without it,BatchMode=yesmakes the script fail instead of hanging on a password prompt.SVI_AFmust match the SVI syntax you used in Lab 5. From Cumulus Linux 5.15, NVUE usesinterface vlan110 ipv4 vrr. Earlier 5.x releases useip vrr.- F6 uses the Lab 12 connection settings (
nvidia.nvue.apiover the NVUE REST API on port 8765). If your Lab 12 inventory uses a different project path or host names, adjustANSIBLE_DIRandhosts. - F7, F7b and F8 assume that Lab 8 started the master OpenSM with priority 14 (
-p 14) and the standby with priority 12 (-p 12), and that the partition file is~/ibsim/partitions.conf. The[o]pensmpattern stopspkillfrom matching its own shell. OpenSM re-reads the partition file on its next heavy sweep, whichSIGHUPtriggers (check on your OpenSM version; if the change doesn't appear, restart the master OpenSM the Lab 8 way). - The ibsim console reads commands from the named pipe
~/ibsim/simcmdthat Lab 8 creates. If nothing is reading the pipe,echo … > simcmdblocks. In the Lab 8 topology, leaf ports 33 and above are spine uplinks; confirm the port numbers in~/ibsim/base-links.txt.
Task 3 — Run the challenge¶
With an instructor: the instructor runs one scenario at a time (~/capstone/inject.sh F3, for example), hands the student the ticket text from the Break and fix section, and starts a timer.
On your own: run ~/capstone/inject.sh random. The script picks a fault, writes the answer to /root/capstone-answers (readable only with sudo), and tells you nothing else. Read the ticket list in Break and fix, decide which ticket matches what you see, and start.
Rules:
- Target 15–20 minutes per scenario. After 10 minutes stuck, take Hint 1; after 15, Hint 2.
- Change one thing at a time. On switches, review
nv config diffbefore everynv config apply. - Prove the fix with the same test that showed the fault.
- Restore before the next scenario. Either apply the fix from the solution, or start again from the
pre-capstonecheckpoint (Lab 6). A checkpoint restore is the only guaranteed clean slate, but it takes several minutes and you must restart ibsim afterwards.

Task 4 — Keep troubleshooting notes¶
Good notes are part of the score. Copy this template into ~/capstone/notes-F<n>.md at the start of each scenario and fill it in as you go:
# Ticket F<n> — <one-line summary in your own words>
Start: 14:05 Finish: 14:22 Hints used: 0
## Symptom and scope
- What fails: gpu02 eth1 (172.16.10.102) cannot reach its gateway 172.16.10.1
- What still works: gpu02 → gpu01 172.16.10.101 (same VLAN, other leaf)
- Scope: one leaf (hxb-leaf-r3), one VLAN (110); leaf-r1 VLAN 110 fine
## Evidence (command → key line)
- gpu02: ip neigh show 172.16.10.1 → FAILED
- leaf-r3: nv show interface vlan110 → no vrr address
- leaf-r1: nv show interface vlan110 → vrr 172.16.10.1/24, 00:00:5e:00:01:01
## Root cause
VRR (anycast gateway) removed from vlan110 on hxb-leaf-r3 (nv config history: rev 41).
## Fix (exact commands)
nv set interface vlan110 ipv4 vrr address 172.16.10.1/24 ... ; nv config diff ; nv config apply
## Verification (same test as the symptom)
- gpu02: ping 172.16.10.1 → 0% loss; ping 172.16.11.101 → 0% loss
## Prevention
Gateway settings come from the Lab 11 template; add a post-change ping check to the pipeline.
What makes notes good:
| Good notes | Weak notes |
|---|---|
| State scope: what works as well as what fails | "Network is broken" |
| Quote the one line of output that proves each point | Paste 200 lines of output with no comment |
| Name the device, interface, VLAN/VNI or port and LID | "Something on the leaf" |
| Separate the root cause from the symptom | "Fixed by rebooting" |
| Show the exact fix and the exact verification | "Changed config, works now" |
| End with prevention: template, check or alert | No lesson recorded |
Task 5 — Score yourself¶
Score each scenario out of 10 with this rubric. An instructor scores from the student's notes and a short verbal walk-through.
| Criterion | 0 | 1 | 2 | 3 |
|---|---|---|---|---|
| Scope and isolation (max 2) | No scope stated | Right area (Ethernet or IB, one tenant or rail) | Exact device and interface, VLAN or port | — |
| Root cause (max 3) | Wrong or none | Right layer, wrong detail | Right cause, weak evidence | Right cause, proved with the key output line |
| Fix (max 2) | None, or it broke something else | Worked but untidy (several changes, no diff review) | Minimal, reviewed with nv config diff or equivalent |
— |
| Verification (max 2) | None | A different test from the symptom | Same test as the symptom, plus a check that nothing else broke | — |
| Notes (max 1) | Missing or unreadable | Complete template, including prevention | — | — |
| Adjustment | Points |
|---|---|
| Each hint used | −1 (Hint 3 counts as two) |
| Solved in under 10 minutes | +1 |
| Needed the solution | Score the notes and verification only |
| Total over ten scenarios | Rating |
|---|---|
| 85–110 | Exam-ready for Domain 5: fast, evidence-based, tidy |
| 65–84 | Solid. Revisit the chapters for the scenarios where you lost points |
| Below 65 | Repeat the lowest-scoring scenarios after re-reading their chapters and labs |
Break and fix¶
Each scenario is written in four parts: Ticket (all the student gets), Injection (instructor only), Symptoms (what the student should find), and Hints. Full solutions follow in the next section.
F1 — "Rail 3 is running on one leg"¶
Ticket: "Monitoring says hxb-leaf-r3 has only one path to the spines. Jobs still run. Can you check before tonight's training run?"
Injection: inject.sh F1. On hxb-spine02, the unnumbered neighbour on swp3 (the link to hxb-leaf-r3 swp32) is changed from remote-as external to remote-as 65200.
Symptoms:
- On
hxb-leaf-r3,sudo vtysh -c "show bgp summary"showsswp31Established butswp32inActive,ConnectorIdle, with 0 prefixes received. sudo vtysh -c "show ip route 10.255.0.11"on leaf-r3 shows one next hop (via swp31) instead of two.- The physical link is up:
nv show interface swp32showsup, and LLDP showshxb-spine02 swp3. - On
hxb-spine02,show bgp neighbor swp3shows an OPEN error / bad peer AS as the last reset reason (wording varies by FRR version).
Hints:
- The link is up. Which layer is left, and which end would you look at?
- Compare
nv show vrf default router bgp neighbor swp3on spine02 withswp1on the same spine. - With BGP unnumbered, both ends normally say
remote-as external. What happens if one end expects a specific AS?
Exam objective: 2.3 (the BGP underlay that EVPN runs on) · 6.1 (finding the change with nv config history) · 1.2 (rail-optimised design: why ECMP to both spines matters).
F2 — "Small pings work, big transfers hang"¶
Ticket: "From hxb-gpu01 to hxb-gpu02 on AURORA VLAN 111, ping works but copying a large file with scp stalls. Other pairs seem fine."
Injection: inject.sh F2. On hxb-spine01, swp2 (the link to hxb-leaf-r2 swp31) gets link mtu 1500. The leaf side stays at 9216.
Symptoms:
ping 172.16.11.102fromhxb-gpu01(over eth2, leaf-r2 → spine → leaf-r4) works with default-size packets.- A full-size, don't-fragment ping fails or partly fails:
ping -M do -s 1472 -c 5 172.16.11.102. VXLAN adds 50 bytes, so a 1500-byte host packet becomes a 1550-byte underlay packet, which doesn't fit a 1500-byte link. Whether it fails depends on which spine the flow hashes to. - BGP stays Established on every link. Unlike OSPF, BGP doesn't check MTU, so the control plane looks healthy.
- An underlay test from
hxb-leaf-r2tohxb-spine01's loopback with a jumbo packet fails, while the same test tohxb-spine02works.
Hints:
- Small works, large fails: think about packet size before you think about routing.
- Test the underlay between leaf-r2 and each spine with
ping -M do -s 8972to the spine loopbacks. - Compare the MTU on both ends of each leaf-r2 uplink.
Exam objective: 2.1 (link settings on Spectrum-X switches, MTU 9216) · 2.3 (VXLAN overhead in an EVPN fabric) · 5.3 (verifying the data path, not just the control plane).
F3 — "gpu02 can't see gpu01 on VLAN 111"¶
Ticket: "After last night's maintenance on rail 4, hxb-gpu02 eth2 (172.16.11.102) can't reach hxb-gpu01 eth2 (172.16.11.101). It can reach its gateway."
Injection: inject.sh F3. On hxb-leaf-r4, VLAN 111 is re-mapped from VNI 10111 to VNI 10101.
Symptoms:
- From
hxb-gpu02:ping 172.16.11.1works (the gateway is local VRR on leaf-r4).ping 172.16.11.101fails. - Traffic between different subnets may still work (for example gpu02 eth2 to 172.16.10.101), because symmetric routing uses the L3 VNI 104001, not the L2 VNI. That partial success is a strong clue.
- On leaf-r4,
sudo vtysh -c "show evpn vni"lists VNI 10101 with no remote VTEPs, and no 10111. - On leaf-r2,
sudo vtysh -c "show evpn vni 10111"lists remote VTEPs 10.255.0.11 and .13 as usual, but not 10.255.0.14.
Hints:
- Same subnet fails, gateway works. Which part of EVPN carries same-subnet traffic between leaves?
- Run
show evpn vnion leaf-r4 and on leaf-r2 and compare the VNI lists. - Look at
nv show bridge domain br_default vlan 111on leaf-r4.
Exam objective: 2.3 (L2 VNI mapping, route type 2/3, symmetric IRB).
F4 — "One server can't leave its subnet"¶
Ticket: "hxb-gpu02 eth1 (172.16.10.102) can't reach anything outside 172.16.10.0/24. It can still ping hxb-gpu01 on 172.16.10.101."
Injection: inject.sh F4. On hxb-leaf-r3, the VRR (anycast gateway) configuration is removed from vlan110.
Symptoms:
- From
hxb-gpu02:ping 172.16.10.101works (L2 across VNI 10110).ping 172.16.10.1andping 172.16.11.101fail. ip neigh show 172.16.10.1onhxb-gpu02showsFAILEDorINCOMPLETE: nothing answers ARP for the gateway on leaf-r3.- On leaf-r3,
nv show interface vlan110has no VRR address. On leaf-r1, the same command shows172.16.10.1/24and MAC00:00:5e:00:01:01. - BGP and EVPN sessions are all healthy.
Hints:
- The host can bridge but not route. Who is the host's router?
- Check the host's ARP entry for 172.16.10.1.
- Compare
nv show interface vlan110on leaf-r3 and leaf-r1.
Exam objective: 2.3 (distributed anycast gateway with VRR on every leaf).
F5 — "RoCE audit failed on rail 2"¶
Ticket: "The weekly RoCE audit flagged hxb-leaf-r2. perftest still runs. Is it a real problem?"
Injection: inject.sh F5. On hxb-leaf-r2, nv set qos roce mode lossy is applied. (Variant for a harder round: change the switch's QoS trust so DSCP is no longer trusted, for example nv set qos mapping default-global trust l2; check on your release whether VX accepts it and how it interacts with the RoCE profile.)
Symptoms:
- Soft-RoCE perftest still works. VX has no Spectrum ASIC, so there is no PFC, pause or buffer behaviour to observe, and nothing drops.
nv show qos roceon leaf-r2 showsmode lossyand no PFC on switch priority 3. Leaf-r1 showsmode losslesswith PFC on priority 3, ECN enabled and trustpcp,dscp.nv config historyon leaf-r2 shows a recent revision that changed it.
Hints:
- VX can't show you drops. What can it show you?
- Diff
nv show qos rocebetween leaf-r2 and leaf-r1. - Which single NVUE command sets the lossless defaults, and what does
mode lossyremove?
Exam objective: 2.1 (configure RoCE with nv set qos roce) · 2.2 (PFC and ECN) · 5.3 (verify the low-latency, lossless configuration end to end).
F6 — "BOREALIS host sees AURORA broadcast traffic"¶
Ticket: "hxb-gpu03 eth2 (BOREALIS, 172.17.11.103) lost its gateway this morning. tcpdump on it shows ARP requests for 172.16.11.x addresses, which belong to another tenant."
Injection: inject.sh F6. A new playbook, committed to the Lab 12 Git repository as "Tidy access ports", sets hxb-leaf-r2 swp2 to access VLAN 111 (AURORA) instead of 211 (BOREALIS).
Symptoms:
- From
hxb-gpu03:ping 172.17.11.1andping 172.17.11.104fail. sudo tcpdump -ni eth2 arponhxb-gpu03shows ARP for 172.16.11.x. The host is in the wrong broadcast domain, which is a tenant isolation breach, not only an outage.- On leaf-r2,
nv show interface swp2 bridge domain br_defaultshows access VLAN 111. On leaf-r1,swp2shows 210, and on leaf-r4,swp2shows 211. nv config historyon leaf-r2 shows a revision applied through the API by the automation user, not an interactive CLI user.git log --oneline -3in the Ansible project shows the "Tidy access ports" commit.
Hints:
- Which VLAN is gpu03's port in, and which should it be in?
- Who changed it?
nv config historytells you whether it came from a person or from the API. - Fix the source of truth, not just the switch. Otherwise the next playbook run puts the fault back.
Exam objective: 6.2 (Ansible playbooks for VLANs) · 6.1 (NVUE configuration management and history) · 2.3 (tenant isolation).
F7 — "Confirm Subnet Manager HA before the change freeze"¶
Ticket: "Change freeze starts tomorrow. Please confirm the Helix-A SM pair is healthy: one master, one standby." (If the student reports HA as healthy, the instructor later runs inject.sh F7b and sends a second ticket: "New nodes on Helix-A are stuck in Initializing and sminfo times out.")
Injection: inject.sh F7 stops the standby OpenSM (priority 12). inject.sh F7b later stops the master (priority 14).
Symptoms (stage 1):
ibsim-run sminfostill shows a master (priority 14, stateSMINFO_MASTER). Everything looks fine at first glance.ibsim-run sminfo <standby LID>(the standby LID frombase-sm-ports.txt) doesn't return a standby SM (it times out or shows it not active), andps -ef | grep [o]pensmshows only one OpenSM.
Symptoms (stage 2, after F7b):
ibsim-run sminfotimes out. There is no master SM.- Existing routes keep working, because the switches keep their programmed forwarding tables, but ports that come up now stay in
Initializing.
Hints:
- "One master" is only half the question. How do you query a specific SM rather than the master?
- Compare
ps -ef | grep [o]pensmwith how Lab 8 started the HA pair. - SM election uses the highest priority, then the lowest GUID. Which priority must the restarted standby have so that it doesn't take over?
Exam objective: 3.1 (SM high availability) · 5.4 (SM and UFM health) · 5.5 (sminfo, saquery).
F8 — "One PRODUCTION node can't join the job"¶
Ticket: "A node on Helix-A can't talk to the other PRODUCTION nodes over the PRODUCTION partition. Everything else about it looks normal."
Injection: inject.sh F8 removes one member (VICTIM_GUID) from the PRODUCTION line of ~/ibsim/partitions.conf and makes the master SM re-read the file.
Symptoms:
- The node's port is
Activeinibsim-run iblinkinfo, and it has a LID. ibsim-run smpquery pkeys <LID>for that node lists the default partition but not 0x8002 (PRODUCTION, full member). A healthy PRODUCTION node lists0x8002.diff ~/ibsim/base-partitions.conf ~/ibsim/partitions.confshows the missing GUID (and there is apartitions.conf.bak).
Hints:
- The port is up and has a LID, so the SM has seen it. What else does the SM program into a port?
- Read the node's PKey table with
smpquery pkeysand compare it with a healthy PRODUCTION node. - Where does OpenSM get partition membership from?
Exam objective: 3.2 (configure PKeys: partitions.conf, full and limited membership, 0x8002 = full member of 0x0002).
F9 — "Errors climbing on an SU1 uplink"¶
Ticket: "UFM-style health alerts every few minutes about hxa-su1-leaf-r1. No outage yet."
Injection: inject.sh F9 sets SymbolErrorCounter on hxa-su1-leaf-r1 port 33 (a spine uplink) to 400, 1500, 4200 and 9800 at two-minute intervals, using the ibsim console command PerformanceSet "hxa-su1-leaf-r1"[33] PortCounters.SymbolErrorCounter=N.
Symptoms:
ibsim-run ibqueryerrorslistshxa-su1-leaf-r1port 33 with a non-zeroSymbolErrorCounter. Running it again a few minutes later shows a higher value.- The link is still
LinkUp/Activeat full width inibsim-run iblinkinfo --switches-only -l. base-errors.txtfrom Task 1 shows no errors on that port.
Hints:
- The link is up. Which tool lists only ports with non-zero error counters?
- One reading tells you little. What does a rising symbol error count mean?
- Which end of the link, and which cable, would you replace, and how do you confirm the fix in the simulator?
Exam objective: 5.5 (ibqueryerrors, perfquery, iblinkinfo) · 3.4 and 5.4 (monitoring link health and acting on alerts).
F10 — "SU2 lost bandwidth to the spines"¶
Ticket: "Jobs spanning both SUs slowed down about 3% after lunch. Nothing is down, as far as we can see."
Injection: inject.sh F10 sends Unlink "hxa-su2-leaf-r2"[34] to the ibsim console through the named pipe (echo 'Unlink "hxa-su2-leaf-r2"[34]' > ~/ibsim/simcmd).
Symptoms:
ibsim-run iblinkinfo --switches-only -lshowshxa-su2-leaf-r2port 34 asDown(physical statePollingorDisabled, depending on the ibsim build), with no peer.diff ~/ibsim/base-links.txt <(ibsim-run iblinkinfo --switches-only -l)highlights that one line.ibsim-run ibqueryerrorsmay show nothing useful: in the simulator, unlinking doesn't necessarily raise error counters. That is why the baseline diff matters.- The fabric keeps working. After its next sweep, the SM routes around the missing uplink, so only bandwidth drops.
Hints:
- "Slower, nothing down" often means a missing link in an ECMP-like fabric. Which tool shows every link?
- Compare the current
iblinkinfo --switches-only -lwith your baseline. - Which spine and port was leaf-r2 port 34 connected to? Your baseline knows.
Exam objective: 5.5 (iblinkinfo, ibnetdiscover) · 3.1 (fabric provisioning and link validation).
Solutions¶
F1 solution — wrong remote-as on spine02¶
- Scope: only leaf-r3's
swp32session is down. Its other uplink is fine, and other leaves' sessions to spine02 are fine, so the problem is this one link or its two ends. - The physical layer is fine (
nv show interface swp32is up, LLDP shows the right neighbour). -
On
hxb-spine02, compare neighbours:cumulus@hxb-spine02:mgmt:~$ nv show vrf default router bgp neighbor swp3 cumulus@hxb-spine02:mgmt:~$ nv show vrf default router bgp neighbor swp1 cumulus@hxb-spine02:mgmt:~$ nv config historyExpected output (illustrative):
swp3showsremote-as 65200andswp1showsremote-as external. The history shows the revision that changed it. -
Fix and verify:
cumulus@hxb-spine02:mgmt:~$ nv set vrf default router bgp neighbor swp3 remote-as external cumulus@hxb-spine02:mgmt:~$ nv config diff cumulus@hxb-spine02:mgmt:~$ nv config apply -y cumulus@hxb-leaf-r3:mgmt:~$ sudo vtysh -c "show bgp summary" cumulus@hxb-leaf-r3:mgmt:~$ sudo vtysh -c "show ip route 10.255.0.11"
Checkpoint: swp32 is Established with prefixes received, and the route to 10.255.0.11 has two next hops again.
Root cause: the peer AS on spine02 swp3 didn't match leaf-r3's real AS (65103). The leaf's OPEN carried AS 65103, the spine expected 65200, and the session was rejected. remote-as external accepts any AS different from the local one, which is why unnumbered fabrics use it.
F2 solution — MTU mismatch on spine01 swp2¶
-
Prove it's size-related from the host:
-
Test the underlay from
hxb-leaf-r2to each spine loopback. Your SSH session is in the management VRF (the:mgmt:in the prompt), so run the ping in the default VRF:cumulus@hxb-leaf-r2:mgmt:~$ sudo ip vrf exec default ping -M do -s 8972 -c 3 10.255.0.1 cumulus@hxb-leaf-r2:mgmt:~$ sudo ip vrf exec default ping -M do -s 8972 -c 3 10.255.0.2Expected output (illustrative): the ping to 10.255.0.2 (spine02) gets replies. The ping to 10.255.0.1 (spine01) shows 100% packet loss or
message too long. -
Compare both ends of the leaf-r2 → spine01 link:
cumulus@hxb-leaf-r2:mgmt:~$ nv show interface swp31 link | grep -i mtu cumulus@hxb-spine01:mgmt:~$ nv show interface swp2 link | grep -i mtuExpected output (illustrative): leaf-r2 swp31 shows 9216, spine01 swp2 shows 1500.
-
Fix and verify:
Checkpoint: both pings succeed with no loss.
Root cause: one side of a routed link had MTU 1500. BGP stayed up because its packets are small and BGP doesn't negotiate MTU. VXLAN-encapsulated full-size frames were dropped on that link only, so failures depended on the ECMP hash.
Version note
How Cumulus VX handles an oversized received frame (drop, or accept) depends on the virtual NIC. If the jumbo ping from leaf-r2 to spine01 still works in your simulation, test in the other direction (from spine01 to 10.255.0.12), where spine01's 1500-byte egress MTU applies.
F3 solution — VNI mismatch on leaf-r4¶
cumulus@hxb-leaf-r4:mgmt:~$ sudo vtysh -c "show evpn vni"
cumulus@hxb-leaf-r4:mgmt:~$ nv show bridge domain br_default vlan 111
cumulus@hxb-leaf-r2:mgmt:~$ sudo vtysh -c "show evpn vni 10111"
Expected output (illustrative): leaf-r4 lists VNI 10101 (L2, 0 remote VTEPs) and no 10111. Leaf-r2's VNI 10111 lists remote VTEPs 10.255.0.11 and 10.255.0.13 only.
Fix and verify:
cumulus@hxb-leaf-r4:mgmt:~$ nv unset bridge domain br_default vlan 111 vni 10101
cumulus@hxb-leaf-r4:mgmt:~$ nv set bridge domain br_default vlan 111 vni 10111
cumulus@hxb-leaf-r4:mgmt:~$ nv config diff
cumulus@hxb-leaf-r4:mgmt:~$ nv config apply -y
cumulus@hxb-leaf-r4:mgmt:~$ sudo vtysh -c "show evpn vni 10111"
ubuntu@hxb-gpu02:~$ ping -c 3 172.16.11.101
Checkpoint: leaf-r4 shows VNI 10111 with remote VTEPs, and the same-subnet ping works.
Root cause: VLAN 111 on leaf-r4 was mapped to the wrong VNI, so its type-3 (inclusive multicast) and type-2 routes were in a VNI no other leaf used. Same-subnet traffic had no remote VTEPs to go to. Routed traffic could still work through the L3 VNI. This is exactly the typo from Chapter 22's opening story, and the reason for the Ansible assert in site-vlans.yml.
F4 solution — missing VRR on leaf-r3¶
ubuntu@hxb-gpu02:~$ ip neigh show 172.16.10.1
cumulus@hxb-leaf-r3:mgmt:~$ nv show interface vlan110
cumulus@hxb-leaf-r1:mgmt:~$ nv show interface vlan110
cumulus@hxb-leaf-r3:mgmt:~$ nv config history
Fix (use the same syntax family you used in Lab 5; the 5.15+ form is shown):
cumulus@hxb-leaf-r3:mgmt:~$ nv set interface vlan110 ipv4 vrr address 172.16.10.1/24
cumulus@hxb-leaf-r3:mgmt:~$ nv set interface vlan110 ipv4 vrr mac-address 00:00:5e:00:01:01
cumulus@hxb-leaf-r3:mgmt:~$ nv set interface vlan110 ipv4 vrr state enabled
cumulus@hxb-leaf-r3:mgmt:~$ nv config diff
cumulus@hxb-leaf-r3:mgmt:~$ nv config apply -y
On releases before 5.15, the equivalent is nv set interface vlan110 ip vrr address 172.16.10.1/24, … ip vrr mac-address 00:00:5e:00:01:01 and … ip vrr state up.
Verify:
Checkpoint: both pings work, and ip neigh shows 172.16.10.1 as REACHABLE with MAC 00:00:5e:00:01:01.
Root cause: the anycast gateway exists on every leaf, so each host's gateway is always its local leaf. Removing it from one leaf breaks routing only for hosts attached to that leaf, while bridging across the fabric still works.
F5 solution — RoCE lossy on leaf-r2¶
cumulus@hxb-leaf-r2:mgmt:~$ nv show qos roce
cumulus@hxb-leaf-r1:mgmt:~$ nv show qos roce
cumulus@hxb-leaf-r2:mgmt:~$ nv config history
Expected output (illustrative): leaf-r2 shows mode lossy and no PFC on switch priority 3. Leaf-r1 shows mode lossless, PFC on priority 3, ECN enabled and trust pcp,dscp.
Fix and verify:
cumulus@hxb-leaf-r2:mgmt:~$ nv set qos roce mode lossless
cumulus@hxb-leaf-r2:mgmt:~$ nv config apply -y
cumulus@hxb-leaf-r2:mgmt:~$ nv show qos roce
For the trust variant, remove the trust override (nv unset qos mapping default-global trust) and confirm that nv show qos roce reports trust pcp,dscp again (check on your release).
Checkpoint: nv show qos roce on leaf-r2 matches leaf-r1 line for line.
Root cause: one leaf's RoCE profile had been changed to lossy (PFC removed from priority 3). On real Spectrum hardware, RoCE traffic through that leaf would lose its lossless guarantee and drop under congestion, which causes retransmissions and slow collectives. VX can't show that, so the audit (comparing configuration against the standard) is the right detection method. Chapter 5: hosts mark DSCP 26 (traffic class 106), switches trust DSCP and map it to switch priority 3 with PFC and ECN.
F6 solution — Ansible pushed the wrong access VLAN¶
-
Find what changed and who changed it:
cumulus@hxb-leaf-r2:mgmt:~$ nv show interface swp2 bridge domain br_default cumulus@hxb-leaf-r2:mgmt:~$ nv config history cumulus@hxb-leaf-r2:mgmt:~$ nv config diff <previous-rev> <latest-rev> ubuntu@oob-mgmt-server:~/helix-ansible$ git log --oneline -3 ubuntu@oob-mgmt-server:~/helix-ansible$ git show HEAD --statExpected output (illustrative): access VLAN 111 on swp2. The history shows an API revision by the automation user.
git logshowsTidy access ports, which addedplaybooks/access-ports.yml. -
Fix the source of truth first, then the switch. Revert the commit and correct the port through Ansible (or NVUE, if your Lab 12 project doesn't manage access ports):
-
Verify the host and check that the automation now agrees with the fabric:
Checkpoint: gpu03 reaches its gateway and its BOREALIS peer. No AURORA ARP is seen on eth2. The check run reports no changes.
Root cause: an unreviewed playbook applied a wrong access VLAN to one port. It broke gpu03 and put a BOREALIS host into the AURORA broadcast domain. Prevention: peer review (pull requests) for playbooks, --check --diff in CI against the Air twin (Lab 6 SDK script), and an assert step that checks access ports against the tenant plan.
F7 solution — standby SM missing, then master failure¶
Stage 1 (the proactive check):
ubuntu@hxa-ibsim:~/ibsim$ ibsim-run sminfo
ubuntu@hxa-ibsim:~/ibsim$ cat base-sm-ports.txt
ubuntu@hxa-ibsim:~/ibsim$ ibsim-run sminfo <standby LID>
ubuntu@hxa-ibsim:~/ibsim$ ps -ef | grep [o]pensm
Expected output (illustrative):
The query to the standby's LID fails or doesn't show a standby state, and ps shows one OpenSM (priority 14).
Fix: start the standby exactly as Lab 8 does, with priority 12 on its own port GUID. Lab 8's command has this shape (use yours):
ubuntu@hxa-ibsim:~/ibsim$ ibsim-run opensm -g <standby port GUID> -p 12 -P ~/ibsim/partitions.conf -B
ubuntu@hxa-ibsim:~/ibsim$ ibsim-run sminfo <standby LID>
Expected output (illustrative): … priority 12 state 2 SMINFO_STANDBY.
Stage 2 (after F7b, if HA was missed): ibsim-run sminfo times out. Start the standby first, which becomes master because it is the only SM. Then prove the failover path works:
ubuntu@hxa-ibsim:~/ibsim$ ibsim-run sminfo
ubuntu@hxa-ibsim:~/ibsim$ ibsim-run opensm -g <master port GUID> -p 14 -P ~/ibsim/partitions.conf -B
ubuntu@hxa-ibsim:~/ibsim$ ibsim-run sminfo
Checkpoint: one SMINFO_MASTER (priority 14) and one SMINFO_STANDBY (priority 12). Optional: stop the master and confirm the standby takes over within a few sweeps, then restore it.
Root cause: HA had silently degraded to a single SM. Nobody noticed until the master failed. With no master, existing routes kept forwarding but new or flapping ports couldn't become Active. Prevention: alert on SM state (UFM health, Chapter 15), and check the standby explicitly, not only the master.
Field note
Watch priorities when you restart SMs. If the standby is restarted with a higher priority than the master (15, say), it takes over mastership. That is legal but not what the design says, and it makes the next failover go the wrong way. OpenSM's election rule is the highest priority, then the lowest GUID. OpenSM's -p accepts 0–15.
F8 solution — missing partition member¶
ubuntu@hxa-ibsim:~/ibsim$ ibsim-run iblinkinfo | grep <node name>
ubuntu@hxa-ibsim:~/ibsim$ ibsim-run smpquery pkeys <LID of affected node>
ubuntu@hxa-ibsim:~/ibsim$ ibsim-run smpquery pkeys <LID of a healthy PRODUCTION node>
ubuntu@hxa-ibsim:~/ibsim$ diff base-partitions.conf partitions.conf
Expected output (illustrative): the healthy node lists 0xffff (or 0x7fff) and 0x8002. The affected node lists only the default partition. The diff shows the node's GUID missing from the PRODUCTION=0x0002 line.
Fix: add the GUID back (restore the line from the baseline copy, or edit it), then make OpenSM re-read the file and check the PKey table:
ubuntu@hxa-ibsim:~/ibsim$ cp base-partitions.conf partitions.conf
ubuntu@hxa-ibsim:~/ibsim$ pkill -HUP -f '[o]pensm.*-p 14'
ubuntu@hxa-ibsim:~/ibsim$ ibsim-run smpquery pkeys <LID of affected node>
Checkpoint: 0x8002 appears in the node's PKey table. If it doesn't after a minute, restart the master OpenSM the Lab 8 way.
Root cause: the node's port GUID had been removed from the PRODUCTION partition, so the SM didn't program 0x8002 into its PKey table. Packets carrying PKey 0x8002 to or from it are dropped by PKey enforcement, while the port looks perfectly healthy. Remember the encoding: 0x8002 is partition 0x0002 with the full-membership bit set.
F9 solution — rising symbol errors on an uplink¶
Expected output (illustrative):
Run it again two minutes later and the count has risen (4200). Confirm that the link is still up and find its far end:
ubuntu@hxa-ibsim:~/ibsim$ ibsim-run iblinkinfo --switches-only -l | grep -E '"hxa-su1-leaf-r1".*\b33\b'
ubuntu@hxa-ibsim:~/ibsim$ ibsim-run perfquery <leaf LID> 33
Diagnosis: a steadily rising SymbolErrorCounter on a link that is still up means a marginal physical link: a dirty or damaged cable or optic, or a bad port. Symbol errors are corrected or retried, so there's no outage yet, but throughput and latency suffer, and the link may go down or retrain at a lower width. In production you would check the physical layer with mlxlink (BER, eye), look at both ends, and replace the cable or optic in a maintenance window. UFM would flag it through its health checks.
Fix in the simulator (the equivalent of "replace the cable and clear the counters"):
ubuntu@hxa-ibsim:~/ibsim$ echo 'PerformanceSet "hxa-su1-leaf-r1"[33] PortCounters.SymbolErrorCounter=0' > ~/ibsim/simcmd
ubuntu@hxa-ibsim:~/ibsim$ ibsim-run ibqueryerrors
The injection loop runs on oob-mgmt-server, so the instructor should also stop it there (kill %1 in the shell that ran it, or wait until it finishes after eight minutes). Clearing counters with perfquery -R <LID> 33 also works on real fabrics (check how your ibsim build handles it).
Checkpoint: ibqueryerrors shows no errors on port 33, and a second run a few minutes later is still clean.
F10 solution — failed spine uplink on SU2 leaf-r2¶
ubuntu@hxa-ibsim:~/ibsim$ ibsim-run iblinkinfo --switches-only -l > now-links.txt
ubuntu@hxa-ibsim:~/ibsim$ diff base-links.txt now-links.txt
ubuntu@hxa-ibsim:~/ibsim$ grep -E '"hxa-su2-leaf-r2".*\b34\b' base-links.txt
Expected output (illustrative): the diff shows hxa-su2-leaf-r2 port 34 changed from 4X … LinkUp with a spine peer to Down with no peer. The baseline line names the peer, for example "hxa-spine02" port 2.
Fix: reconnect the link through the ibsim console using the peer from your baseline:
ubuntu@hxa-ibsim:~/ibsim$ echo 'Link "hxa-su2-leaf-r2"[34] "hxa-spine02"[2]' > ~/ibsim/simcmd
ubuntu@hxa-ibsim:~/ibsim$ sleep 15; ibsim-run iblinkinfo --switches-only -l | grep -E '"hxa-su2-leaf-r2".*\b34\b'
Some ibsim builds also have ReLink "hxa-su2-leaf-r2", which restores the node's disconnected links. Type Help at the ibsim console to see the commands your build supports.
Checkpoint: port 34 is LinkUp / Active at 4X again, and diff base-links.txt <(ibsim-run iblinkinfo --switches-only -l) shows nothing. The SM brings the port to Active on its next sweep (default sweep interval about 10 seconds).
Root cause: one spine uplink failed. The SM routed around it, so nothing broke, but SU2's bandwidth to the spines dropped by one uplink's share. Only a comparison with a known-good baseline (or UFM link monitoring) shows it quickly. In production, also check LinkDownedCounter on both ends and the physical layer before you declare it fixed.
Verify¶
- You captured Ethernet and InfiniBand baselines and a
pre-capstonefavourite checkpoint. - For each scenario you solved, the notes follow the template, with scope, evidence, root cause, fix, verification and prevention.
- Every fix was proved with the same test that showed the fault.
- After the last scenario, all switches match the baseline:
nv config diff applied startupis empty andshow bgp summary,show evpn vniandnv show qos rocematch the baseline files. - On
hxa-ibsim,sminfoshows a master (priority 14), the standby answers (priority 12),ibqueryerrorsis clean andiblinkinfomatchesbase-links.txt. - Your rubric total is recorded, with the chapters to revisit for any scenario that scored below 7.
Clean-up / save state¶
- If you fixed each fault by hand, run a final comparison with the baseline (the Task 1 loop into a new folder, then
diff -r). - If anything is uncertain, stop the simulation without a checkpoint and start it from
pre-capstone(Lab 6). Restart ibsim and OpenSM afterwards. - Remove the capstone playbook if it's still in the Ansible project, and check
git statusis clean. - Stop the simulation with Stop Simulation and Store Checkpoint. Keep
pre-capstoneas a favourite: you can run the challenge again before the exam.
Exam tie-in¶
- Domain 5 (20%) questions look like these tickets: a symptom, a few lines of output, and "what is the most likely cause?" or "what should you check next?". Practise reading the one line that matters.
- 5.5: know which InfiniBand tool answers which question:
sminfo(who is master),saquery/smpquery(SM-capable ports, PKey tables),iblinkinfo(every link's state, width and speed),ibqueryerrorsandperfquery(error counters),ibnetdiscover(topology for baseline diffs). - 2.1, 2.2, 2.3: control plane up does not mean data plane healthy (F2). Same subnet broken but routing working points to the L2 VNI (F3). Bridging working but routing broken points to the gateway (F4). RoCE settings must be identical on every switch (F5).
- 3.1, 3.2: SM HA means checking the standby too (F7). A port can be Active and still be missing from a partition (F8).
- 6.1, 6.2:
nv config historyshows whether a change came from a person or from automation. Fix the source of truth, then the device (F6).
Review questions¶
- BGP is Established on every link, but large packets between two hosts on different leaves fail and small ones succeed. What do you check first, and why doesn't BGP show the problem?
- A host can ping another host in the same subnet on a different leaf, but not its default gateway. Which configuration is most likely missing, and where?
sminfoshows a healthy master SM. Why is that not enough to confirm SM high availability, and how do you check the standby?ibqueryerrorsshowsSymbolErrorCounter == 4200on a leaf uplink that is still Active, and the value was 1500 five minutes ago. What does this indicate and what is the usual fix?- A leaf's VLAN is wrong, and
nv config historyshows the change came through the API from the automation account. Why should you fix the Git repository before (or together with) the switch?
Answers¶
- Check MTU on both ends of every link in the path (and test with
ping -M do -s 8972in the underlay). BGP doesn't negotiate or test MTU and its packets are small, so sessions stay up while full-size VXLAN frames (inner frame plus 50 bytes) are dropped on the smaller link. - The anycast gateway (VRR) on that SVI on the host's local leaf. In EVPN with a distributed gateway, every leaf answers for the gateway address. Bridging uses the L2 VNI, so it still works without the SVI's VRR address.
sminfowithout arguments shows only the current master. HA needs a standby that is running and inSMINFO_STANDBYstate. Find the SM-capable ports withsaquery -s, then query the standby directly withsminfo <standby LID>, and check that its priority is lower than the master's.- A marginal physical link: a dirty or damaged cable or optic, or a bad port. It hasn't failed yet but is getting worse. Check the physical layer (
mlxlinkon real hardware), reseat or replace the cable or optic in a maintenance window, clear the counters, and confirm they stay at zero. - The repository is the source of truth. If you fix only the switch, the next playbook run pushes the wrong VLAN again. Revert or correct the commit, run the playbook in check and diff mode, and add review and an
assertso it can't happen again.
Go deeper
The matching book chapters cover the exam objectives for this lab in full, with a Q&A pack of about 40 exam-style questions per chapter.