Frequent InfiniBand disconnects are a symptom, not a diagnosis. In a 10-node cluster, first correlate each incident with the affected host/HCA port and switch port, then check port state, subnet-manager availability, switch link diagnostics, changing error counters, and end-to-end connectivity. The evidence will usually separate an administrative problem from a physical-link, firmware, configuration, power, or thermal fault.
What frequent disconnects can mean
The number of nodes does not identify the fault. A single cable or HCA port can fail intermittently, while a missing subnet manager or shared switch/configuration problem can affect many paths. NVIDIA documents these useful clues:
| Observed evidence | What it can indicate | Qualification |
|---|---|---|
PORT_DOWN |
Disabled switch port or disconnected cable | Mapping documented for the WinOF-2 troubleshooting scenario |
PORT_INITIALIZED |
Possible missing subnet manager | Check the fabric for a running SM before replacing hardware |
PORT_ARMED |
Firmware issue in the documented case | Confirm versions and vendor guidance; it is not a universal diagnosis |
| Rising symbol, recovery, or down counters | Developing link, cable, or switch problem | Compare timestamped readings rather than relying on one stale value |
| Active-looking port but failed traffic | Possible end-to-end connectivity problem | Test with ibping |
See NVIDIA’s InfiniBand Related Troubleshooting for the documented port-state mappings and firmware cases.
1. Build an incident map before changing anything
For every disconnect, record the event time and the complete path:
#1 Best Overall
- The 25Gb dual-port SFP+ network card is based on the Mellanox ConnectX-5 Ex controller, which provide the highest performing and most flexible interconnect solution.
- Technical Support:PXE、 RDMA、UEFI、SR-IOV、1588 PTP、Jumbo Frames(9.5KB)
- Windows 10/11、Windows Server 2016/2019/2022、Deepin 15.11/20/20.6/20.9、VMware ESXi 6.5/6.7、Ubuntu 18.04.5/20.04.1、Ubuntu 22.04.2/22.04.3、RHEL/CentOS 7.6/7.9/8.2/8.3、ZTE New Fulcrum 3.2.2/5.0.5、SUSE 12.5/15.4、FreeBSD 13.2、NeoKylin 7.6、OpenKylin 0.7.5、Mikrotik、iKuai route、Galaxy Kylin v10、Zhongke Fangde desktop OS、Zhongke Fangde server OS、Tongxin UOS 20、Emind OS
- install the operating system with its driver CD, or download it from the official website. Includes low-profile and full-height stands to support standard and ultra-thin computers/servers.
- Enjoy 24/7 customer service, 30-day free returns, 1-year free warranty, and lifetime technical support for your peace of mind.
- Hostname, HCA, and HCA port
- Switch name and switch port
- Host-reported state from
ibstatoribstatus - Switch-reported link-down reason or diagnostic code
- Whether other nodes or links failed at the same time
- Recent reboots, reconfiguration, firmware changes, or management actions
A fault confined to one path suggests a local component or port. Simultaneous failures on links sharing a switch, rail, or configuration point toward a common dependency. Keep the timestamps; they are needed to compare status changes and counter growth.
2. Check host port state and the subnet manager
Inspect each affected HCA port
ibstat
ibstatus
Determine whether the port is down, initialized, armed, or active. A down state can correspond to a disabled switch port or disconnected cable; initialized can occur when no subnet manager is available; and armed is associated with a firmware issue in NVIDIA’s documented troubleshooting case. Treat these as search directions, not final proof.
Verify that an SM is running
sudo sminfo
InfiniBand fabrics require a Subnet Manager (SM) to be running. If sminfo fails or reports no SM, ensure that one is active on the fabric, commonly through an opensm service. Service names and deployment choices vary by operating system and installation, so use the service management method appropriate to your environment. The requirement and check are described in NVIDIA’s NCCL networking troubleshooting guide.
Rank #2
- Host Interface: PCI Express 5.0 x16
- Total Number of Ports: 1
- Expansion Slot Type: OSFP
- Media Type Supported: Optical Fiber
- Maximum Data Transfer Rate: 400 Gbit/s
3. Read switch-side link diagnostics
Host status alone cannot show every physical-layer or management event. On NVIDIA NVOS InfiniBand switches, use the command documented for your installed release; the v25.02 manual includes:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
nv show interface <interface-id> link diagnostics
The manual also documents a view across interfaces. Consult the version-matched manual before interpreting syntax or fields.
Diagnostic categories that narrow the search
- Auto-negotiation or link-training failure
- Logical mismatch between link partners
- Bad signal integrity
- Cable compliance-code mismatch
- Unplugged or unsupported cable
- Module thermal shutdown
- Power budget exceeded
- Port closed by a management command
NVOS link-down reasons can additionally include high SER/BER, loss of block lock or alignment, a credit-monitoring watchdog, cable-access problems, a remote fault, a thermal event, or too many link-error recoveries. A code describes the affected port and event; it does not prove that every disconnect in the cluster has the same cause. Refer to NVIDIA’s Link Diagnostic Per Port documentation.
4. Track counters over time
Take at least two readings—ideally before, during, and after an incident—and save the timestamps. NVIDIA’s NCCL guide identifies these fields as useful indicators:
sudo perfquery -x <lid>
SymbolErrorCounterLinkErrorRecoveryCounterLinkDownedCounter
Values that increase on the affected path during failures strengthen the case for a link, cable, HCA, or switch issue. A nonzero value by itself is not a failure-rate statistic and may be historical; interpret it with the port state, switch reason, and event timing.
Recommended Free Tools
5. Test node-to-node connectivity when status looks healthy
If ibstat and ibstatus show normal states but applications still lose connectivity, use ibping to test the fabric path. Start the server on the remote node, then query it from the local node with the remote node’s LID:
Rank #4
- DUAL-PROTOCOL 100G: ConnectX-4 VPI (MCX456A-ECAT) runs EDR InfiniBand 100Gb/s or 100GbE per QSFP28 port with 100G/50G/40G/25G/10G auto-negotiation — one card serves IB and Ethernet fabrics.
- PCIe 3.0 x16, FULL BANDWIDTH: Dual ports sustain line-rate 100Gb/s each for HPC, AI training nodes and high-throughput storage fabrics.
- RDMA WITHOUT CPU COPIES: Native InfiniBand RDMA plus RoCE accelerate MPI, NVMe-oF and distributed storage; hardware offloads cut latency and free CPU cycles.
- HEAVY VIRTUALIZATION: SR-IOV with up to 127 VFs per port (254 per card) plus VXLAN/GENEVE/NVGRE overlay offload for multi-tenant clouds and dense VM hosts.
- DATA CENTER FEATURES: PXE/UEFI boot, NC-SI management, DCB, jumbo frames; Linux (MLNX_OFED), Windows (WinOF) and VMware ESXi support; brackets for any chassis.
# Remote node
sudo ibping -S
# Local node; replace with the remote node's LID
sudo ibping <remote_lid>
A failed test while ports appear active points to an end-to-end or routing/control-plane problem rather than a simple “port is down” condition. Stop the temporary server when testing is complete, according to your operational procedures.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.6. Isolate a physical component only after collecting evidence
When diagnostics and rising counters point to the physical path, inspect cable and transceiver seating, the HCA port, and the corresponding switch port. If maintenance policy allows, change one variable at a time:
- Reseat the suspected cable or module and record the result.
- Replace it with one known-compatible component and repeat the same test.
- If the symptom remains, test the HCA port or switch port using the approved maintenance procedure.
- Confirm whether the fault follows the cable/module, stays with the HCA port, or stays with the switch port.
Do not buy a generic QSFP cable based only on the cluster size. Confirm connector type, InfiniBand generation and speed, cable length, transceiver requirements, and vendor or switch/HCA compatibility. The NVOS diagnostics can identify cable or module indications, but the correct replacement depends on the actual hardware.
How to compare competing explanations
| Comparison | Interpretation |
|---|---|
| One host/port versus several links | One path favors a local component; shared failures favor a common switch, configuration, or control-plane dependency. |
| Simultaneous versus unrelated times | Simultaneous events suggest a shared event; isolated intermittent events suggest individual links or components. |
| Host state versus switch code | Agreement narrows the cause; disagreement warrants checking logs, timing, and release-specific meanings. |
| Counter trend | Increasing symbol, recovery, or down counters are more informative than one old nonzero reading. |
| Controlled component swap | A fault that follows a component identifies a stronger suspect than an unrecorded replacement. |
| Connectivity test | ibping failure with active-looking ports exposes an end-to-end problem. |
What information is still needed for a definitive diagnosis
The root cause cannot be selected from the symptom alone. A useful escalation bundle contains:
- HCA and switch models
- Fabric topology and affected paths
- Operating-system, driver, firmware, and NVOS versions
- Cable and transceiver part numbers
- Timestamped host states and switch link-down codes
sminfooutput- Repeated
perfquery -x <lid>snapshots ibpingresults between an affected pair
This evidence lets an administrator distinguish a missing SM, administrative shutdown, firmware/configuration incompatibility, physical signal problem, thermal or power event, and a failing media or port without guessing.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

