NVIDIA Quantum-2 Fabric Planning with QM9700 InfiniBand Switches

NVIDIA Quantum-2 Fabric Planning with QM9700 InfiniBand Switches

Plan a QM9700 fabric from the number of endpoint ports, rail count, oversubscription target, and failure domains before selecting switch quantities. A Quantum-2 switch provides substantial radix and bandwidth, but a port count on a data sheet does not determine the topology. The same 64-port switch can serve as a leaf, a spine, a compact single-switch fabric, or part of a larger multi-tier design.

A QM9700 fabric planning exercise should account for the NVIDIA Quantum-2 switch's 64 non-blocking NDR 400Gb/s ports, or up to 128 NDR200 200Gb/s ports through port splitting, in a 1U chassis. NVIDIA specifies 51.2Tb/s of aggregate bidirectional throughput and more than 66.5 billion packets per second. Those capabilities make the switch suitable for dense AI and HPC fabrics, but they also make cabling, power, airflow, management, and topology errors expensive to correct after installation.

Translate compute nodes into network endpoints

Start with endpoint ports, not server count. A server may use one network port, two ports for a dual-rail design, or multiple adapters aligned to GPU, storage, or service domains. For each node type, record the adapter model, number of active ports, required link speed, protocol, and expected communication pattern.

A simple first calculation is:

endpoint ports = nodes x active fabric ports per node

Add non-compute endpoints such as storage systems, login or service nodes, gateways, and management appliances only when they belong on the high-performance fabric. Keep ordinary management Ethernet out of the NDR port count unless the architecture has a specific reason to combine it.

Choose one rail or two rails deliberately

A single-rail design minimizes adapters, switch ports, cables, and operational state. A dual-rail design can provide path diversity and additional aggregate bandwidth when the application stack is designed to use both fabrics. It also doubles the number of endpoint links and introduces a second topology that must remain consistent.

For dual rail, avoid creating two nominal rails that share the same critical power feed, cable tray, or upstream switch failure domain. Document rail A and rail B from adapter port through leaf and spine to power and management. Then validate that NCCL, MPI, UCX, or the selected communication stack uses the intended interfaces.

Use port allocation to expose oversubscription

In a two-tier leaf-spine fabric, each leaf divides its 64 ports between endpoints and spine uplinks. A balanced 32-downlink and 32-uplink allocation provides equal raw downlink and uplink bandwidth when every link operates at the same rate. It supports 32 endpoint ports per leaf at a nominal 1:1 ratio.

The useful formula is:

oversubscription = total leaf downlink bandwidth / total leaf uplink bandwidth

A 48-downlink and 16-uplink allocation produces a 3:1 ratio at equal link speeds. That may be acceptable for workloads with limited cross-leaf traffic, but synchronized AI collectives can expose the reduction quickly. Use application traces or an approved communication model before selecting an oversubscription target.

Leaf allocationEndpoint ports per leafSpine uplinks per leafNominal ratio at equal rates
32 down / 32 up32321:1
40 down / 24 up40241.67:1
48 down / 16 up48163:1

These examples are planning arithmetic, not a universal topology recommendation. Port splitting, mixed link rates, rail design, growth reserve, and routing can change the result.

Size the spine layer from leaf uplinks

Once the leaf allocation is fixed, count the total leaf-to-spine links and distribute them so that a single spine or link failure has a known effect. A fully balanced design typically connects every leaf to the same spine set. Keep enough free spine ports for approved growth, replacement strategy, and any topology-specific links.

Do not use average traffic to justify uneven spine connectivity. Collective traffic can involve every leaf at once. If phased growth is required, define which switch and cable additions preserve the routing pattern at each phase rather than adding links opportunistically.

QM9700 and QM9790 differ in management model

NVIDIA identifies QM9700 and QM9701 as internally managed switches with an onboard subnet manager and MLNX-OS. The product data sheet states that this integrated management supports straightforward bring-up for fabrics up to 2,000 nodes. QM9790 is externally managed and can use NVIDIA UFM for provisioning, monitoring, troubleshooting, and maintenance.

The choice should follow the operating model. An internally managed switch can simplify a smaller or self-contained deployment. External management may be preferred when the fabric team requires centralized UFM capabilities or separates switch hardware from the management plane. Confirm the exact model suffix and airflow direction; do not substitute QM9700 and QM9790 solely because both provide the same nominal port rate.

Examples in the site catalog include the internally managed QM9700-NS2F and the externally managed QM9790-NS2F.

Map 64 logical ports to 32 physical OSFP cages

The QM9700 series exposes 64 400G ports through 32 twin-port OSFP connectors. Cabling documentation must therefore distinguish the front-panel cage from the two logical ports it carries. A label that records only "OSFP 7" is incomplete; it also needs the branch or logical-port identity and the far-end device.

For direct endpoint connections, the switch-side twin-port assembly may split toward two adapter ports. For spine links, a twin-port cable or optical arrangement may connect corresponding switch cages. Use a port map that shows lane assignment, cable part number, length, far-end cage, logical port, rail, and intended link rate.

Plan the interconnect medium with the topology

NVIDIA lists passive and active copper, active fiber, optical modules, and splitter options for Quantum-2. Short in-rack paths may favor qualified DAC or active copper assemblies. Longer or structured routes may need AOCs or pluggable optics and fiber.

Measure routed distance, not rack-center distance. Include vertical managers, overhead trays, service loops, and the bend radius of twin-port assemblies. Verify that cable bulk does not restrict fan trays, power supplies, or adjacent cages. The Mellanox cable catalog can be reviewed after the port map is complete.

Power, airflow, and service access are topology inputs

Switch placement affects cable reach and failure domains, but it must also respect airflow direction and power redundancy. Match front-to-rear or rear-to-front airflow to the rack design and verify the full ordering suffix. Document separate power feeds and make sure a service technician can remove a fan, power supply, or cable without disturbing unrelated fabric links.

High-density OSFP cabling requires disciplined strain relief and labeling. A topology that fits in a spreadsheet but blocks service access is not deployable.

Firmware and software are part of the bill of materials

Record the approved switch software, adapter firmware, OFED or inbox driver, UFM version where used, and communication libraries. Compatibility must be verified across the complete release set. Avoid updating one layer during commissioning without retesting link stability, routing, congestion behavior, and collective performance.

Commission the fabric in stages

  1. Verify every cable and module against the port map before applying workload traffic.
  2. Confirm link rate, width, firmware, logical-port identity, and error counters.
  3. Validate subnet management, routing, partitions, and management-plane resilience.
  4. Run point-to-point bandwidth and latency tests to find local link problems.
  5. Run multi-node and collective tests that exercise every leaf and spine.
  6. Introduce a planned link or switch failure and measure path recovery and workload effect.
  7. Save topology, counters, optical readings, and performance results as the operating baseline.

Questions that arise during QM9700 planning

Does one QM9700 support 64 servers?

It can provide 64 logical 400G ports, but the number of servers depends on ports per server and whether any ports are reserved for uplinks, storage, gateways, or future expansion. A single-switch fabric also creates a different failure domain from a leaf-spine design.

Can 200G endpoints be connected to a 400G Quantum-2 fabric?

The QM9700 family supports up to 128 NDR200 ports through port splitting, and NVIDIA documents backward connectivity options to 200G and 100G infrastructure. The specific breakout cable, adapter capability, firmware, and topology must be qualified.

Should a small cluster use QM9700 or QM9790?

Base the decision on management architecture, not cluster size alone. QM9700 provides internal management; QM9790 is externally managed. Consider existing UFM operations, high-availability requirements, staff skills, and the planned growth path.

Planning summary: count endpoint ports, define rails, select an oversubscription target, and allocate leaf and spine ports before ordering switches. Then bind every logical port to a qualified adapter, cable, management model, power feed, and validation step. That process turns a QM9700 data sheet into an operable Quantum-2 fabric.