06 Aug
|
World Wide Technology
|
Sydney
06 Aug
World Wide Technology
Sydney
The Domain Architect - AI Network acts as the primary technical authority for the physical and logical lifecycle of high-performance interconnect fabrics across diverse client environments, who bridges the gap between architectural design and hands-on execution. You are a "doer" who is as comfortable configuring Adaptive Routing on leaf and spine switches in the CLI as you are explaining that configuration to a C-level client.
As a System Integrator, we do not simply manage a static cloud; we design and deliver bespoke, high-scale AI factories for the world's leading enterprises. In this role, you will define the "Gold Standard" for network infrastructure, moving beyond single- switch management to architect repeatable, scalable, and automated lossless fabrics. You will serve as the technical lead for NVIDIA Cloud Provider (NCP) and private enterprise AI cloud deployments, owning the "
Network " in the critical "Compute-Network-Storage" triad.
In this role, you will operate with a 60/40 split between delivering (60) complex AI infrastructure and providing Pre-Sales (40) Subject Matter Expertise (SME). You will lead the physical provisioning of InfiniBand and Spectrum-X Ethernet fabrics for NVIDIA SuperPOD, NVIDIA BasePOD, and Cisco AI Factory environments , ensuring our clients receive "Day 2" ready AI factories, while assisting the sales team in defining the scope and cost of future deployments.
Key Responsibilities
1. Delivery & Implementation (60%)
- High-Performance Fabric Design:
- Architect and deploy Non-Blocking Fat-Tree (Clos) topologies using NVIDIA Quantum-2 (NDR) / Quantum-3 (XDR) InfiniBand switches and Spectrum-4 Ethernet switches .
- Implement Rail-Optimised network designs to ensure GPU-to-GPU traffic is perfectly aligned with compute node PCIe trees, minimising latency for NCCL collectives.
- Configure Adaptive Routing , Congestion Control , and Quality of Service (QoS) to prevent "Head-of-Line Blocking" and ensure lossless data delivery.
- In-Network Computing & Offload Strategy:
- InfiniBand (Switch-Based): Enable and tune SHARP (Scalable Hierarchical Aggregation and Reduction Protocol) on Quantum switches to offload collective operations (AllReduce, ReduceScatter) from the GPU to the switch silicon.
- Ethernet (Endpoint-Based):
Architect high-performance Ethernet AI fabrics utilising SmartNIC/DPU offloading (e.g., NVIDIA BlueField-3 ) to accelerate collective operations and isolate management traffic from the data path.
- Future Stacks: Evaluate and roadmap emerging Ultra Ethernet Consortium (UEC) standards and hardware capabilities to transition In-Network Collectives from proprietary InfiniBand to open Ethernet standards.
- Tune RoCEv2 and adaptive routing algorithms on Ethernet fabrics and SmartNICs to approximate InfiniBand-like lossless behaviour.
- Fabric Management & Telemetry:
- Deploy NVIDIA UFM (Unified Fabric Manager) to manage InfiniBand subnets and NetQ for Ethernet telemetry.
- Monitor fabric health to identify "Slow Receivers" and link degradation in real-time.
- Implement PTP (Precision Time Protocol) for nanosecond-level clock synchronisation across the entire cluster.
- Layer 1 Precision:
- Oversee the physical cabling strategy, defining strict cable schedules (DAC vs. AOC vs. OSFP transceivers) to meet signal integrity requirements over specific distances.
- Validate the "Link Budget" to ensure optical loss falls within acceptable limits for 400G/800G links.
2. Pre-Sales SME & Consulting (40%)
- Technical Scoping & Sizing:
- Calculate the required Bisectional Bandwidth for client workloads (e.g., Training requires 1:1 non-blocking, Inference may tolerate 3:1 oversubscription).
- Produce accurate Labour Estimates (LOE) for complex network integrations.
- BoM Validation:
- Own the Network Bill of Materials (BoM), validating that every transceiver, cable, and switch is on the NVIDIA/OEM HCL .
- Manage the complexities of breakout cables (e.g., 800G to 2x400G) and connector types (OSFP, QSFP112) to prevent in office installation failures.
- Client Strategy Workshops:
- Educate clients on the difference between standard Enterprise Ethernet (Lossy, TCP-based)
and AI Fabric (Lossless, RDMA-based).
- Design integration points between the high-speed "Backend" AI fabric and the client's existing "Frontend" management network (BGP/EVPN handoffs).
Technical Competencies
Essential Skills
- NVIDIA Networking (Mellanox):
- Deep expertise in NVIDIA Quantum InfiniBand switch architecture and management via MLNX-OS and UFM.
- Mastery of Spectrum-X Ethernet platforms running NVIDIA Cumulus Linux or SONiC (Software for Open Networking in the Cloud) .
- Proficiency in configuring RoCEv2 (RDMA over Converged Ethernet), including PFC (Priority Flow Control) and ECN (Explicit Congestion Notification).
- Architectural understanding of NVIDIA BlueField DPU and DOCA use cases for offloading network functions.
- Network Automation (NetDevOps):
- Proficiency in Python and Ansible for switch configuration management.
- Experience with Infrastructure as Code (IaC) principles to treat network config as software.
- Fabric Telemetry:
- Hands-on experience with UFM , Prometheus , and Grafana for visualising fabric congestion and "tail latency."
Desirable Experience
- Multi-Vendor Exposure: Experience with Arista (EOS) for high-performance AI Ethernet or Cisco (Nexus Dashboard) for AI infrastructure.
- Routing Protocols: Solid understanding of BGP , EVPN , and VXLAN for multi-tenant isolation.
- Compute Integration: Understanding of how the network interacts with the host OS (IPoIB, Netlink) and optimisation of GPU Direct RDMA (GDR) and NCCL communication patterns .
- Scale-Across : Extending an InfiniBand fabric across multiple data centre halls ( NVIDIA Metro-X ).
- Certifications:
- NVIDIA-Certified Professional: AI Networking (NCP-AIN)
- NVIDIA-Certified Associate: AI Infrastructure and Operations (NCA-AIIO)
- Arista ACE (Cloud Engineer).
- Cisco CCNP/CCIE Data Center.
Success Metrics (KPIs)
1. Fabric Performance: Achieving >95% effective bandwidth efficiency on NCCL-test benchmarks across the full cluster.
2. Deployment Accuracy: Zero "Layer 1" rework required due to incorrect transceiver/cable specification in the BoM.
3. Stability: Zero topology or design-induced link instability (e.g., cabling mismatches) in production.
📌 AI Network Domain Architect (Sydney)
🏢 World Wide Technology
📍 Sydney