Domain Architect - Ai Network (New South Wales)

Domain Architect - Ai Network (New South Wales)

27 Aug
|
World Wide Technology
|
New South Wales

27 Aug

World Wide Technology

New South Wales

The
Domain Architect - AI Network
acts as the primary technical authority for the physical and logical lifecycle of
high-performance interconnect fabrics
across diverse client environments, who bridges the gap between architectural design and hands-on execution. You are a "doer" who is as comfortable configuring
Adaptive Routing on leaf and spine switches
in the CLI as you are explaining that configuration to a C-level client.
As a System Integrator, we do not simply manage a static cloud; we design and deliver bespoke, high-scale AI factories for the world's leading enterprises. In this role, you will define the "Gold Standard" for
network
infrastructure, moving beyond single-
switch
management to architect repeatable, scalable, and automated
lossless
fabrics. You will serve as the technical lead for NVIDIA Cloud Provider (NCP) and private enterprise AI cloud deployments, owning the "
Network
" in the critical "Compute-Network-Storage" triad.
In this role, you will operate with a
***** split
between delivering (60) complex AI infrastructure and providing Pre-Sales (40) Subject Matter Expertise (SME). You will lead the physical provisioning of
InfiniBand and Spectrum-X Ethernet fabrics for NVIDIA SuperPOD, NVIDIA BasePOD, and Cisco AI Factory environments
, ensuring our clients receive "Day 2" ready AI factories, while assisting the sales team in defining the scope and cost of future deployments.
Key Responsibilities
1. Delivery & Implementation (60%)
High-Performance Fabric Design:
Architect and deploy
Non-Blocking Fat-Tree (Clos)
topologies using
NVIDIA Quantum-2 (NDR) / Quantum-3 (XDR) InfiniBand switches
and
Spectrum-4 Ethernet switches
.
Implement
Rail-Optimised
network designs to ensure GPU-to-GPU traffic is perfectly aligned with compute node PCIe trees, minimising latency for NCCL collectives.
Configure
Adaptive Routing
,
Congestion Control
, and
Quality of Service (QoS)
to prevent "Head-of-Line Blocking" and ensure lossless data delivery.
In-Network Computing & Offload Strategy:
InfiniBand (Switch-Based):
Enable and tune
SHARP (Scalable Hierarchical Aggregation and Reduction Protocol)
on Quantum switches to offload collective operations (AllReduce,



ReduceScatter) from the GPU to the switch silicon.
Ethernet (Endpoint-Based):
Architect high-performance Ethernet AI fabrics utilising
SmartNIC/DPU offloading
(e.g.,
NVIDIA BlueField-3
) to accelerate collective operations and isolate management traffic from the data path.
Future Stacks:
Evaluate and roadmap emerging
Ultra Ethernet Consortium (UEC)
standards and hardware capabilities to transition In-Network Collectives from proprietary InfiniBand to open Ethernet standards.
Tune
RoCEv2
and adaptive routing algorithms on Ethernet fabrics and SmartNICs to approximate InfiniBand-like lossless behaviour.
Fabric Management & Telemetry:
Deploy
NVIDIA UFM (Unified Fabric Manager)
to manage InfiniBand subnets and
NetQ
for Ethernet telemetry.
Monitor fabric health to identify "Slow Receivers" and link degradation in real-time.
Implement
PTP (Precision Time Protocol)
for nanosecond-level clock synchronisation across the entire cluster.
Layer 1 Precision:
Oversee the physical cabling strategy, defining strict cable schedules (DAC vs. AOC vs. OSFP transceivers) to meet signal integrity requirements over specific distances.
Validate the "Link Budget" to ensure optical loss falls within acceptable limits for 400G/800G links.
Technical Scoping & Sizing:
Calculate the required
Bisectional Bandwidth
for client workloads (e.g., Training requires 1:1 non-blocking, Inference may tolerate 3:1 oversubscription).
Produce accurate
Labour Estimates (LOE)
for complex network integrations.
BoM Validation:
Own the Network Bill of Materials (BoM), validating that every transceiver, cable, and switch is on the
NVIDIA/OEM HCL
.
Manage the complexities of breakout cables (e.g., 800G to *****G) and connector types (OSFP, QSFP112) to prevent on-site installation failures.




Educate clients on the difference between standard Enterprise Ethernet (Lossy, TCP-based) and AI Fabric (Lossless, RDMA-based).
Design integration points between the high-speed "Backend" AI fabric and the client's existing "Frontend" management network (BGP/EVPN handoffs).
Deep expertise in
NVIDIA Quantum InfiniBand
switch architecture and management via
MLNX-OS
and UFM.
Mastery of
Spectrum-X
Ethernet platforms running
NVIDIA Cumulus Linux
or
SONiC (Software for Open Networking in the Cloud)
.
Proficiency in configuring
RoCEv2
(RDMA over Converged Ethernet), including
PFC
(Priority Flow Control) and
ECN
(Explicit Congestion Notification).
Architectural understanding
of
NVIDIA BlueField DPU
and
DOCA
use cases for offloading network functions.
Network Automation (NetDevOps):
Proficiency in
Python
and
Ansible
for switch configuration management.
Experience with
Infrastructure as Code (IaC)
principles to treat network config as software.
Fabric Telemetry:
Hands-on experience with
UFM
,
Prometheus
, and
Grafana
for visualising fabric congestion and "tail latency."
Multi-Vendor Exposure:
Experience with
Arista (EOS)
for high-performance AI Ethernet or
Cisco (Nexus Dashboard)
for AI infrastructure.
Routing Protocols:
Solid understanding of
BGP
,
EVPN
, and
VXLAN
for multi-tenant isolation.
Compute Integration:
Understanding of how the network interacts with the host OS (IPoIB, Netlink) and optimisation of
GPU Direct RDMA (GDR)
and
NCCL communication patterns
.
Scale-Across
: Extending an InfiniBand fabric across multiple data centrehalls (NVIDIA Metro-X).
Certifications:
NVIDIA-Certified Skilled: AI Networking (NCP-AIN)
NVIDIA-Certified Associate: AI Infrastructure and Operations (NCA-AIIO)
Fabric Performance:
Achieving >95% effective bandwidth efficiency on
NCCL-test
benchmarks across the full cluster.
Deployment Accuracy:
Zero "Layer 1" rework required due to incorrect transceiver/cable specification in the BoM.
Stability:
Zero
topology or design-induced link instability
(e.g., cabling mismatches) in production.
#J-*****-Ljbffr

📌 Domain Architect - Ai Network (New South Wales)
🏢 World Wide Technology
📍 New South Wales

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: domain architect - ai network (new south wales) / new south wales

Subscribe to this job alert:

Get the latest job offers by email for: domain architect - ai network (new south wales) / new south wales