09 Aug
|
World Wide Technology
|
Sydney
09 Aug
World Wide Technology
Sydney
About the Role The Field Solutions Engineer is the onshore hands-on execution engine that closes the gap between the offshore engineering squad and the physical reality of an Australian and New Zealand data centre floor. While the Domain Architects (Compute, Network, Storage) design the “Gold Standard” from the regional hub and the offshore HPC Engineer squad executes remote configuration and automation from India, you are the hands that cannot be replaced by a remote session: pulling a faulty transceiver, walking a rack elevation against the LLD, re-seating a cable, or supervising a burn-in test in person.
As a System Integrator, we do not simply manage a static cloud; we design and deliver bespoke, high-scale AI factories for the world’s leading enterprises. In this role, you sit inside the AI Infrastructure team and work across NVIDIA SuperPOD, BasePOD, and Cisco AI Factory deployments as a generalist across the Compute-Network-Storage triad, rather than as a single-domain specialist, and you are the primary point of RMA/DOA diagnosis and remediation on the ground.
You operate with a 100% focus on Delivery, executing across Low-Level Designs (LLDs) assigned by whichever Domain Architect owns the active engagement, and providing Layer 1 QA support and OOB (out-of-band) standup ahead of AI Factory commissioning.
Key Responsibilities
1. Physical Infrastructure & Layer 1 Execution
- Rack & Stack Execution:
- Physically install and verify DGX/HGX/MGX nodes, switches, and PDUs against the current rack elevation and LLD.
- Confirm floor-loading, bolting, and levelling before energisation; escalate any structural discrepancy to the Domain Architect – AI Facilities.
- Structured Cabling & P2P Validation:
- Execute the point-to-point (P2P) cabling schedule; confirm transceiver type and MPO cable size against the code on the box, not the colour, before patching.
- Clean and inspect optical connectors on every patch; validate seating and troubleshoot link-down, miswire, and link-flap faults by elimination (reseat, swap to a known-good port, clean, replace).
- OOB Standup:
- Bring up and validate the out-of-band (BMC/IPMI) management network ahead of in-band and compute-fabric activation, keeping it physically segregated per design.
- Confirm node power state and basic health via BMC before handing off to HPC configuration.
2. HPC Compute, Network & Storage Configuration
- Bare-Metal & Fabric Provisioning:
- Execute NVIDIA Base Command Manager (BCM) provisioning workflows and Ansible playbooks supplied by the Domain Architects to bring compute nodes, switches, and storage clients into service.
- Configure host-side networking (IPoIB, Netplan) and mount high-performance storage clients (VAST, WEKA, Lustre) to the current LLD.
- Firmware & Driver Lifecycle:
- Execute SBIOS, BMC, GPU VBIOS, and NVSwitch firmware upgrades per the NVIDIA firmware recipe across compute, network, and storage tiers.
- Apply OS hardening, kernel patching, and driver updates during scheduled maintenance windows.
- Performance Validation:
- Execute and log HPL, NCCL-tests, ib_write_bw/ib_send_bw, and IOR/FIO benchmark suites; compare results against the Gold Standard and flag deviations to the relevant Domain Architect.
3. RMA/DOA Diagnosis & Operations Support
- Fault Isolation:
- Triage Dead-on-Arrival hardware, Xid errors, flapping links, and stale storage mounts, isolating whether the fault sits in compute, network, or storage before escalating.
- Distinguish a “Node Issue” from a “Network Issue” or a “Storage Issue” using nvidia-smi/dmesg, ibstat/ibdiagnet, and iostat/iotop.
- RMA Coordination:
- Raise and track Return Merchandise Authorisation cases with OEM and vendor support; coordinate physical swap-out with minimal schedule impact.
- L2 Ticket Resolution:
- Handle on-site escalations passed from L1 support; close the loop with the offshore HPC Engineer squad on root cause and remediation.
Technical Competencies
Essential Skills
- Data Centre Background: Prior Data Centre Technician, Field Engineer, or rack-and-stack experience within a System Integrator, OEM, or colocation environment.
- Fault Diagnosis:
- Comfortable triaging GPU faults (nvidia-smi, dmesg),
link faults (ibstat, ibdiagnet, ethtool), and storage mount issues (iostat, iotop) to isolate a fault domain before escalating.
- Physical Infrastructure (Layer 0/1):
- Rack-and-stack procedures, structured cabling (OS2/OM4/DAC/AOC), MPO and OSFP transceiver handling, and cable pathway standards.
- Site Acceptance Test (SAT) support and As-Built documentation capture.
Desirable Experience
- Linux & Automation:
- Solid RHEL/Ubuntu administration; ability to execute and troubleshoot Ansible playbooks.
- Git workflow familiarity (pulling code, branching, committing configuration changes).
- HPC & AI Infrastructure:
- Working proficiency with NVIDIA Base Command Manager (BCM) for bare-metal provisioning.
- Familiarity with DGX/HGX/MGX hardware architecture and standard benchmark suites (HPL, NCCL-tests, IOR/FIO).
- Network Affinity: Hands-on exposure to InfiniBand/RoCEv2 cabling and switch-side transceiver handling.
- Storage Integration: Parallel filesystem client mounting experience (VAST, WEKA, Lustre).
- Container Platforms: Kubernetes/Red Hat OpenShift node-level troubleshooting (NVIDIA GPU Operator, Network Operator).
Certifications
Highly Desirable:
- NVIDIA-Certified Associate: AI Infrastructure and Operations (NCA-AIIO)
- Red Hat Certified System Administrator (RHCSA)
- CDCP (Certified Data Centre Professional)
Success Metrics (KPIs)
- RMA/DOA Turnaround: Time from fault identification to RMA lodgement and physical swap-out completion.
- Layer 1 Accuracy: Zero unresolved miswires or “ghost links” handed to the validation/stress-test phase; cabling matches the current P2P schedule at handover.
- Benchmark Pass Rate: Percentage of nodes and fabric segments passing HPL/NCCL/IOR-FIO validation on the first attempt.
- Ticket Efficiency: Consistently meeting SLAs for L2 escalations requiring on-site hands.
Role overview: It is a hybrid HPC Engineer + Data Centre Technician (DCT).Candidate with good Data Centre Technician (DCT) experience & some knowledge or background with HPC is required for this role
Travel Requirement: Regular travel to data centre sites across Australia and New Zealand; extended in office presence during active builds is a firm requirement, not a preference.
📌 Field Solution Engineer or Data Centre Technician (Sydney)
🏢 World Wide Technology
📍 Sydney