02 Oct
|
World Wide Technology
|
Melbourne
02 Oct
World Wide Technology
Melbourne
About the role The Principal Architect – AI Infrastructure is the design authority for the "Compute-Network-Storage" triad and the governing technical owner of the physical AI factory. Where the Domain Architects own the engineering of each layer and the offshore squads own the build, you own the reference architecture, the standards against which every cluster is judged, and the technical coherence of the triad as a single system. You are the technical judiciary for the infrastructure stream: you adjudicate conflicting engineering approaches, approve deviations from the "Gold Standard", and delegate the L3 fix to the responsible Domain Architect.
As a System Integrator, you design and deliver bespoke, high-scale AI factories for the world's leading enterprises. You are a consultative and governing authority, not an engineer-of-record. You reason fluently across silicon, fabric, and data platform – as comfortable interrogating a rail-optimised topology or a parallel filesystem sizing as defending a Bill of Materials against budget – but you draw hands-on depth from the Domain Architect pool and the offshore Senior HPC Engineers rather than owning the keyboard yourself.
You own the "Gold Standard" reference architectures (NVIDIA SuperPOD, NVIDIA BasePOD, Cisco AI Factory) for NVIDIA Cloud Provider (NCP) and private enterprise AI cloud deployments, and the deviation gates that protect them.
On infrastructure-led engagements – the "Turnkey AI Data Center" – you are the Prime. You own the final acceptance criteria and the Architecture Review Board (ARB) approval, defend the solution before the client's Design Authority, and direct the Platforms, Solutions, and Facilities Principals as internal stakeholders and suppliers to the build. You are the originating technical owner who converts the sanctioned Horizon 2 and Horizon 3 programs handed down by the Enterprise Architect into a buildable, governed reference design.
This position operates with a 60/40 split between Technical Authority & Delivery Governance (60%) and Pre-Sales, Commercial & Practice Development (40%). The 60% is governance, design authority,
and delivery oversight exercised through the Domain Architects and squads – not personal hands-on implementation.
Key responsibilities
- Own and maintain the "Gold Standard" reference architectures for the triad (NVIDIA SuperPOD, NVIDIA BasePOD, Cisco AI Factory), keeping them repeatable, scalable, and commercially defensible across NCP and private enterprise AI cloud builds
- Chair the internal Architecture Review Board (ARB), approving or denying deviations from the reference architecture and owning the record of why each deviation was permitted
- Ensure HLD/LLD coherence across Compute, Network, and Storage so the three layers integrate as one fabric: NUMA/PCIe affinity aligned to rail-optimised topology, storage clients aligned to the RDMA fabric, and validation targets (NCCL/HPL, IOR/FIO) consistent across the triad
- Act as the Design Authority on infrastructure-led engagements: defend the architecture before the client's Design Authority / ARB and own the final technical acceptance criteria as Prime
- Approve reference-architecture deviations affecting the global fabric, security posture, or supportability, and hold logical-design authority for changes the Field Solutions Engineer and Domain Architects escalate from site
- Coach the Domain Architects on "the Story": translating engineering decisions into a narrative a client CTO and a procurement function will both sign
- Line-manage and technically develop the Compute, Network, and Storage Domain Architects, owning their technical bench health, skills plan, and succession
- Arbitrate cross-domain technical disputes within the triad and across streams, preventing "design-by-committee" stalls
- Govern the layered delivery model: Domain Architects direct the offshore Senior HPC Engineers,
who direct the HPC Engineers, holding the architects accountable for HLD/LLD quality and first-pass validation success
- Validate the consolidated infrastructure Bill of Materials against budget and against the NVIDIA HCL / OEM compatibility matrices before commitment, and own margin on infrastructure delivery
About you
- Demonstrated authority across Compute, Network, and Storage, with the ability to reason about their interaction as a single system; depth in at least one layer and credible breadth across the other two
- Command of SuperPOD and BasePOD reference designs; NVL72/DGX/HGX/MGX compute; Quantum InfiniBand and Spectrum-X Ethernet fabrics; and high-performance parallel storage (VAST, WEKA, DDN, Pure)
- Working knowledge of NCP and Cisco AI Factory build models
- ARB facilitation, reference-architecture and deviation control, and BoM/HCL governance, holding a design defensible under a client Design Authority and an internal margin review at once
- LOE/SOW construction and review, infrastructure delivery margin, BoM-versus-budget, and TCO/utilisation reasoning sufficient to underwrite a build commercially
- Line management and technical coaching of senior architects, and governance of multi-squad delivery through a layered architect / senior engineer / engineer model
- Sufficient depth in Linux systems engineering and Infrastructure as Code (Ansible, Python, Terraform) to govern automation quality and review the squads without owning implementation
- Prior tenure as a Domain Architect or lead infrastructure architect within a System Integrator (SI) or MSP, with at least one greenfield AI factory delivered end to end (desirable)
- Working understanding of the "Day 2" stack (Kubernetes / Red Hat OpenShift, Rafay, Run:AI) and the MLOps/LLMOps handover (desirable)
- Familiarity with structured service management practices (e.g. incident, change, and problem management)
About us World Wide Technology is a System Integrator designing and delivering bespoke, high-scale AI factories for the world's leading enterprises.
📌 Principal Architect AI Infrastructure (Melbourne)
🏢 World Wide Technology
📍 Melbourne