09 Oct
|
World Wide Technology
|
Sydney
09 Oct
World Wide Technology
Sydney
The Principal Architect - AI Infrastructure is the design authority for the "Compute-Network-Storage" triad and the governing technical owner of the physical AI factory. You own the reference architecture, the standards against which every cluster is judged, and the technical coherence of the triad as a single system. As a System Integrator, you design and deliver bespoke, high-scale AI factories for the world’s leading enterprises. You are a consultative and governing authority, not an engineer-of-record. You reason fluently across silicon, fabric, and data platform and own the "Gold Standard" reference architectures (NVIDIA SuperPOD, NVIDIA BasePOD, Cisco AI Factory) for NVIDIA Cloud Provider (NCP) and private enterprise AI cloud deployments. On infrastructure-led engagements - the "Turnkey AI Data Center" - you are the Prime, owning the final acceptance criteria and the Architecture Review Board (ARB) approval. This position operates with a 60/40 split between Technical Authority & Delivery Governance (60%) and Pre-Sales, Commercial & Practice Development (40%).
Key responsibilities
- Own and maintain the "Gold Standard" reference architectures for the triad (NVIDIA SuperPOD, NVIDIA BasePOD, Cisco AI Factory), keeping them repeatable, scalable, and commercially defensible across NCP and private enterprise AI cloud builds
- Chair the internal Architecture Review Board (ARB), approving or denying deviations from the reference architecture and owning the record of why each deviation was permitted
- Ensure HLD/LLD coherence across Compute, Network, and Storage so the three layers integrate as one fabric
- Act as the Design Authority on infrastructure-led engagements: defend the architecture before the client’s Design Authority / ARB and own the final technical acceptance criteria as Prime
- Line-manage and technically develop the Compute, Network, and Storage Domain Architects, owning their technical bench health, skills plan, and succession
- Arbitrate cross-domain technical disputes within the triad and across streams, preventing "design-by-committee" stalls
- Validate the consolidated infrastructure Bill of Materials against budget and against the NVIDIA HCL / OEM compatibility matrices before commitment, and own margin on infrastructure delivery
- Define the interfaces between the triad and the adjacent Facilities and Edge Infrastructure domains, reconciling the physical envelope with the logical design
- Receive qualified Horizon 2 and Horizon 3 infrastructure programs from the Enterprise Architect and shape them into buildable programs with an explicit reference architecture, charter, and acceptance criteria
- Own the Partner Selection Matrix for the triad (compute OEM, fabric, and storage vendors) and the technical side of the NVIDIA and OEM alliance relationships
About you
- Demonstrated authority across Compute, Network, and Storage,
with the ability to reason about their interaction as a single system; depth in at least one layer and credible breadth across the other two
- Command of SuperPOD and BasePOD reference designs; NVL72/DGX/HGX/MGX compute; Quantum InfiniBand and Spectrum-X Ethernet fabrics; and high-performance parallel storage (VAST, WEKA, DDN, Pure)
- Working knowledge of NCP and Cisco AI Factory build models
- ARB facilitation, reference-architecture and deviation control, and BoM/HCL governance, holding a design defensible under a client Design Authority and an internal margin review
- LOE/SOW construction and review, infrastructure delivery margin, BoM-versus-budget, and TCO/utilisation reasoning sufficient to underwrite a build commercially
- Line management and technical coaching of senior architects, and governance of multi-squad delivery through a layered architect / senior engineer / engineer model
- Sufficient depth in Linux systems engineering and Infrastructure as Code (Ansible, Python, Terraform) to govern automation quality and review the squads without owning implementation
- Prior tenure as a Domain Architect or lead infrastructure architect within a System Integrator (SI) or MSP, with at least one greenfield AI factory delivered end to end (desirable)
- Working understanding of the "Day 2" stack (Kubernetes / Red Hat OpenShift, Rafay, Run:AI) and the MLOps/LLMOps handover (desirable)
- NVIDIA-Certified Professional: AI Infrastructure (NCP-AII), NVIDIA-Certified Qualified: AI Networking (NCP-AIN), or NVIDIA-Certified Associate: AI Infrastructure and Operations (NCA-AIIO) (preferred)
#J-18808-Ljbffr
📌 Principal Architect – AI Infrastructure (Sydney)
🏢 World Wide Technology
📍 Sydney