Domain Architect Ai Storage (New South Wales)

Domain Architect Ai Storage (New South Wales)

03 Aug
|
World Wide Technology
|
New South Wales

03 Aug

World Wide Technology

New South Wales

TheDomain Architect - AI Storageacts as the primary technical authority for the physical and logical lifecycle ofhigh-performance data platformsacross diverse client environments, who bridges the gap between architectural design and hands‐on execution.
You are a "doer" who is as comfortable configuringan NVMe-over-Fabrics connectionin the CLI as you are explaining that configuration to a C‐level client.
As a System Integrator, we do not simply manage a static cloud; we design and deliver bespoke, high‐scale AI factories for the world's leading enterprises.
In this role, you will define the "Gold Standard" forstorageinfrastructure, moving beyond single-arraymanagement to architect repeatable, scalable, and automateddata fabrics.
You will serve as the technical lead for NVIDIA Cloud Provider (NCP) and private enterprise AI cloud deployments, owning the "Storage" in the critical "Compute‐Network‐Storage" triad.
In this role, you will operate with a***** splitbetween delivering (60) complex AI infrastructure and providing Pre‐Sales (40) Subject Matter Expertise (SME).
You will lead the physical provisioning ofhigh‐throughput storage clustersforNVIDIA SuperPOD, NVIDIA BasePOD,andCisco AI Factory environments, ensuring our clients receive "Day 2" ready AI factories, while assisting the sales team in defining the scope and cost of future deployments.
Key Responsibilities
1. Delivery & Implementation (60%)
Parallel File System (PFS) Deployment:
Lead the installation and configuration of high-performance storage clusters using technologies such asWEKA,DDN (Lustre),VAST Data, orPure Storage.
Optimise storage client configurations on compute nodes, managing kernel modules and mounting parameters to ensure stability at scale.
ImplementGPUDirect Storage (GDS)technologies to bypass the CPU and enable direct data paths between NVMe drives and GPU memory.
Container Storage Integration:




Deploy and configureContainer Storage Interface (CSI)drivers for Kubernetes/Red Hat OpenShift, ensuring persistent storage is dynamically provisioned for AI workloads.
Design storage classes that differentiate between "Scratch" (High Performance) and "Home/Project" (General Purpose) tiers.
Performance Tuning & I/O Profiling:
Execute synthetic benchmark suites (IOR,FIO,mdtest) to validate throughput (GB/s) and metadata performance (IOPS) against agreed SLAs.
Troubleshoot "straggler" issues where slow I/O starves GPU utilisation, analysing client‐side logs and fabric counters.
Data Lifecycle Management:
Implement automated data tiering strategies to move datasets between Hot (NVMe), Warm (QLC Flash), and Cold (Object/S3/Tape) tiers based on access frequency.
Technical Scoping & Sizing:
Lead the architectural sizing for storage opportunities.
Move the conversation beyond "How many Petabytes?" to "How many Terabytes per Second per GPU?".
Calculate required performance for specific workloads (e.g., massive small‐file ingestion for Computer Vision vs. large‐file streaming for LLMs) to create accurate estimates.
BoM Validation:
Own the Storage Bill of Materials (BoM), ensuring the correct ratio of Storage Servers to Compute Nodes/GPUs.
Validate interoperability between Storage Controllers, Host Channel Adapters (HCAs), and Transceivers against theNVIDIA Hardware Compatibility List (HCL)and vendor compatibility matrices.
Advise clients on the transition from legacy Enterprise NAS (NFSv3) to modern AI‐Native Storage protocols (NFS over RDMA, NVMe‐oF).




Design namespace architectures that support multi‐tenancy, data sovereignty requirements, and potentially hybrid architectures for workload bursting capability.
High-Performance Storage Ecosystem:
Expert-level knowledge ofParallel File Systems(WEKA, Lustre, BeeGFS, GPFS).
Deep understanding ofObject Storage(S3 protocols) for model checkpointing and archiving.
Mastery of storage protocols:NVMe-over-Fabrics (NVMe‐oF),NFS over RDMA, andNVIDIA GPUDirect Storage (GDS).
Deep understanding of the Linux I/O stack, including Block Device drivers, file system tuning (xfs, ext4), and client-side caching mechanisms.
Infrastructure as Code (IaC):
Proficiency in one ofPython,Ansible, orTerraformfor automating storage cluster deployment and client configuration management.
Industry Background:Experience working within aSystem Integrator (SI), Storage Vendor (e.g., NetApp, Dell, Pure), or MSP setting.
Platform Integration:Hands‐on experience integrating storage with bare metal, andKubernetes/Red Hat OpenShift.
Network Affinity:Understanding ofInfiniBandandRoCEv2fabrics from a storage perspective (Congestion Control, Quality of Service).
LLM Inference & Caching Stack:
Understanding of technologies such asNVIDIA Dynamo,vLLM,SGLang, for distributed inference serving and KV cache management.
Understanding of Distributed Pre-fill concepts and their storage I/O requirements.
Understanding LMCache for KV cache offloading and sharing.
I/O Saturation:Achieving >90% of the theoretical wire-speed throughput on compute clients during validation testing (IOR/FIO).
Deployment Velocity:Successful "Day 1" mount availability across all compute nodes using automated playbooks.
Storage BoM Accuracy: Usable capacity vs. Raw capacity calculations are accurate to within 5% of client requirements (i.e. accounting for product functional overhead).
#J-*****-Ljbffr

📌 Domain Architect Ai Storage (New South Wales)
🏢 World Wide Technology
📍 New South Wales

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: domain architect ai storage (new south wales) / new south wales