09 Aug
|
P&C Partners
|
Queensland
09 Aug
P&C Partners
Queensland
We're working with an Australian technology company that builds and operates its own GPU infrastructure, delivering dedicated GPU clusters to customers across Australia. As the business continues to grow, they're looking for a GPU Infrastructure Lead to establish and lead their customer delivery capability.
This is a founding technical role responsible for taking GPU hardware from powered racks through to fully operational customer environments. You'll own the delivery process from cluster build and configuration through to customer acceptance, ongoing technical ownership and future delivery standards.
What You'll Work On
Build, provision and configure production GPU clusters
Deploy Linux, firmware, drivers, CUDA and supporting software across customer environments
Stand up Slurm or Kubernetes clusters based on customer requirements
Configure container runtimes, shared storage and customer software environments
Validate multi-node performance across Ethernet, RoCE or InfiniBand networks, including NCCL performance
Produce benchmark reports, runbooks and customer handover documentation
Provide ongoing technical ownership following customer acceptance, including upgrades, monitoring and technical escalations
Support pre-sales engagements through workload sizing, cluster specification and proof-of-concept activities
We're working with an Australian technology company that builds and operates its own GPU infrastructure, delivering dedicated GPU clusters to customers across Australia. As the business continues to grow,
they're looking for a GPU Infrastructure Lead to establish and lead their customer delivery capability.
This is a founding technical role responsible for taking GPU hardware from powered racks through to fully operational customer environments. You'll own the delivery process from cluster build and configuration through to customer acceptance, ongoing technical ownership and future delivery standards.
What You'll Work On
Build, provision and configure production GPU clusters
Deploy Linux, firmware, drivers, CUDA and supporting software across customer environments
Stand up Slurm or Kubernetes clusters based on customer requirements
Configure container runtimes, shared storage and customer software environments
Validate multi-node performance across Ethernet, RoCE or InfiniBand networks, including NCCL performance
Produce benchmark reports, runbooks and customer handover documentation
Provide ongoing technical ownership following customer acceptance, including upgrades, monitoring and technical escalations
Support pre-sales engagements through workload sizing, cluster specification and proof-of-concept activities
What We're Looking For
Demonstrated experience building and delivering GPU clusters for customers or end users
Strong Linux systems engineering experience
Production experience with Slurm and/or Kubernetes
Strong understanding of the CUDA ecosystem, including drivers, NCCL and container tooling
Experience tuning high-performance networking for distributed GPU workloads
Infrastructure as Code experience using tools such as Ansible, Terraform or similar
Comfortable working directly with customers throughout deployment and acceptance
Highly Regarded
Experience within a GPU cloud, neocloud or hyperscaler GPU environment
InfiniBand operations
High-performance storage platforms such as Lustre, Weka or VAST
Remote or edge infrastructure experience
Eligibility for Australian Government security clearance
Why You'll Love It
Brisbane preferred | Australia-wide considered
Up to $300k + Super + Equity
Founding technical leadership opportunity with genuine ownership of the delivery capability
Prospect to define the tooling, delivery processes and technical standards for future deployments
Work on large-scale GPU infrastructure supporting AI and high-performance computing workloads
Join an Australian business building and operating its own GPU infrastructure from the ground up
If you love the deep technical work of making GPU clusters hum, and want to be the person who defines how it's done rather than inheriting someone else's playbook, this could be the role for you.
Additional information
Up to $300k + Super + Equity
Founding technical leadership role
Slurm, Kubernetes, CUDA & GPU clusters
#J-*****-Ljbffr
📌 Gpu Infrastructure Lead - Customer Delivery (Queensland)
🏢 P&C Partners
📍 Queensland