29 Aug
|
Sharon AI
|
Sydney
AI Platform Engineer L2 (GPUaaS – AI Neocloud)
? Sydney | Hybrid
About Sharon AI
Sharon AI is building the infrastructure powering the next generation of artificial intelligence.
Operating across AI infrastructure, high-performance compute, cloud platforms and large-scale technology environments, Sharon AI delivers scalable, secure and reliable infrastructure for demanding AI, ML and HPC workloads.
The Role
As an AI Platform Engineer L2, you'll help build, operate and support the platform layer powering Sharon AI's GPU-as-a-Service (GPUaaS) offering. You'll work across Kubernetes, Slurm, container orchestration, MLOps tooling and model serving infrastructure to enable customers to train and run AI/ML workloads reliably and efficiently on Sharon AI's neocloud platform.
Reporting to the Head of Operations, you'll work closely with Network Engineering, Infrastructure and customer-facing teams to implement, automate and troubleshoot the platform services sitting above Sharon AI's underlying GPU and network fabric. This is a hands-on opportunity for a platform, DevOps, MLOps or SRE engineer looking to deepen their expertise in GPU infrastructure and AI-native platform operations.
Key Responsibilities
- Build and operate the AI platform layer, including Kubernetes, Slurm and container orchestration, supporting Sharon AI's GPU infrastructure
- Develop and maintain CI/CD pipelines for model training, fine-tuning and inference workloads
- Implement and support MLOps tooling for experiment tracking, model registry and deployment
- Configure and manage multi-tenant GPU resource scheduling and quota management across customer workloads
- Support model serving infrastructure for training and inference,
ensuring reliability and performance
- Build monitoring, logging and alerting to track platform health, GPU utilisation and workload performance
- Automate platform provisioning and configuration using Infrastructure-as-Code tools such as Terraform and Ansible
- Collaborate with Network Engineering and Infrastructure teams to ensure the platform layer aligns with the underlying InfiniBand/RDMA fabric
- Troubleshoot platform-level issues affecting customer AI/ML workloads
- Participate in on-call rotations and incident response for platform-related issues
- Contribute to platform documentation, runbooks and the internal knowledge base
- Support customer onboarding onto the GPUaaS platform, including workload configuration and troubleshooting
Skills & Experience
- 2–4 years' experience in platform engineering, DevOps, MLOps or SRE, ideally supporting GPU or AI/ML workloads
- Bachelor's degree in Computer Science or a related field
- Hands-on production experience with Kubernetes
- Experience with CI/CD and Infrastructure-as-Code
- Solid understanding of Kubernetes and GPU scheduling frameworks, including Slurm, Kubernetes device plugins and NVIDIA GPU Operator
- Experience with MLOps tooling and ML pipeline orchestration
- Proficiency in scripting and automation using Python and Bash
- Working knowledge of GPU infrastructure and distributed training concepts, including NCCL and data/model parallelism
- Experience with observability tools such as Prometheus and Grafana
- Understanding of Linux systems administration and networking fundamentals
- Strong troubleshooting and problem-solving skills
- Strong communication and collaboration skills, with the ability to work across Infrastructure, Network Engineering and customer-facing teams
- Exposure to GPU-based infrastructure or high-performance computing environments
Experience with Slurm, NVIDIA GPU Operator, NCCL and distributed training frameworks such as PyTorch or TensorFlow is advantageous, as is experience with MLOps platforms including MLflow, Kubeflow or Ray. Kubernetes certifications such as CKA/CKAD, cloud certifications across AWS/GCP/Azure, and exposure to InfiniBand/RDMA networking concepts are also advantageous.
Why Join Sharon AI
- Help build and operate the platform powering a growing GPU-as-a-Service and AI neocloud business
- Work hands-on with GPU infrastructure, Kubernetes, Slurm and AI-native platform technologies
- Develop deeper expertise across AI/HPC infrastructure and distributed workloads
- Work closely with Network Engineering and Infrastructure teams across the underlying GPU and network fabric
- Help enable customers to reliably train, fine-tune and run AI/ML workloads at scale
- Contribute to automation, observability and platform capabilities in a quick-moving environment
- Join a highly technical and ambitious team operating at the forefront of AI infrastructure
Our Values
Integrity | Innovation | Collaboration | Wellbeing | Inclusion
📌 AI Platform Engineer (Sydney)
🏢 Sharon AI
📍 Sydney