06 Oct
|
Westbury Partners
|
Sydney
06 Oct
Westbury Partners
Sydney
Engineer and optimize large-scale HPC infrastructure across compute, storage, networking, and Linux environments, developing software tooling and distributed systems that enable reliable, high-performance quantitative research.
What You'll Do:
- Design and maintain high-performance compute and storage infrastructure.
- Optimize systems, networks, and storage for maximum performance.
- Build monitoring and fault-detection solutions across infrastructure.
- Develop tooling for software deployment and upgrades at scale.
- Collaborate with researchers to optimize HPC workloads.
- Support global infrastructure projects and production environments.
Your responsibilities will include:
- Build and maintain distributed HPC systems and parallel filesystems.
- Develop software and automation using languages such as Go, Python, or C.
- Profile, debug, and optimize application and infrastructure performance.
- Implement configuration management and scalable deployment solutions.
- Conduct root cause analysis across complex system failures.
- Participate in coordinated maintenance and operational support.
- Collaborate with technology teams, researchers, and external vendors.
- Develop comprehensive systems and user documentation.
Why Join Us:
- Work with sophisticated, globally distributed HPC infrastructure.
- Solve challenging performance and reliability problems at scale.
- Collaborate directly with quantitative researchers and technology specialists.
- Work across compute, networking, storage, and software engineering.
- Contribute to global infrastructure projects with significant technical scope.
- Gain exposure to highly customized production computing environments.
About You:
- 5+ years of professional HPC experience.
- 5+ years of Linux systems administration experience.
- Strong programming or scripting skills in Go, Python, C, or similar.
- Experience with complex, distributed, interdependent systems.
- Robust debugging, profiling, and performance optimization capabilities.
- Experience with configuration management tools such as Ansible, Puppet, or SaltStack.
- Familiarity with parallel filesystems and batch scheduling systems.
- Strong analytical mindset and commitment to root cause analysis.
- Reliable, hands‑on, and comfortable supporting production environments.
#J-18808-Ljbffr
📌 Senior HPC Infrastructure & Systems Engineer (Sydney)
🏢 Westbury Partners
📍 Sydney