23 Sep
|
Westbury Partners
|
New South Wales
23 Sep
Westbury Partners
New South Wales
Job Description
Support and optimise large-scale Linux HPC environments, resolving complex infrastructure issues, automating operations, and maintaining reliable compute, storage, and networking systems around the clock.
n
What You'll Do:
n
n
- Join a highly technical infrastructure team responsible for keeping large-scale Linux HPC environments running reliably and efficiently. You'll provide hands‐on operational support across compute, storage, networking, and interconnected infrastructure while responding to challenging and unpredictable production issues.
n
- You'll work closely with researchers, engineers, vendors, and infrastructure teams to investigate problems, implement solutions, automate repetitive tasks, and continuously improve the reliability and performance of the environment.
n
n
Your responsibilities will include:
n
n
- Provide front-line operational support for 24/7 Linux HPC compute, storage, and network infrastructure.
n
- Troubleshoot complex issues across RDMA fabrics, parallel filesystems, batch schedulers, FUSE filesystems, hardware, and software.
n
- Manage incidents and problem reports through their entire lifecycle, escalating when required.
n
- Respond quickly and effectively to infrastructure alerts.
n
- Participate in coordinated maintenance activities, including scheduled evening and weekend maintenance windows.
n
- Support global infrastructure projects across a broad range of technologies.
n
- Develop code and tooling to diagnose, troubleshoot, triage, and resolve difficult operational problems.
n
- Automate frequently performed operational tasks.
n
- Contribute to testing infrastructure and codebases across multiple programming languages.
n
- Develop and maintain performance and fault‐monitoring systems.
n
- Create and improve technical and user documentation.
n
- Work with external technology vendors and support vendor relationships.
n
- Participate in an on‐call rotation and provide operational support as a core responsibility.
n
- Follow cybersecurity and IT policies when managing infrastructure and production systems.
n
n
Why Join Us:
n
n
- This is an prospect to work directly with large-scale, high‐performance computing infrastructure where every day brings new technical challenges.
n
- You'll gain hands‐on exposure to Linux, HPC, distributed storage, high-speed networking, automation, monitoring, hardware, and production operations while working alongside highly technical infrastructure and research teams.
n
- The role is ideal for engineers who genuinely enjoy operational work and want to develop their expertise by solving difficult problems in a fast-moving environment. You'll also have opportunities to contribute to global projects, work with leading technology vendors, and build automation that makes complex infrastructure easier to operate.
n
n
About You:
n
n
- You're a hands‐on Linux engineer who enjoys operational work and thrives when solving unpredictable technical problems.
n
- You have at least two years of professional Linux experience and strong programming or scripting skills in a language such as Python, Go, or C. HPC experience is valuable but not essential — what matters is your ability to learn quickly and tackle unfamiliar technologies.
n
- You're comfortable performing root cause analysis, managing multiple workstreams, and communicating clearly with both technical colleagues and external vendors.
n
- You bring a strong sense of urgency, excellent collaboration skills, reliable availability, and the flexibility to support scheduled maintenance and on‐call responsibilities when required.
n
📌 Linux HPC Operations Engineer (New South Wales)
🏢 Westbury Partners
📍 New South Wales