LLM Pre-training & Distributed Engineer (AI Infrastructure) (Australia)

LLM Pre-training & Distributed Engineer (AI Infrastructure) (Australia)

16 Aug
|
Hyphen Connect
|
Australia

16 Aug

Hyphen Connect

Australia

We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure effective and reliable training processes.

Responsibilities:

- Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
- Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
- Automate checkpointing and failure recovery during month-long training runs.

Required Skills:

- Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
- Experience managing SLURM or Kubernetes-based GPU clusters.
- Strong systems engineering background (C++, CUDA, Python).

📌 LLM Pre-training & Distributed Engineer (AI Infrastructure) (Australia)
🏢 Hyphen Connect
📍 Australia

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: llm pre-training & distributed engineer (ai infrastructure) (australia) / australia

Subscribe to this job alert:

Get the latest job offers by email for: llm pre-training & distributed engineer (ai infrastructure) (australia) / australia