Llm Pre Training & Distributed Engineer Ai Infrastructure Launceston

Llm Pre Training & Distributed Engineer Ai Infrastructure Launceston

08 Oct
|
Hyphen Connect
|
Launceston

08 Oct

Hyphen Connect

Launceston

Job Description

We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure effective and reliable training processes.

Responsibilities:
Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, Deep Speed, or Megatron-LM.
Optimize networking (Infini Band/RDMA) and memory management to prevent out-of-memory errors.
Automate checkpointing and failure recovery during month-long training runs.

Required Skills:
Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
Experience managing SLURM or Kubernetes-based GPU clusters.
Robust systems engineering background (C++, CUDA, Python).

📌 Llm Pre Training & Distributed Engineer Ai Infrastructure Launceston
🏢 Hyphen Connect
📍 Launceston

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: llm pre training & distributed engineer ai infrastructure launceston / launceston

Subscribe to this job alert:

Get the latest job offers by email for: llm pre training & distributed engineer ai infrastructure launceston / launceston