Llm Pre-Training & Distributed Engineer (Ai Infrastructure) (Central Coast)

Llm Pre-Training & Distributed Engineer (Ai Infrastructure) (Central Coast)

07 Oct
|
Hyphen Connect
|
Central Coast

07 Oct

Hyphen Connect

Central Coast

Job Description
We are seeking a highly skilled LLM Pre-training & Distributed Systems Engineer. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure. The ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure effective and reliable training processes.
n
n
Responsibilities:
n
n
n
n
Orchestrate distributed training runs across 1,000+ GPUs using PyTorch, DeepSpeed, or Megatron-LM.
n
n
Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors.
n
n
Automate checkpointing and failure recovery during month-long training runs.
n
n
n
n
Required Skills:
n
n
n
n
Deep expertise in 3D parallelism (Data, Tensor, Pipeline).
n
n
Experience managing SLURM or Kubernetes-based GPU clusters.
n
n
Strong systems engineering background (C++, CUDA, Python).
n
n
n

📌 Llm Pre-Training & Distributed Engineer (Ai Infrastructure) (Central Coast)
🏢 Hyphen Connect
📍 Central Coast

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: llm pre-training & distributed engineer (ai infrastructure) (central coast) / central coast

Subscribe to this job alert:

Get the latest job offers by email for: llm pre-training & distributed engineer (ai infrastructure) (central coast) / central coast