Senior Engineer – Distributed Training Systems & Performance
Eternity Quests LLP- Posted 3 hours ago
- Be among the first 10 applicants
Job Description
Senior Engineer – Distributed Training Systems & Performance
Squeeze every FLOP out of a very large cluster. Sharding, numerics, checkpointing, MFU.
Location: Delhi NCR or Bengaluru | Experience: 4 – 12 years
About the opportunity
We are retained by a well-funded AI lab building foundation models from the ground up in India — real compute, real pretraining, not fine-tuning on top of someone else's model. The team is small and the work is unusually direct: what you build ships into the model.
The client's identity is confidential at this stage and will be shared on first call.
What you will do
• Optimise core training infrastructure across the JAX/XLA stack.
• Design tensor, pipeline and data sharding strategies across large TPU/GPU clusters.
• Implement FP8 numerics and architect reliable high-speed checkpointing.
• Profile workloads continuously to maximise Model FLOPs Utilisation.
What we are looking for
• 4+ years in distributed systems, HPC, or ML infrastructure.
• Deep expertise in parallel computing and hardware accelerators.
• Low-level stack optimisation experience.
What is on offer
• Compensation is open and benchmarked to the top of the Indian market for this profile, with meaningful equity.
• Delhi NCR is the primary base, but the client is flexible on Bengaluru for the right person.
• Relocation support, including for candidates returning from outside India.
How to apply
Apply here, or write to [Confidential Information] with your CV and links to anything you have published, shipped or open-sourced. We reply to every serious application within 48 hours.
Spotlight
- Joining bonus, Performance bonus, Mobile bill reimbursements, Relocation benefits, Retirement benefits, Health insurance
Bachelor Of Technology (B.Tech/B.E), Masters in Technology (M.Tech/M.E), Master of Science (MS/M.Sc), PhD, Bachelor in Management Studies (B.M.S.), Doctor of Ministry, Master in Landscape Architecture, Doctor of Public Health (DrPH)
More Info
Key Skills
XLA
Model Parallelism
Tensor Parallelism
Pipeline Parallelism
FSDP
NCCL
GPU Optimization
Kernel Optimization
Mixed Precision
FP8
Checkpointing
MFU
ML Infrastructure
