Senior Research Engineer – Post-Training & Reinforcement Learning
Eternity Quests LLP- Posted 5 hours ago
- Be among the first 10 applicants
Job Description
Senior Research Engineer – Post-Training & Reinforcement Learning
Take a base model and make it reason, follow instructions, and stay aligned — at frontier scale.
Location: Delhi NCR or Bengaluru | Experience: 4 – 12 years
About the opportunity
We are retained by a well-funded AI lab building foundation models from the ground up in India — real compute, real pretraining, not fine-tuning on top of someone else's model. The team is small and the work is unusually direct: what you build ships into the model.
The client's identity is confidential at this stage and will be shared on first call.
What you will do
• Drive model alignment and capability gains through supervised fine-tuning.
• Implement and scale preference optimisation (RLHF, DPO and successors) and reasoning-focused RL.
• Train robust reward models.
• Engineer verifier environments that support iterative capability scaling.
What we are looking for
• 4+ years of applied deep learning, concentrated in reinforcement learning, alignment, or post-training of foundation models.
• Working command of the modern post-training stack, not just API-level fine-tuning.
• Evidence of depth: publications, open-source contributions, or models you have shipped.
What is on offer
• Compensation is open and benchmarked to the top of the Indian market for this profile, with meaningful equity.
• Delhi NCR is the primary base, but the client is flexible on Bengaluru for the right person.
• Relocation support, including for candidates returning from outside India.
How to apply
Apply here, or write to [Confidential Information] with your CV and links to anything you have published, shipped or open-sourced. We reply to every serious application within 48 hours.
Spotlight
- Rewards & recognition, Joining bonus, Performance bonus, Mobile bill reimbursements, Relocation benefits, Retirement benefits, Health insurance
Master of Design (M.Des.), Bachelor of Computer Science (BCS), Master of Computer Science (MCS), Executive Master OF Business Administration (E.M.B.A), Bachelor Of Technology (B.Tech/B.E), Doctor of Ministry, Master in Landscape Architecture
More Info
Key Skills
RLHF
DPO
Supervised Fine Tuning
SFT
Reward Modelling
Model Alignment
GRPO
Preference Optimization
LLM Post Training
TRL
Reasoning Models
Large Language Models
