Search Jobs

Search by job, company or skills

Senior Research Engineer – Post-Training & Reinforcement Learning

Senior Research Engineer – Post-Training & Reinforcement Learning

Eternity Quests LLP
4-12 Years
45 - 150 LPA
Early Applicant
Quick Apply
  • Posted 16 days ago
  • Be among the first 10 applicants

Job Description

  • Location: Delhi NCR
  • Experiwnce - 4+ years

  • Core Mandate: SFT, preference optimization, reasoning RL, reward modelling, verifier environments.

  • Responsibilities:
  • Drive model alignment and capability enhancement through Supervised Fine-Tuning (SFT).
  • Implement and scale preference optimization techniques (e.g., RLHF, DPO) and advanced reasoning RL.
  • Train robust reward models and engineer complex verifier environments to support iterative capability scaling.

  • Requirements:
  • 4+ years of applied experience in deep learning, specifically focusing on Reinforcement Learning, alignment methodologies, or post-training foundation models.

More Info

Job Type:
Function:
Employment Type:

Key Skills

RLHF

DPO

LLM Post Training

Model Alignment

Reward Modelling

Reasoning Models

About Company

Similar Jobs

4-12 yrs
Bengaluru, Delhi NCR
Skills:
Deep Learning, Pytorch, Machine Learning, RLHF, DPO, Supervised Fine Tuning, SFT, reinforcement learning, Reward Modelling, Model Alignment, ppo, GRPO, Preference Optimization, LLM Post Training, TRL, Reasoning Models, Large Language Models, Llm, Ai