Search by job, company or skills

AI Research Engineer (Benchmarking)

AI Research Engineer (Benchmarking)

soch street
3-5 Years
Not Disclosed
Early Applicant
  • Posted 17 days ago
  • Be among the first 10 applicants

Job Description

What you'll own

Evaluation design

Build rigorous benchmarks for reasoning, coding, and agentic capabilities across C/C++, Rust, embedded systems, hardware specifications, and engineering tasks

Verifiable ground truth

Develop deterministic evaluators using compilers, emulators, simulators, static analysis, formal verification, and hardware-in-the-loop systems

Evals at scale

Build distributed infrastructure to run evaluations reliably against models and live training checkpoints

Model behavior analysis

Diagnose regressions and anomalous results, separating model failures from issues in prompts, data, evaluators, or infrastructure

Evaluation methodology

Design robust metrics, difficulty curricula, contamination-resistant tasks, and experiments around prompting, sampling, and scaffolding

Closing the training loop

Turn evaluation failures into targeted datasets, reward signals, and new training objectives in collaboration with RL researchers

Researcher-facing tooling

Create dashboards, experiment tracking, and evaluation libraries that make model performance easy to understand and trust

Required experience

  • Strong Python programming and software engineering fundamentals
  • Experience building benchmarks, evaluation systems, automated graders, research infrastructure, or similar testbeds
  • Strong understanding of LLMs and modern evaluation methodologies
  • Strong analytical and experimental thinking — you care deeply about whether a metric actually measures the intended capability
  • Experience debugging complex systems and investigating unexpected experimental results
  • Strong written and verbal communication

Strong-to-have

  • Experience with LLM coding/reasoning evaluations or agent benchmarks
  • C/C++/Rust and systems programming experience
  • Experience with QEMU, Renode, Verilator, Spike, or embedded/RTOS environments Familiarity with formal verification, static analysis, fuzzing, differential testing, or ASTbased program transformation
  • Experience with evaluation frameworks such as SWE-bench, EvalPlus, HumanEval, lmevaluation-harness, or BigCode
  • Background in statistics, experimental design, observability, or large-scale ML infrastructure

What makes this role special

  • Evaluations directly drive training: Your benchmarks and verifiers become part of the SFT/RL feedback loop
  • Verifiable intelligence: Evaluate models through executable and formally checkable ground truth — not just LLM judges
  • Frontier research: Work on measuring reasoning, coding, and agentic capabilities in domains where correctness genuinely matters
  • High ownership: Define evaluation methodology and own projects end-to-end
  • Research impact: Opportunity to publish and open-source benchmark methodologies and tooling

Key Skills

large-scale ML infrastructure

Renode

embedded RTOS environments

Verilator

observability

differential testing

LLMs

SWE-bench

compilers

HumanEval

lmevaluation-harness

distributed infrastructure

evaluation frameworks

EvalPlus

hardware-in-the-loop systems

BigCode

About Company

Similar Jobs

3-6 yrs
Bengaluru, India
Skills:
Java, Github, Rust, Git, Typescript, Javascript, Docker, Python, Integration Tests, SWE-bench, Go, Terminal-Bench, DeepSWE, unit tests, Linux shell scripts, CI workflows, SWE-Bench Pro, test harnesses