- Posted 17 days ago
- Be among the first 10 applicants
Job Description
What you'll own
Evaluation design
Build rigorous benchmarks for reasoning, coding, and agentic capabilities across C/C++, Rust, embedded systems, hardware specifications, and engineering tasks
Verifiable ground truth
Develop deterministic evaluators using compilers, emulators, simulators, static analysis, formal verification, and hardware-in-the-loop systems
Evals at scale
Build distributed infrastructure to run evaluations reliably against models and live training checkpoints
Model behavior analysis
Diagnose regressions and anomalous results, separating model failures from issues in prompts, data, evaluators, or infrastructure
Evaluation methodology
Design robust metrics, difficulty curricula, contamination-resistant tasks, and experiments around prompting, sampling, and scaffolding
Closing the training loop
Turn evaluation failures into targeted datasets, reward signals, and new training objectives in collaboration with RL researchers
Researcher-facing tooling
Create dashboards, experiment tracking, and evaluation libraries that make model performance easy to understand and trust
Required experience
- Strong Python programming and software engineering fundamentals
- Experience building benchmarks, evaluation systems, automated graders, research infrastructure, or similar testbeds
- Strong understanding of LLMs and modern evaluation methodologies
- Strong analytical and experimental thinking — you care deeply about whether a metric actually measures the intended capability
- Experience debugging complex systems and investigating unexpected experimental results
- Strong written and verbal communication
Strong-to-have
- Experience with LLM coding/reasoning evaluations or agent benchmarks
- C/C++/Rust and systems programming experience
- Experience with QEMU, Renode, Verilator, Spike, or embedded/RTOS environments Familiarity with formal verification, static analysis, fuzzing, differential testing, or ASTbased program transformation
- Experience with evaluation frameworks such as SWE-bench, EvalPlus, HumanEval, lmevaluation-harness, or BigCode
- Background in statistics, experimental design, observability, or large-scale ML infrastructure
What makes this role special
- Evaluations directly drive training: Your benchmarks and verifiers become part of the SFT/RL feedback loop
- Verifiable intelligence: Evaluate models through executable and formally checkable ground truth — not just LLM judges
- Frontier research: Work on measuring reasoning, coding, and agentic capabilities in domains where correctness genuinely matters
- High ownership: Define evaluation methodology and own projects end-to-end
- Research impact: Opportunity to publish and open-source benchmark methodologies and tooling
More Info
Key Skills
large-scale ML infrastructure
Renode
embedded RTOS environments
Verilator
observability
differential testing
LLMs
SWE-bench
compilers
HumanEval
lmevaluation-harness
distributed infrastructure
evaluation frameworks
EvalPlus
hardware-in-the-loop systems
BigCode
