Search by job, company or skills

Post-Training Optimization Engineer (LLM Inference & Efficiency)

Post-Training Optimization Engineer (LLM Inference & Efficiency)

Nava
Fresher
Not Disclosed
Early Applicant
  • Posted 2 months ago
  • Be among the first 20 applicants

Job Description

Role & Responsibilities

  • Optimize LLM inference pipelines for latency, throughput, and memory efficiency across GPU/TPU hardware—using quantization, pruning, kernel fusion, and runtime scheduling.
  • Implement and benchmark model compression techniques (INT4/INT8/FP8, LoRA, Sparse attention) to reduce model size while preserving performance.
  • Integrate and tune LLM serving frameworks (vLLM, TensorRT-LLM, HuggingFace TGI, ONNX Runtime) for high-throughput, low-latency inference in cloud and on-premise environments.
  • Collaborate with ML researchers and infrastructure engineers to profile bottlenecks and implement hardware-aware optimizations using CUDA, Triton, or custom kernels.
  • Design and automate model benchmarking suites to track performance regression, memory footprint, and cost-per-token across model versions and hardware targets.
  • Document optimization playbooks and contribute to internal tooling for repeatable, scalable model deployment workflows.

Skills & Qualifications

Must-Have

  • PyTorch
  • TensorRT
  • vLLM
  • Quantization (INT4/INT8/FP8)
  • CUDA
  • ONNX Runtime
  • Triton Inference Server
  • LLM Inference Optimization

Preferred

  • Experience with Mixture-of-Experts (MoE) models
  • Familiarity with HuggingFace Transformers and TGI
  • Knowledge of NPU/ASIC inference backends (e.g., Qualcomm, Groq, Cerebras)

Skills: cuda,ml,learning,compression,optimization,distillation,decoding,research,machine learning,training

More Info

Job Type:
Industry:
Employment Type:

Key Skills

Triton Inference Server

Quantization INT4 INT8 FP8

ONNX Runtime

LLM Inference Optimization

TensorRT

vLLM

About Company