Search by job, company or skills

Inference Systems Engineer

Early Applicant
  • Posted a month ago
  • Be among the first 10 applicants

Job Description

Role & Responsibilities

  • Design and deploy low-latency inference pipelines for LLMs, diffusion models, and vision transformers across CPU, GPU, and NPU architectures.
  • Optimize model serving stacks using Triton Inference Server, vLLM, or TGI—tuning batch sizes, quantization, and memory layout for peak performance.
  • Containerize and orchestrate inference services via Docker and Kubernetes, ensuring high availability and auto-scaling under fluctuating workloads.
  • Implement model monitoring, health checks, and A/B testing frameworks to validate performance and drift in production.
  • Collaborate with ML Engineers to convert trained models into production-ready formats (ONNX, TensorRT, GGUF) with minimal accuracy loss.
  • Build observability dashboards (Prometheus/Grafana) and alerting logic to detect and mitigate inference bottlenecks in real time.

Skills & Qualifications

Must-Have

  • Python
  • Docker
  • Kubernetes
  • Triton Inference Server
  • ONNX Runtime
  • PyTorch
  • TensorRT
  • Prometheus
  • Grafana
  • CI/CD (GitHub Actions, GitLab CI)

Preferred

  • Experience with vLLM or TGI
  • Knowledge of Model Quantization (AWQ, GPTQ, GGUF)
  • Familiarity with NVIDIA Triton model ensemble pipelines

Skills: cuda,architecture,building,foundation,ml,management,models,distributed systems,infrastructure,inference

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151274187

Beware of Scammers

We don’t charge money for job offers