Search Jobs

Search by job, company or skills

Member of Technical Staff

Member of Technical Staff

eBay
  • Posted 16 hours ago
  • Be among the first 10 applicants

Job Description


Job Description

As an LLM Inference Engineer on our AI Platform team, you'll remove the compute-scaling bottleneck for production LLMs. Your job is to make frontier-model inference fast, efficient, reliable, and observable—the last mile from GPUs to APIs that products depend on. This role sits at the intersection of HPC, GPU systems, and MLOps, and requires strong intuition for how model architecture, runtimes, and hardware interact.

What You'll Do

  • Own production inference: Take models from handoff to production-grade serving, including release engineering, capacity planning, cost optimization, and incident response.
  • Tune inference performance: reduce end-to-end latency and increase throughput across real production traffic patterns.
  • Optimize runtimes and servers: Scale inference across heterogeneous GPU fleets; optimize stacks such as vLLM, Triton, and related components (e.g., schedulers, KV cache, batching, memory).
  • Benchmark and measure: Build benchmarking suites, metrics, and tooling to quantify latency, throughput, GPU utilization, memory, and cost.
  • Reliability and observability: Improve monitoring, tracing, and alerting; participate in incident response and postmortems to harden systems.
  • Apply and ship new optimizations: Evaluate research and implement pragmatic inference optimizations (e.g., quantization, paging, kernel/runtimes improvements).
  • Partner with cross-functional teams: Work with data science and product teams to translate business requirements into performance and availability SLOs.

What We're Looking For

  • 8+ years of strong development experience
  • Experience deploying and operating LLM inference services in production.
  • Strong production coding skills in Python plus Go or Rust (systems-level implementation and debugging).
  • Experience with ML frameworks and runtimes: PyTorch, vLLM, SGLang (and/or TensorRT).
  • Knowledge of GPU architecture and performance (profiling, memory bandwidth/latency tradeoffs); CUDA/kernel programming is a strong plus.
  • Solid understanding of LLM inference and optimization techniques: continuous batching, KV cache management, quantization, speculative decoding (nice-to-have), etc.
  • 3+ years hands-on experience in performance optimization and systems programming for AI/ML workloads.
  • Demonstrated ability to deliver measurable production improvements (e.g., 2X throughput, lower p95/p99 latency, reduced GPU cost).
  • Proven skill in root-cause analysis: finding bottlenecks across model, runtime, networking, and infrastructure.
  • Demonstrated proficiency in applying autonomous AI coding agents to speed up software delivery pipelines. This includes advanced prompting and careful human-in-the-loop code review to improve development speed and code accuracy.

More Info

Key Skills

Continuous batching

ML frameworks and runtimes

Benchmarking suites

TensorRT

Autonomous AI coding agents

Quantization

SGLang

Software delivery pipelines

GPU architecture and performance profiling

KV cache management

Speculative decoding

Kernel runtimes improvements

Root-cause analysis

CUDA kernel programming

vLLM

About Company

Similar Jobs

10-12 yrs
Bengaluru, India
Skills:
data engineering JavaHibernateSqlSpringNosqlRestful ApisPythonApplication development infrastructure managementMicroservices architectureMessaging queue conceptsAutomated testing frameworks
5-8 yrs
Bengaluru, India
Skills:
UnixMachine LearningSqlNlpPythonKubernetesComputer VisionGenAI ModelsDeep Learning ModelsEKSAWS EcosystemNeural Network ArchitecturesKubeflow
5-10 yrs
Bengaluru, India
Skills:
IpsecUnix Shell ScriptingPythonVirtualizationEspEmbedded development on X86 64 processorsEmbedded multi-core processorsIKEV2EAP-AKAx86 64 processor architectureVPPMulti-processor system developmentIntel DPDK
5-10 yrs
Bengaluru, India
Skills:
PythonRtl DesignSVA assertionssystemverilogLow-power design techniquesFPGA prototypingSoC Methodology
7-9 yrs
Hyderabad, Bengaluru, India
Skills:
.Net CoreC#SamlMicroservicesAzure FunctionsDockerAzure AdministrationAWSOauthPowerShellPython ScriptingAzureKubernetesIdentity ServersOperational support software development and deployment methodologiesOpenID ConnectStorage AccountsAPI RESTEvent GridsAzure APIsLogic AppsServerless ArchitectureKey VaultsService Bus