Search Jobs

Search by job, company or skills

Network Engineer GPU Infrastructure

Network Engineer GPU Infrastructure

Nava
  • Posted 4 hours ago
  • Be among the first 10 applicants

Job Description

About Nava

Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the bare metal, on infrastructure built to keep GPUs available and models serving.

About The Team

The Compute team owns Nava's GPU infrastructure end-to-end: from bare-metal provisioning of NVIDIA GPU systems through cluster software, RDMA integration, and production operation. We keep thousands of GPUs healthy and highly utilized so training and inference workloads run fast and reliably.

Responsibilities

  • Provision and configure GPU nodes: OS, drivers, CUDA, GPU Operator, and RDMA connectivity.
  • Write and maintain automation using Python, Ansible, and Terraform.
  • Support cluster deployments, run health checks and benchmarks, and document results.
  • Monitor GPU infrastructure, respond to alerts, and resolve or escalate issues.
  • Maintain runbooks and operational documentation.
  • Collaborate with cross-functional engineering teams to understand business and technical requirements and deliver software solutions.

Required Qualifications

  • Strong hands-on experience with Kubernetes administration and architecture.
  • Bare Metal GPU cluster deployment and management experience.
  • Kubernetes networking, storage, and security.
  • Infrastructure automation (Ansible, Terraform, etc.).
  • Strong Linux systems expertise and working knowledge of the NVIDIA GPU stack (CUDA, drivers, GPU Operator) and RDMA (RoCEv2/InfiniBand).
  • 2–4 years in systems/infrastructure engineering with strong Linux fundamentals and a learning-oriented mindset.

Preferred Qualifications

  • NVIDIA GPU ecosystem exposure.
  • AI/ML infrastructure experience.
  • Monitoring tools such as Prometheus, Grafana, Dynatrace, Datadog, or Zabbix.
  • Hybrid cloud and datacenter infrastructure experience.
  • NCCL and NVLink tuning, Slurm, and distributed training/inference workloads.

Technology environment

NVIDIA GPU systems (DGX/HGX-class, Blackwell), NVLink, CUDA, GPU Operator; RoCEv2 / InfiniBand RDMA fabrics; Kubernetes (networking, storage, security), Slurm; Ansible, Terraform, Python; monitoring with Prometheus, Grafana, Dynatrace, Datadog, or Zabbix; Linux at scale.

Skills: nvidia,kubernetes,cluster,rdma,metal,software,cuda,ansible,infrastructure,linux

More Info

Job Type:
Industry:
Employment Type:

Key Skills

About Company

Similar Jobs

5-8 yrs
Bengaluru, India
Skills:
PrometheusGrafanaDatadogZabbixCudaTerraformLinuxAnsibleDynatracePythonKubernetesNVIDIA GPU systemsRDMAinfinibandGitOpsNCCLSlurm