Search by job, company or skills

Infrastructure Engineer

8-10 Years
  • Posted a day ago
  • Be among the first 10 applicants

Job Description

Key Responsibilities:

• Deploy, manage, and optimize AI/ML and LLM inference workloads across GPU clusters, HPC infrastructure, and cloud environments.

• Build and maintain scalable AI platform infrastructure using Kubernetes, containers, and enterprise orchestration platforms.

• Administer and optimize Linux servers including system configuration, patching, security hardening, performance tuning, and troubleshooting.

• Manage physical infrastructure including servers, storage, networking, and bare metal environments within enterprise data centers.

• Implement and maintain CI/CD and automation workflows for platform and infrastructure deployments.

• Optimize infrastructure performance, GPU utilization, resource allocation, and distributed workloads to meet operational requirements. • Benchmark and evaluate AI workloads for scalability, latency, throughput, and resource efficiency.

• Collaborate with infrastructure, SRE, and platform engineering teams to provision compute resources and maintain enterprise-scale AI environments.

• Implement monitoring, logging, observability, and alerting solutions for platform reliability and operational visibility.

• Apply security patches, upgrades, compliance controls, and operational best practices for Linux and Kubernetes environments.

• Troubleshoot issues across hardware, networking, operating systems, Kubernetes clusters, and AI/ML workloads.

• Support enterprise operations through efficient incident, change, and ticket management processes.

• Automate infrastructure operations using scripting and infrastructure automation tools.

Required Qualifications:

• 8+ years of experience in Linux systems administration, cloud-native infrastructure, HPC environments, or platform engineering.

• At least 4 years of experience supporting AI/ML workloads or large-scale distributed compute environments in production.

• Comfortable leveraging AI-assisted tools for collaborative development, code generation, refactoring, and productivity enhancement.

• Strong hands-on expertise with Linux administration (RHEL, Ubuntu, or similar).

• Experience with Kubernetes administration, container orchestration, and cloud native infrastructure platforms.

• Strong understanding of GPU infrastructure, distributed computing, and HPC systems.

• Hands-on experience with bare-metal infrastructure, servers, storage systems, and enterprise networking.

• Strong understanding of networking fundamentals including TCP/IP, DNS, load balancing, and firewalls.

• Experience with scripting and infrastructure automation with Bash, Python etc.

• Experience with CI/CD, DevOps, or infrastructure deployment workflows.

• Experience with monitoring, observability, and logging platforms.

• Strong troubleshooting and performance optimization skills across Linux, Kubernetes, networking, and infrastructure stacks.

• Excellent problem-solving, communication, and collaboration skills.

• Ability to work effectively in fast-paced, mission-critical production environments.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152911679

Similar Jobs

Pune, India

Skills:

shell scriptingPrometheusKubernetesLinux AdministrationGrafanaDockerTerraformCloud BuildGitHub ActionsCloud KMSOpenTelemetry

Beware of Scammers

We don’t charge money for job offers