

Search by job, company or skills

This job is no longer accepting applications
Senior DevOps Engineer – AI & HPC Infrastructure
We are looking for an exceptional Senior DevOps Engineer to lead the design, operation, and evolution of a cutting-edge AI and High-Performance Computing (HPC) infrastructure.
This is not a traditional DevOps role. We're looking for someone who has built and operated mission-critical infrastructure at scale, thrives in complex environments, and can take full ownership of platform reliability, automation, and performance. You will work at the intersection of AI, cloud infrastructure, HPC, and MLOps, enabling world-class engineering and machine learning teams.
What You'll Do
What We're Looking For
Required Qualifications
Preferred Qualifications
Ideal Candidate
We're looking for a highly experienced engineer who has worked in demanding production environments and is comfortable taking ownership of critical infrastructure. The ideal candidate combines deep technical expertise with a pragmatic approach to engineering, automation, and operational excellence, and enjoys solving complex infrastructure challenges that support advanced AI workloads.
Job ID: 151775779
Skills:
Unix, Elk, Prometheus, Grafana, Jenkins, Git, DevSecOps, Linux, Bitbucket, Docker, Terraform, Ansible, Openshift, Gitlab, Dynatrace, Splunk, Azure, Kubernetes, Python, AWS, Agile delivery
Skills:
Shell Scripting, Git, Gcp, Cloud, Docker, Terraform, Linux, Ansible, Azure, Kubernetes, Azure DevOps, AWS, monitoring and logging tools
Skills:
Continuous Integration, Continuous Deployment, security best practices in software development and deployment, AWS cloud platforms, AWS container orchestration technologies, Infrastructure as Code using Terraform, DevOps Architecture
Skills:
Unix, Prometheus, Bash, Grafana, Elk Stack, Jenkins, Git, Gcp, Linux, Docker, shell scripting, Terraform, Ansible, Azure, Kubernetes, Python, AWS, CI CD, GitLab CI, GitHub Actions
Skills:
Terraform, Datadog, Python, Kubernetes, AWS, GitOps, Braintrust, LLM systems, Arize, Observability, LangSmith