Job Summary :
We are seeking a GCP DevOps Engineer with 4–5 years of experience, specializing in AI/ML workloads and GPU-based infrastructure. The ideal candidate should have strong hands-on expertise in Google Cloud Platform (GCP), with the ability to design, deploy, optimize, and troubleshoot AI-driven architectures, including GPU/TPU environments.
This role requires deep technical knowledge in DevOps practices, cloud automation, and high-performance computing (HPC), along with excellent communication skills to collaborate with cross-functional teams and clients
Key Responsibilities
- Design, implement, and manage scalable cloud infrastructure on Google Cloud Platform (GCP).
- Build, deploy, and maintain Kubernetes-based environments using Google Kubernetes Engine (GKE).
- Manage and optimize GPU/TPU-based infrastructure for AI/ML workloads.
- Configure and administer Compute Engine, Cloud Storage, IAM, VPC, and networking services.
- Develop and maintain CI/CD pipelines for application and machine learning model deployments.
- Monitor cloud infrastructure using Cloud Monitoring, Cloud Logging, Prometheus, and Grafana.
- Implement Infrastructure as Code (IaC) and deployment automation wherever applicable.
- Manage cloud governance, organization policies, project hierarchy, security controls, and cost optimization strategies.
- Troubleshoot infrastructure, networking, Kubernetes, and GPU performance issues.
- Optimize cluster performance, workload scheduling, and cloud resource utilization.
- Ensure high availability, security, compliance, backup, and disaster recovery for cloud environments.
- Collaborate with AI/ML Engineers and Data Scientists to deploy, scale, and monitor production machine learning models.
- Create technical documentation, operational runbooks, and architecture diagrams.
- Provide technical support during production incidents and client engagements.
Required Skills
- Strong hands-on experience with Google Cloud Platform (GCP), including cloud governance, organization policies, IAM, project hierarchy, and cost optimization.
- Expertise in Google Kubernetes Engine (GKE) for deploying, managing, and scaling containerized applications in production environments.
- Proficiency in Docker, Kubernetes, and Helm for containerization, orchestration, and application deployment automation.
- Experience managing GPU/TPU-based infrastructure, including workload scheduling, resource allocation, performance tuning, and cost optimization for AI/ML applications.
- Strong understanding of Compute Engine, Cloud Storage, BigQuery, Vertex AI/AI Platform, and other core GCP services.
- Hands-on experience building and maintaining CI/CD pipelines for cloud-native applications and machine learning model deployments.
- Knowledge of Cloud Monitoring, Cloud Logging, Prometheus, and Grafana to monitor infrastructure health, troubleshoot issues, and improve system reliability.
- Strong Linux administration skills, including system configuration, troubleshooting, scripting, and performance optimization.
- Experience implementing cloud security best practices, Identity & Access Management (IAM), Security Command Center, and compliance controls.
- Ability to troubleshoot Kubernetes clusters, networking issues, storage performance, and GPU utilization bottlenecks in production environments.
- Experience collaborating with Data Scientists and AI/ML Engineers to deploy, scale, and manage production-ready machine learning workloads.
- Strong understanding of high availability, disaster recovery, infrastructure automation, and cloud cost management in enterprise environments.
Qualifications:
- Bachelor's degree in Computer Science, Information Technology, or a related field
- 4–5 years of experience in DevOps / Cloud Engineering
- Hands-on experience with Google Cloud Platform (GCP)
- Experience supporting AI/ML workloads and GPU/accelerator-based infrastructure
Preferred Certification:
- Google Cloud Professional DevOps Engineer (Preferred)