We are looking for an experienced Azure Devops & Kubernetes Engineer to manage, troubleshoot, and optimize our container orchestration platform. The ideal candidate will handle day-to-day cluster operations, deep-dive troubleshooting, and platform reliability tasks while working closely with L1 support and L3/Platform Architects to ensure highly available, secure, and performant Kubernetes environments.
Experience: 5-9 Years
Location: Chennai
Work Timing: 1 pm - 11 pm & 10 pm - 6 am
Key Responsibilities
Cluster Operations & Support
- Manage and support Kubernetes clusters (on Azure AKS) including day-to-day operations, incident and problem management (L2 escalations from L1), and RCA within defined SLAs.
- Perform cluster upgrades (control plane and worker nodes) including version planning, compatibility checks, rollback strategy, and minimizing downtime during upgrades.
- Manage Control Plane components — kube-apiserver, kube-scheduler, kube-controller-manager, cloud-controller-manager — including health checks, HA configuration, and troubleshooting component failures.
- Monitor, backup, restore, and troubleshoot etcd (cluster state store), including snapshotting, defragmentation, quorum issues, and disaster recovery.
- Perform hands-on operations using kubectl — get/describe/logs/exec/apply/patch/rollout/drain/cordon/top/debug — for day-to-day administration and troubleshooting.
Troubleshooting
- Troubleshoot Pod-level issues: CrashLoopBackOff, ImagePullBackOff, OOMKilled, pending pods, failed scheduling, readiness/liveness probe failures, and node pressure conditions.
- Diagnose and resolve Networking issues: DNS resolution (CoreDNS), service discovery, pod-to-pod communication, service ClusterIP/NodePort/LoadBalancer issues, and network policy conflicts.
- Debug CNI (Container Network Interface) issues — Calico, Cilium, Flannel, Azure CNI, or AWS VPC CNI — including IP exhaustion, plugin failures, and overlay/underlay network issues.
- Troubleshoot Storage issues: PV/PVC binding failures, StorageClass misconfigurations, CSI driver issues, volume mount errors, and data persistence problems across StatefulSets.
- Resolve Ingress issues: Ingress controller (NGINX, Traefik, HAProxy, Azure App Gateway) misconfigurations, TLS/certificate issues, routing rules, and load balancing problems.
- Investigate node-level issues (disk pressure, memory pressure, kubelet health) and cluster autoscaler behavior.
Security & Access
- Implement and manage RBAC (Roles, ClusterRoles, RoleBindings, ClusterRoleBindings) and Service Accounts for secure, least-privilege access.
- Manage Network Policies for pod-to-pod traffic control and micro-segmentation.
- Support Pod Security Standards/Admission, Secrets management, and integration with external secret stores (Vault, Azure Key Vault, AWS Secrets Manager).
- Ensure compliance with security best practices, image scanning, and Kubernetes CIS benchmark adherence.
Platform & Automation
- Contribute to Platform Design decisions — cluster architecture, multi-tenancy models, namespace strategy, resource quotas/limits, and capacity planning.
- Design and maintain Helm charts, Kustomize overlays, and GitOps workflows (ArgoCD/Flux) for application and platform deployments.
- Work with Infrastructure as Code (Terraform, Bicep, CloudFormation) for provisioning clusters and supporting cloud resources.
- Implement and maintain CI/CD pipelines integrating with Kubernetes deployments (Azure DevOps, GitHub Actions, Jenkins, GitLab CI).
- Configure and maintain observability stack — Prometheus, Grafana, Loki/ELK, Azure Monitor/Container Insights — for cluster and workload monitoring.
- Manage service mesh (Istio/Linkerd) where applicable for traffic management, mTLS, and observability.
- Document runbooks, SOPs, architecture diagrams, and knowledge base articles for recurring issues.
- Participate in on-call rotation and support during planned maintenance/outages.
Required Skills & Qualifications
- Bachelor's degree in computer science, IT, or related field (or equivalent experience).
- 4-9 years of hands-on experience in Kubernetes administration/engineering.
- Strong command over kubectl and Kubernetes object model (Deployments, StatefulSets, DaemonSets, Jobs/CronJobs, ConfigMaps, Secrets).
- Deep troubleshooting experience across Pods, Networking, Storage, and Ingress.
- Solid understanding of Control Plane architecture and etcd operations/backup-restore.
Preferred / Good to Have
- Certified Kubernetes Administrator (CKA) or Certified Kubernetes Security Specialist (CKS).
- Experience with Platform Design for multi-cluster/multi-tenant environments.
- Exposure to service mesh (Istio, Linkerd) and API gateways.
- Experience with managed Kubernetes services (AKS, EKS, GKE) alongside on-prem/bare-metal clusters.
- Knowledge of cost optimization tools (Kubecost) and cluster autoscaling strategies.
- Experience with chaos engineering / resilience testing tools.
Soft Skills
- Strong communication skills to collaborate with Dev, SRE, Security, and Network teams.
- Ability to work independently under SLA pressure in a 24x7 support environment.