Search by job, company or skills

Senior Site Reliability Engineer (SRE) / DevOps Engineer

7-9 Years
Early Applicant
  • Posted 5 hours ago
  • Be among the first 10 applicants

Job Description

Location: Viman Nagar, Pune – Work From Office

Experience Overall(must have): 8 Years

CTC: Up to ₹25 LPA

Notice Period: Immediate Joiners Only within 15d or (if serving max 30days)

Working Hours: 3:00 PM – 12:00 AM, Monday to Friday

On-Call: 24/7 Production Support – On-Call Rotation Required

Employment Type: Full-Time

About The Role

We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to manage and improve the reliability, scalability, performance, security, and observability of mission-critical production environments.

The role requires strong hands-on expertise in Cloud, Kubernetes, DevOps automation, Monitoring & Observability, Incident Management, and SRE practices. The ideal candidate should be comfortable handling production incidents while also driving long-term initiatives around reliability, automation, scalability, and reduction of operational toil.

Must-Have Skills & Experience1. SRE & Production Operations

  • Relevant 7+ years of relevant experience in SRE / DevOps / Cloud Infrastructure / Production Engineering.
  • Hands-on experience with 24/7 production support and on-call operations.
  • Strong experience in incident management, troubleshooting, RCA, and post-mortems.
  • Good understanding of SLI, SLO, SLA, Error Budgets, MTTR, and reliability engineering.
  • Experience with toil reduction, capacity planning, high availability, disaster recovery, and failover strategies.
  • Ability to improve system availability, performance, scalability, and operational reliability.
  • Cloud & Infrastructure
  • Strong hands-on experience with Microsoft Azure, AWS, and/or GCP.
  • Strong understanding of cloud infrastructure, networking, IAM, storage, compute, and cloud-native services.
  • Hands-on experience with at least one major cloud platform and good exposure to multi-cloud environments.
  • Experience with:
    • Azure: VMs, Networking, Storage, IAM, Azure Monitor, AKS
    • AWS: EC2, S3, RDS, IAM, VPC, CloudWatch, EKS
    • GCP: Compute Engine, Cloud Storage, IAM, VPC, GKE, Cloud Monitoring
  • Kubernetes & Containerization
  • Strong hands-on experience with Kubernetes and containerized workloads.
  • Experience with AKS / EKS / GKE or equivalent Kubernetes environments.
  • Hands-on experience with Helm deployments.
  • Understanding of Kubernetes troubleshooting, scaling, networking, and workload management.
  • Infrastructure as Code & DevOps
  • Hands-on experience with Terraform / Infrastructure as Code (IaC).
  • Experience with Git-based workflows using GitHub, GitLab, or Azure Repos.
  • Strong DevOps automation and CI/CD understanding.
  • Strong scripting skills in Python and/or Bash.
  • Monitoring & Observability
  • Strong hands-on experience with OpenTelemetry.
  • Experience with monitoring and observability tools such as:
    • Prometheus
    • Grafana
    • Datadog
    • Azure Monitor
    • AWS CloudWatch
    • GCP Cloud Monitoring
  • Strong understanding of metrics, logs, distributed tracing, and alerting.
  • Experience implementing monitoring based on Golden Signals:
    • Latency
    • Traffic
    • Errors
    • Saturation
  • Ability to develop symptom-based, user-impact-focused alerting.
  • Linux & Networking
  • Strong knowledge of Linux system administration.
  • Strong understanding of:
    • DNS
    • TCP/IP
    • Load Balancing
    • SSL/TLS
    • Networking fundamentals
  • Experience supporting highly available production environments.
  • Incident & Reliability Engineering
  • Ability to rapidly diagnose and resolve high-severity production incidents.
  • Experience driving MTTR reduction.
  • Strong debugging and analytical problem-solving skills.
  • Ability to identify recurring issues and implement permanent corrective/preventive solutions.
Good-to-Have Skills

  • Experience working across Azure + AWS + GCP in a multi-cloud environment.
  • Knowledge of Go (Golang).
  • Experience with OpenSearch / ELK Stack.
  • Experience supporting AI/ML workloads in production.
  • Exposure to Azure AI Services and Azure AI Foundry.
  • Experience supporting RAG (Retrieval-Augmented Generation) workloads.
  • Experience designing infrastructure for AI/ML platforms.
  • Experience building enterprise-wide OpenTelemetry observability frameworks.
  • Strong understanding of distributed systems architecture.
  • Exposure to advanced cloud-native architectures and reliability patterns.
  • Experience with security, compliance, vulnerability remediation, secrets management, and network segmentation.

Key ResponsibilitiesProduction & Incident Management

  • Participate in the 24/7 on-call rotation.
  • Diagnose, mitigate, and resolve production incidents.
  • Lead RCA and post-incident reviews.
  • Implement corrective and preventive actions.
  • Continuously improve MTTR and production stability.

Reliability Engineering

  • Define and improve SLIs, SLOs, SLAs, and Error Budgets.
  • Identify and eliminate operational toil.
  • Conduct reliability and capacity reviews.
  • Improve redundancy, failover, disaster recovery, and system resilience.

Cloud & Infrastructure

  • Manage and optimize cloud infrastructure across Azure, AWS, and/or GCP.
  • Manage Kubernetes clusters and containerized applications.
  • Implement and maintain Infrastructure as Code using Terraform.
  • Support CI/CD and Git-based development workflows.

Observability & Performance

  • Build and improve monitoring, logging, metrics, and tracing.
  • Implement OpenTelemetry and distributed tracing.
  • Establish Golden Signals-based monitoring and alerting.
  • Identify and resolve infrastructure and application performance bottlenecks.

Security

  • Implement cloud security best practices around IAM, network segmentation, and secrets management.
  • Support vulnerability remediation and compliance initiatives.
  • Collaborate with Development, Security, and Infrastructure teams.

Ideal Candidate

We Are Looking For Someone With

  • Strong SRE mindset and production ownership.
  • Excellent troubleshooting and incident-management skills.
  • Hands-on expertise in Cloud + Kubernetes + Terraform + Observability.
  • Strong understanding of OpenTelemetry and Golden Signals.
  • Experience working in highly available, production-critical environments.
  • Ability to remain calm and make effective decisions during critical incidents.
  • Strong communication and cross-functional collaboration skills.
  • Passion for automation, scalability, reliability, and continuous improvement.

Important Hiring Criteria

Must Be

  • 7+ years relevant experience
  • Immediate joiner
  • Willing to work from office in Viman Nagar, Pune
  • Comfortable with 3:00 PM – 12:00 AM shift
  • Comfortable with 24/7 on-call rotation
  • Strong hands-on SRE/DevOps experience
  • Strong Cloud + Kubernetes + Observability experience
  • Strong production incident management experience

Good To Have

  • Multi-cloud: Azure + AWS + GCP
  • OpenTelemetry
  • AI/ML or RAG production workloads
  • Azure AI / AI Foundry
  • Go
  • OpenSearch / ELK
  • Distributed systems

Skills: ms azure,aws,sla,production engineering,mttr,terraform,python,golang,troubleshooting,sre,iac,gitlab,24/7 production support,capacity planning,sre & production operations,toil reduction,linux system,github,golden signals,incident management,opentelemetry,error budgets,rca,devops,kubernetes & containerization,slo,gcp,cloud infrastructure,sli

More Info

Job Type:
Industry:
Function:
Employment Type:

About Company

Job ID: 153727921

Beware of Scammers

We don’t charge money for job offers