

Search by job, company or skills
Position - SR. Site Reliability Engineer
Experience - 12+ years
Location - Remote for India
Employment type- Full time
Project- UAE project
Skills- Azure, Devops, Python/shell,
Job Responsibilities:
• Design, implement, and maintain highly reliable, scalable, and observable systems on Microsoft Azure.
• Define and operationalize SLIs, SLOs, and SLAs to measure and improve service reliability.
• Build and enhance observability platforms using tools like Grafana, Prometheus, ELK stack, and OpenTelemetry.
• Drive adoption of SRE principles (error budgets, toil reduction, automation-first mindset).
• Implement proactive monitoring, alerting, and incident response frameworks.
• Lead incident management, root cause analysis (RCA), and postmortems with a blameless culture. • Automate infrastructure and workflows using Infrastructure as Code (IaC).
• Collaborate with engineering teams to improve system resilience, performance, and deployment practices.
• Develop and maintain runbooks, playbooks, and operational standards.
• Advocate for DevOps and SRE culture adoption across teams
Desired Skill:
SRE Practices
• Hands-on experience implementing: o SLIs, SLOs, SLAs o Error budgets o Toil reduction strategies
• Strong understanding of incident management lifecycle
Relevant Exp:
Programming & Automation
• Proficiency in Python (automation, tooling, scripting)
• Experience building internal tools for reliability and observability Infrastructure as Code (IaC)
• Strong experience with Terraform
• Familiarity with infrastructure automation and configuration management
• Experience with GitHub Actions (or similar CI/CD tools)
• Knowledge of deployment strategies (blue-green, canary, rolling updates)
Value Add:
Good to Have Experience with AI/ML Observability (monitoring models, drift detection, LLM observability)
• Familiarity with: o Service Mesh (Istio, Linkerd) o Chaos Engineering tools (e.g., Chaos Monkey, Litmus) o Distributed tracing tools (Jaeger, Tempo)
• Exposure to FinOps practices (cost optimization in cloud)
• Experience with multi-cloud or hybrid environments
Comment:
Technical Skills: Cloud & Infrastructure
• Strong hands-on experience with Microsoft Azure
• Experience with Kubernetes (AKS) and containerized environments
• Knowledge of networking, load balancing, and distributed systems Observability & Monitoring
• Experience with: o Grafana (dashboards, alerting) o Prometheus / OpenTelemetry o ELK Stack (Elasticsearch, Logstash, Kibana)
• Ability to design end-to-end observability (metrics, logs, traces)
Job ID: 151284405
Skills:
Jenkins, Terraform, Ansible, Bash, Datadog, Kubernetes, Python, AWS
Skills:
Jenkins, Terraform, Ansible, Bash, Datadog, Kubernetes, Python, AWS
Skills:
Elk, PowerShell, Prometheus, Bash, Grafana, Datadog, Zabbix, Gcp, Terraform, Docker, Ansible, Splunk, Nagios, Puppet, Azure, Python, Kubernetes, AWS, Linux Unix system administration, Chef, Go, Istio
Skills:
Storm, Cassandra, Prometheus, Kafka, Docker, Terraform, Elasticsearch, Shell scripting, Postgres, Gitlab, Python, AWS, Rust, Cloudformation, Redis, Jenkins, Cloudwatch, Gcp, Linux, Ansible, Spark, Kubernetes, Go, Flink, GitHub Actions, ArangoDB, Stackdriver
Skills:
containerization , Kibana, Perforce, Prometheus, Kafka, Tableau, Grafana, Nosql, Docker, Infrastructure Management, System Administration, MySQL, AWS, Automation, Zabbix, Devops, Jenkins, Git, Gcp, Ansible, Elastic Search, Puppet, Azure, Kubernetes, Virtualization, Chef, Filebeat, Monitoring