Search by job, company or skills

Site Reliability Engineer

Site Reliability Engineer

mumba technologies, inc.
5-10 Years
Not Disclosed
  • Posted 17 hours ago
  • Be among the first 10 applicants

Job Description

Job Title: Site Reliability Engineer (SRE)

Job Type: Full Time

Location: Gurgaon (Hybrid)

Job Summary

We are looking for a Senior Site Reliability Engineer (SRE) with 5–10 years of experience to drive reliability, observability, automation, and cloud-native platform engineering. The ideal candidate will have strong hands-on experience with AWS, Kubernetes, Terraform, monitoring/observability, incident management, and AEM environments.

Key Responsibilities

Reliability & Observability

  • Lead production incidents, RCA, problem management, and reliability improvements.
  • Build and maintain observability using Prometheus, Grafana, Dynatrace, Datadog, OpenTelemetry or similar tools.
  • Develop actionable alerting, dashboards, distributed tracing, and performance monitoring.
  • Drive automation and toil reduction across production operations.

Cloud & Platform Engineering

  • Design and manage highly available AWS infrastructure using Terraform/Pulumi.
  • Manage Kubernetes clusters, including upgrades, autoscaling, networking, and resource optimization.
  • Drive cloud cost optimization, capacity planning, and infrastructure reliability.

AEM Administration

  • Manage reliability and availability of Adobe Experience Manager (AEM) environments across Author, Publish, Dispatcher, and AEM as a Cloud Service.
  • Troubleshoot AEM performance, replication queues, OSGi configurations, DAM, and Dispatcher issues.
  • Monitor AEM application and infrastructure health across Dev, QA, and Production.

AI & Automation

  • Apply AIOps/AI-assisted tools for anomaly detection, incident triage, RCA, and operational automation.
  • Leverage LLM-based solutions to improve troubleshooting and reduce MTTR.

Security

  • Support vulnerability remediation across OS, containers, dependencies, and cloud infrastructure.
  • Integrate SAST, DAST, and SCA practices into CI/CD pipelines.

Leadership & Collaboration

  • Mentor junior and mid-level engineers and establish SRE best practices.
  • Participate in architecture reviews, on-call rotations, and cross-functional engineering initiatives.
  • Partner with development, security, and product teams to improve overall platform reliability.

Required Skills

  • 5–10 years of experience in SRE, DevOps, Cloud Engineering, or Platform Engineering.
  • Strong AWS and Kubernetes experience.
  • Hands-on Terraform experience.
  • Strong knowledge of observability, monitoring, alerting, SLO/SLI, and incident management.
  • Experience with AEM Administration, including Author/Publish/Dispatcher.
  • Experience with CDN technologies, preferably Cloudflare.
  • Strong Linux, scripting, troubleshooting, and automation skills.
  • Experience with CI/CD and DevSecOps practices.
  • Strong communication, problem-solving, and stakeholder-management skills.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

Similar Jobs

7-9 yrs
Gurugram, Gurugram, India
Skills:
Elasticsearch, Prometheus, Kafka, Bash, MongoDB, Grafana, Python, Redis, HashiCorp Vault, WSO2 API Manager
4-7 yrs
Gurugram, Gurugram, India
Skills:
Java, Prometheus, Grafana, Sql, Bash Scripting, Git, Gcp, Dynatrace, Splunk, Cyberark, Python, AWS, Hashi Corp Vault, Open Telemetry
9-14 yrs
Gurugram, Gurugram, India
Skills:
Incident Management, change management, DDoS protection using AWS Shield Advanced, AI-assisted Workflow Process Automation, AWS WAF, Site Reliability Engineering, bot mitigation strategies, Prompt Engineering
8-10 yrs
Noida, India
Skills:
Java, Scala, Jenkins, Terraform, Ansible, Kubernetes, Cloud Formation, Scrum, Python, Infrastructure as Code, Go, AWS Networking, Sumo Logic products, Linux systems, CI/CD tooling, Command line, Cloud-native software security, Kanban
7-9 yrs
Gurugram, Gurugram, India
Skills:
Terraform, Azure, Kubernetes, AWS, CI CD tools, code quality tools, secret management tools