Search Jobs

Search by job, company or skills

Site Reliability Engineer

Site Reliability Engineer

mumba technologies, inc.
Early Applicant
  • Posted 25 days ago
  • Be among the first 20 applicants

Job Description

Job Title: Site Reliability Engineer

Job Type: Full Time

Location: Gurgaon (Hybrid 3-Days in the office)

About the Role:

We are looking for an experienced Site Reliability Engineer (SRE) to help build and operate highly reliable, scalable, secure, and observable systems. The role involves hands-on work across AWS, Kubernetes, Infrastructure as Code, observability, automation, security, and AI-driven SRE practices.

Reliability & Availability

  • Lead incident response, conduct RCAs and ensure action items are tracked to closure
  • Build and maintain runbooks, playbooks and escalation frameworks for proactive and reactive response
  • Drive toil reduction by identifying repetitive operational work and engineering it away

Observability

  • Design and own the full observability stack — metrics, logs, traces and events — using tools like Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace or similar
  • Build intelligent alerting that reduces noise, eliminates alert fatigue and surfaces actionable signals
  • Implement distributed tracing and dependency mapping to provide end-to-end visibility across microservices
  • Drive adoption of continuous profiling and real user monitoring (RUM) for proactive performance management

AI Adoption in SRE

  • Leverage AIOps platforms to enable anomaly detection, predictive alerting and automated root cause analysis
  • Implement AI-assisted incident triage — using LLM-powered tools to summarise incidents, suggest fixes and accelerate MTTR
  • Build and maintain ML-powered capacity forecasting models to optimise infrastructure spend and prevent resource saturation

Security & Vulnerability Management

  • Embed security-as-reliability principles — treating security incidents with the same urgency as availability incidents
  • Own issue remediations across infrastructure (OS, containers, dependencies)
  • Integrate SAST, DAST and SCA tools into CI/CD pipelines to shift security left

Infrastructure & Platform Engineering

  • Design, build and maintain cloud-native infrastructure on AWS using Infrastructure as Code (Terraform, Pulumi) & drive rightsizing, reserved capacity planning and cost anomaly detection
  • Own Kubernetes cluster operations — autoscaling, resource management, networking and upgrade strategy

Leadership & Culture

  • Mentor and guide junior and mid-level SREs — conducting technical reviews and pair debugging sessions
  • Define and evolve SRE team standards, best practices and engineering principles
  • Collaborate closely with product, development and security teams as an embedded reliability partner
  • Contribute to on-call rotation and drive continuous improvement of on-call experience
  • Represent SRE in architecture reviews, sprint planning and cross-functional forums

The Competitive Edge

AEM Administration

  • Own end-to-end reliability and availability of AEM environments — Author, Publish, Dispatcher and AEM as a Cloud Service (AEMaaCS) — across dev, staging and production
  • Monitor and manage AEM instance health, optimise Dispatchers, Manage DAM, OSGi Configurations, replication queues.

Exposure to CDN

  • Experience with Cloudflare - CDN, Workers.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

Infrastructure as Code

AEM Administration

ML-powered capacity forecasting

OpenTelemetry

AEM as a Cloud Service (AEMaaCS)

Distributed tracing

LLM-powered tools

Replication queues

Continuous profiling

CI/CD pipelines

Cloud-native infrastructure

Kubernetes cluster operations

Author Publish Dispatcher

OSGi Configurations

Cloudflare - CDN Workers

AIOps platforms

Real user monitoring (RUM)

Similar Jobs

Noida, India
Skills:
Monitoring Tools, Http, Datadog, Https, Gcp, Linux, Terraform, Distributed Systems, Splunk, Kubernetes, AWS, Infrastructure as Code, GitLab CI, observability
2-4 yrs
Gurugram, Gurugram, India
Skills:
PowerShell, Apigee, Prometheus, Bash, New Relic, Git, Docker, Terraform, Apache Kafka, Microsoft Azure, Kubernetes, Python, Azure DevOps, Azure API Management, GitHub Actions, Azure Service Bus, CI/CD pipelines, Event Hubs, FluentBit, Kong
Noida
Skills:
Https, Gcp, Terraform, Linux, Http, Splunk, Datadog, Kubernetes, Python, AWS, GitLab CI
1-3 yrs
Gurugram, India, Gurugram
Skills:
AWS, Grafana, Linux, Prometheus, Kubernetes, Python, Bash, Azure, Docker, Gcp, Helm, Terraform, Git, Networking fundamentals, IaC, Tempo
Gurugram, India, Gurugram
Skills:
Linux Administration, Gcp, AWS, Exception Management, operating a production platform with live services, configuration management tools, observability and monitoring best practices, Production on-call duties, Event Streaming, infrastructure as code, Kubernetes administration, batch-processing frameworks