Search by job, company or skills

Site Reliability Engineer

Site Reliability Engineer

terragig
8-12 Years
Not Disclosed
Early Applicant
  • Posted 2 months ago
  • Be among the first 10 applicants

Job Description

Duration: 6 months

Must-Have Experience:

  • 8 -12 years in enterprise observability, SRE, or APM architecture
  • Proven track record building observability platforms from scratch (not just tool implementation)
  • Hands-on experience with Grafana, and familiarity with tools like Dynatrace, Datadog, Prometheus, or similar APM/monitoring platforms
  • Strategic architecture experience: reference architectures, maturity models, stakeholder requirements gathering
  • Experience with observability frameworks (MELT - Metrics/Events/Logs/Traces correlation)
  • Strong understanding of MTTD/MTTR measurement and improvement
  • Excellent English communication skills for US stakeholder collaboration

Preferred:

  • Azure cloud platform experience (our primary environment)
  • Financial services or professional services industry background
  • Experience working with US-based teams
  • Availability for overlap with US Eastern time zone
  • OpenTelemetry, synthetic monitoring, and AIOps experience

Role Scope: This is a hands-on architect role (player-coach) responsible for designing our enterprise observability strategy, conducting tool evaluations, implementing reference architectures, and building operational maturity frameworks. Not a pure management or pure implementation role.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

OpenTelemetry

Observability frameworks

MTTR measurement

AIOps

MELT - Metrics Events Logs Traces correlation

About Company

Similar Jobs

5-10 yrs
Gurugram, Gurugram, India
Skills:
DAST, Prometheus, Aem, Linux Scripting, Grafana, Datadog, DevSecOps, Terraform, Dynatrace, Kubernetes, AWS, CI CD, SCA, LLM-based solutions, OpenTelemetry, SAST, AIOps
7-9 yrs
Gurugram, Gurugram, India
Skills:
Elasticsearch, Prometheus, Kafka, Bash, MongoDB, Grafana, Python, Redis, HashiCorp Vault, WSO2 API Manager
5-8 yrs
Bengaluru, India
Skills:
Unix, Cloudformation, Prometheus, Bash, Grafana, Cloudwatch, Docker, Terraform, Linux, Ansible, Splunk, Kubernetes, Python, AWS, Open Telemetry, Go, EKS
10-12 yrs
Bengaluru, India
Skills:
Java, Cloudformation, Prometheus, Datadog, Terraform, Ansible, Splunk, Puppet, Kubernetes, Python, AWS, GitOps, Go
5-8 yrs
Hyderabad, India
Skills:
.NET, Java, Prometheus, Nodejs, Azure Log Analytics, Grafana, JIRA, Datadog, Gcp, Docker, Splunk, Azure, Kubernetes, Python, AWS, C#, Https, Storage, Load Balancing, Tcp/ip, Dns, SSL, Tls, File Systems, PagerDuty, CI/CD pipelines, Infrastructure-as-Code