Role Overview
The Observability Engineer leads the enterprise observability and service reliability strategy across infrastructure and application domains. This role drives proactive monitoring, automation, and resilience initiatives to enhance system availability, ensure regulatory compliance, and reduce mean time to resolution (MTTR).
The ideal candidate should have an SRE background with administrator-level experience in Datadog, utilizing it for Application Performance Monitoring (APM), Real User Monitoring (RUM), Synthetic Monitoring, Datadog agent configurations, and logs monitoring (pipeline and parsing).
Key Responsibilities
- Observability Strategy & Governance: Define enterprise observability architecture aligned with operational resilience standards. Deploy and optimize full-stack observability platforms (metrics, logs, traces) and integrate them with ITSM and AIOps systems for predictive alerting.
- Reliability Engineering & Automation: Implement SRE frameworks, define error budget policies, and codify operational reliability. Automate runbooks, self-healing workflows, and auto-remediation actions using Python, Ansible, and Terraform.
- Cloud & Platform Observability: Architect and manage telemetry solutions for cloud-native workloads across AWS and Azure, embedding observability into landing zones and CI/CD pipelines via Infrastructure-as-Code (IaC).
- Operational Excellence: Partner with cross-functional teams to conduct resilience testing, chaos engineering, and capacity validation. Maintain executive dashboards for operational risk indicators and act as a technical advisor during major incidents and audits.
Competencies and Qualifications
- Technical Background: Qualification in Computer Science, Information Technology, or equivalent practical experience.
- Domain Experience: 5 years of experience and track record in Infrastructure, Cloud, or Site Reliability Engineering (SRE), with experience operating as an SRE subject matter expert, ideally within financial services or regulated environments.
- Observability Platforms: Hands-on proficiency with tools - Datadog (Highly desired), Dynatrace, Splunk, or ELK Stack.
- Automation & Cloud: Expertise in Infrastructure-as-Code and scripting (Terraform, Ansible, Python) alongside cloud observability tooling (AWS CloudWatch/X-Ray, Azure Monitor/Log Analytics).
- SRE & Compliance: Deep understanding of SRE principles (SLOs/SLAs, error budgets), financial sector operational resilience frameworks (e.g., MAS TRM, DORA), and automated remediation.
- Certifications: Relevant certifications in observability platforms, IaC, cloud engineering (AWS/Azure), SRE, or ITIL are advantageous.