Search by job, company or skills

SRE Engineer- Data Platform

Early Applicant
  • Posted 4 days ago
  • Be among the first 10 applicants

Job Description

Role Overview:

We are offering an exciting opportunity to contribute to MetLife's digital and AI transformation journey. We are seeking an experienced Site Reliability Engineer (SRE) for the Data domain to ensure highly available, scalable, performant, and reliable data platforms, pipelines, and services. This role will drive operational excellence, strengthen observability, improve incident response, reduce toil through automation, and partner with engineering, data, cloud, and operations teams to enable resilient data services and measurable business outcomes.

Core Responsibilities:

  • Data Platform Reliability: Ensure the availability, performance, scalability, and reliability of data platforms, pipelines, and services through proactive monitoring, troubleshooting, and issue resolution.
  • Service Design and Implementation: Collaborate with engineering teams to design, implement, and operate large-scale data systems, including software and tooling to automate and streamline operations.
  • Automation and Scripting: Develop and maintain scripts, utilities, and automation tools to improve operational efficiency, reduce manual effort, minimize errors, and support repeatable data platform operations.
  • Monitoring and Alerting: Design, implement, and continuously improve monitoring, alerting, dashboards, and observability signals to enable early detection and timely resolution of issues.
  • Collaboration and Communication: Partner with engineering teams, product managers, data stakeholders, cloud teams, and operations teams to ensure services meet business requirements and align with company goals.
  • Incident Response and Management: Participate in incident response, service restoration, root cause analysis, postmortems, and corrective actions to improve resiliency and prevent recurrence.
  • Documentation and Knowledge Sharing: Maintain accurate documentation for systems, services, runbooks, processes, and operational procedures, and promote knowledge sharing across the team.
  • Governance and Day-to-Day Operations: Support day-to-day data platform operations while ensuring alignment with internal standards, global governance requirements, operational controls, and production support processes.

Skills & Experience:

  • Programming, Scripting, and Data Processing: Hands-on experience with Python, Spark, Bash, and PowerShell for automation, diagnostics, data processing support, operational tooling, and repeatable platform operations.
  • Azure Data Platform: Experience supporting Azure Data Lake Gen2, Azure Data Factory, Azure Synapse Analytics including Data Warehouse, Spark, and Pipelines, Azure SQL Database, Cosmos DB, and Databricks in production environments.
  • Observability and Monitoring: Proficiency with Azure Application Insights, Azure Log Analytics, Azure Monitor, Splunk, AppDynamics, and ELK to monitor availability, performance, pipeline health, application behavior, logs, metrics, traces, and alerts.
  • DevOps and Source Control: Working knowledge of Azure DevOps and GitHub Enterprise for repositories, CI/CD pipelines, release management, operational changes, branching strategies, and controlled production deployments.
  • Containers and Cloud-Native Operations: Experience with Docker, Kubernetes, and Azure Kubernetes Service (AKS) to support containerized workloads, platform reliability, deployment operations, scaling, and service health management.
  • ITSM and Operational Governance: Experience using ServiceNow or equivalent ITSM platforms for incident, problem, change, request, and knowledge management in a governed enterprise environment.
  • SRE and Production Reliability: Strong understanding of SLIs, SLOs, SLAs, error budgets, incident response, root cause analysis, postmortems, runbooks, automation, production readiness, and toil reduction practices.
  • Disaster Recovery & Resilience Engineering: Design and architect for multi-region systems to meet strict RTO (Service Availability) and RPO (Data Integrity) targets. You will drive down RTO via automation (Terraform, auto-scaling, DNS failover) and enforce RPO through robust replication and Continuous Data Protection (CDP) strategies.
  • AI-Assisted Engineering: Ability to use GitHub Copilot and Microsoft 365 Copilot responsibly to accelerate investigation, documentation, scripting, knowledge discovery, and operational productivity with appropriate human validation.
  • Collaboration and Execution: Strong communication, documentation, troubleshooting, evidence capture, escalation management, and cross-functional collaboration skills across engineering, data, cloud, product, and operations teams.

Minimum Qualification:

  • 8 – 12 years in production support, DevOps, infrastructure, cloud operations, or software engineering.
  • Experience supporting business-critical systems and working in incident, problem, and change management processes.
  • Ability to script and automate standard operational tasks using Python, PowerShell, Bash, or equivalent.
  • Bachelor's degree in computer science, engineering, or equivalent practical experience.
  • Exposure to regulated enterprise, insurance, banking, or financial services environments preferred.
  • Exposure to Hybrid cloud platforms including on-premises and Azure-hosted services.
  • Business proficiency in English; Japanese language skills are a plus.

Preferred Exposure:

  • Japanese language ability, including reading and writing.
  • Domain knowledge of life insurance business processes, data platforms, and operational support needs.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151581257