Description
We are seeking a Senior Site Reliability Engineer to join our team in India. The ideal candidate will have a strong focus on SLOs and Error Budgets, ensuring the reliability and performance of our web services, batch pipelines, and real-time data flows. You will work closely with development teams to implement best practices for system reliability and automation.
What You'll Do
- Define and manage SLOs, SLIs, and error budgets for web services, batch pipelines, and real-time data flows.
- Automate deployments, batch operations, data transfers, health checks, and common run tasks through self-service tooling.
- Lead incident response across web, application, data, and database tiers; conduct blameless postmortems and implement preventive controls.
- Promote SRE practices, including production ownership, runbooks, readiness reviews, and operational standards.
- Lead Production Readiness Reviews focused on observability, capacity, scalability, and disaster recovery.
- Modernize legacy systems and manual processes using infrastructure-as-code, templated deployments, and centralized observability with Ansible, Helm, GitLab CI/CD, and Elasticsearch.
- Build end-to-end monitoring, dashboards, alerts, and tracing to support effective incident analysis.
- Improve on-call practices and help teams develop reliable, operable, and debuggable services.
What You Bring
- 10+ years of SRE/DevOps experience supporting business-critical systems across Linux/UNIX, some Windows, and multiple environments.
- Strong Linux/UNIX expertise and automation skills using Java, Python, shell scripting, or similar languages.
- Experience operating and tuning web/application tiers, batch workloads, and relational databases.
- Hands-on experience with:
- Ansible for configuration management, infrastructure-as-code, and repeatable deployments.
- GitLab CI/CD for build, test, deployment, gated releases, and automated checks.
- Elasticsearch for centralized logging, observability pipelines, index lifecycle management, and incident analysis.
- Helm and templated deployment practices.
- Ability to manage environment-specific release configurations and implement configuration-as-code practices.
- Experience with financial data workflows such as reference data, pricing, corporate actions, entitlements, calculations, reporting, and publishing.
- Knowledge of secure data exchange patterns.
- Excellent English communication, incident leadership, documentation, and cross-functional collaboration skills.
Skills:
- Legacy/niche middleware and protocols: CORBA/Orbix, IIOP; UDP/TCP messaging patterns.
- Scheduler: $Universe.
- Datastores beyond common RDBMS: Sybase ASE, BerkeleyDB (real-time).
- App stack specifics: Apache Tomcat/HTTPD mod_proxy tuning, Nginx as proxy, Java frameworks (Spring/Struts/JSP/GWT).