Search Jobs

Search by job, company or skills

Lead Site Reliability Engineer

Lead Site Reliability Engineer

Deutsche Borse
Early Applicant
  • Posted 2 months ago
  • Be among the first 50 applicants

Job Description

About Deutsche Börse Group:

Headquartered in Frankfurt, Germany, we are a leading international exchange organization and market infrastructure provider. We empower investors, financial institutions, and companies by facilitating access to global capital markets. Our business areas cover the entire financial market transaction process chain, including trading, clearing, settlement and custody, digital assets and crypto, market analytics, and advanced electronic systems. As a technology-driven company, we develop and operate cutting-edge IT solutions globally.

About Deutsche Börse Group in India:

Our presence in Hyderabad serves as a key strategic hub, comprising India's top-tier tech talent. We focus on crafting advanced IT solutions that elevate market infrastructure and services. Together with our colleagues from across the globe, we are a team of highly skilled capital market engineers forming the backbone of financial markets worldwide. We harness the power of innovation in leading technology to create trust in the markets of today and tomorrow.

Site Reliability Engineer – Cloud Native Platforms (GCP / Azure)

Corporate IT – Deutsche Börse Group

We're building the next generation of cloud-native operations at Deutsche Börse Group. As part of our Corporate IT Cloud Infrastructure Operations team, we're looking for a Site Reliability Engineer (SRE) who's passionate about automation, reliability, and modern cloud technologies.

This is not just another ops role, it's an opportunity to shape how we run and scale our cloud-native platforms (GCP, Azure) in a highly regulated financial environment. You'll work at the intersection of development and infrastructure operations, helping us evolve our platform capabilities while ensuring performance, resilience, and security.

What You will Do:

  • Design, deploy and manage scalable and reliable systems in GCP and Azure, using serverless technologies, containerization, AI / ML and PaaS-based infrastructure as a code, with a strong focus on automation and observability.
  • Integrate AI / ML technologies into internal tools and workflows to drive automation and efficiency.
  • Operate and maintain geo-redundant, business-critical services leveraging automation and observability tools, ensure transparency and fast issue resolution.
  • Collaborate with Development Teams to implement best practice for cloud infrastructure, ensuring high availability and scalability of applications.

What You Bring:

  • Proven experience as a DevOps / SRE Engineer or a similar role
  • Expertise in managing and optimizing GCP or Azure cloud-native services and AI/ML integration.
  • Experience or knowledge of Container technology such as Docker, Buildah and Kubernetes (GKE, AKS)
  • Must have 2+ scripting and programming experience (Python, Bash)
  • Proficiency in infrastructure-as-code tools, particularly Terraform and ArgoCD
  • Familiarity with observability tools such as Prometheus, Grafana, OpenTelemetry
  • Solid understanding of CI/CD concepts

Why Join Us

  • Be part of a growing SRE team that's building new expertise in cloud-native operations.
  • Work with modern technologies in a mission-critical environment.
  • Help shape the future of IT operations at one of Europe's leading financial institutions.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

observability

GKE

Buildah

AKS

OpenTelemetry

ArgoCD

serverless technologies

About Company

Similar Jobs

5-7 yrs
Hyderabad, India
Skills:
Postgres Sql, PowerShell, Bash, Itil, Sql, Datadog, Celery, Docker, Terraform, Arm, Cosmos DB, Kubernetes, Application Insights, LangChain, Checkly, Microsoft Azure Cloud, AI ML-based anomaly detection, Playwright, Log Analytics, Kusto, OpenTelemetry, OpenAI APIs, Bicep, Azure Monitor
5-7 yrs
Hyderabad, India
Skills:
.NET, Datadog, Networking Technologies, Java, Scalability, Continuous Delivery, Grafana, Continuous Integration, ECS, Terraform, Splunk, Spring Boot, Performance, Dynatrace, Gitlab, Prometheus, Kubernetes, Python, Docker, Jenkins, toil reduction, Security, telemetry collection, white and black box monitoring, Reliability, enterprise system architecture, alerting, observability, SLO
5-7 yrs
Hyderabad
Skills:
operational readiness , Firewalls, Ansible, Shell, Vmware Nsx, Routing, Load Balancers, Python, Proxies, Sda, Switching, Cisco ACI Fabrics, AI-assisted reliability workflows, SD-WAN
6-12 yrs
Hyderabad
Skills:
SQL/ MySQL, Java, Unix/Linux, Openshift, Gitlab, Shell Scripting, MongoDB, Python, Git/Github, GCP Cloud, Splunk/Dynatrace