Site Reliability Engineering Specialist
Site Reliability Engineering Specialist
BT Group- Posted 10 hours ago
- Be among the first 10 applicants
Job Description
Job Title: Site Reliability Engineering Specialist
Req ID: 62585
Job Function: Software Engineering
Posting Start Date: 30/09/2026
Posting End Date: 05/10/2026
Division: BT International
Job Location: IND-Bengaluru-Pritech
Advertised Salary: Competitive
Job Req ID: 62585
Posting Date: 10-Sep-26
Function: Software Engineering
Location: Bengaluru
Salary: Competitive
About The Role
As a Site Reliability Engineer (SRE) within the Network Operations team, BTI International, you will be reliable for ensuring the reliability, resilience and performance of our Global Platforms including Global Fabric. You will collaborate closely with Engineering, Product and ASG teams to embed SRE principles such as automation, observability and proactive incident reduction into day to day operations. By improving how we monitor, maintain and evolve our services, you will help reduce risk, improve service quality and increase operational efficiency. Through this role, you will help BTI International's strategy by enabling stable, secure and scalable platforms that help business growth, accelerate delivery of new capabilities, and protect customer experience.
What You'll Be Doing
Req ID: 62585
Job Function: Software Engineering
Posting Start Date: 30/09/2026
Posting End Date: 05/10/2026
Division: BT International
Job Location: IND-Bengaluru-Pritech
Advertised Salary: Competitive
Job Req ID: 62585
Posting Date: 10-Sep-26
Function: Software Engineering
Location: Bengaluru
Salary: Competitive
About The Role
As a Site Reliability Engineer (SRE) within the Network Operations team, BTI International, you will be reliable for ensuring the reliability, resilience and performance of our Global Platforms including Global Fabric. You will collaborate closely with Engineering, Product and ASG teams to embed SRE principles such as automation, observability and proactive incident reduction into day to day operations. By improving how we monitor, maintain and evolve our services, you will help reduce risk, improve service quality and increase operational efficiency. Through this role, you will help BTI International's strategy by enabling stable, secure and scalable platforms that help business growth, accelerate delivery of new capabilities, and protect customer experience.
What You'll Be Doing
- Provide end-to-end SRE ownership for the Global Fabric service, ensuring platform reliability, performance, resilience, and operational excellence.
- Own incidents across the customer journey, from detection through to resolution, working with the appropriate ASGs and support teams.
- Act as the primary operational escalation point for CF-related incidents participating in an on-call rota.
- Escalate complex technical issues appropriately and coordinate resolution activities
- Manage incidents through ServiceNow and track defects and improvements through Jira.
- Perform root cause analysis and implement preventative actions to prevent recurrence along with providing documentation and knowledge transfer sessions.
- Design, implement, operate, and continuously improve observability and monitoring solutions using Dynatrace.
- Drive automation initiatives using Ansible, scripting, CI/CD pipelines, and GitOps practices to reduce operational toil and improve service reliability.
- Define, measure, and report service health metrics, SLIs, SLOs, error budgets, and operational performance dashboards.
- Fulfil operational service requests in line with agreed SLAs.
- Support onboarding of new customers and operational readiness activities.
- Analyse platform, network, and service trends to identify reliability and optimisation opportunities.
- Experience supporting large-scale, high-availability services in an ISP / NaaS / network-centric environment.
- Experience delivering changes through CI/CD and GitOps processes, including release validation, deployment governance, monitoring verification, and rollback planning.
- Experience operating customer-facing applications, APIs, and distributed services with a strong focus on reliability, availability, performance, and customer experience.
- Proven ability to troubleshoot end-to-end customer fulfilment journeys across UI, APIs, middleware, event platforms, and downstream systems.
- Strong observability skills for CF services, including journey‑based monitoring and synthetic checks.
- Knowledge of Infrastructure as Code tools like Terraform or Ansible.
- Knowledge of event-driven architectures, including Kafka concepts such as message delivery, lag monitoring, loss detection, replay, and troubleshooting.
- Working knowledge of incident/problem management in ServiceNow and delivery tracking in Jira (Scrum / PI planning).
