Search Jobs

Search by job, company or skills

DevOps and Reliability Lead

DevOps and Reliability Lead

erelego technologies pvt ltd (etpl)
Early Applicant
  • Posted 22 hours ago
  • Be among the first 10 applicants

Job Description

About ETPL

eReleGo Technologies Private Limited (ETPL) is a Bengaluru-based technology company

building scalable digital platforms across AdTech, EdTech, SaaS ERP, and FinTech

domains. The company focuses on applying advanced technologies such as artifici

intelligence and data-driven systems to solve real operational problems for businesses

and educational institutions.

As ETPL scales into a firm of larger size,revenue, and complexity, it is investi

deliberately in building strong internal capabilities across leadership, operations,

technology, people, and culture. The objective is not justrapid growth, but sustainable,

future-ready growth supported by structured systems, strong talent, and accountable

teams.

Role Overview

We are seeking a DevOps and Reliability Lead with 6 to 9 years of experience to own

ETPL's infrastructure, deployment, and reliability engineering practice across both ETPL

Digital and ETPL AI. This is a seniortechnical leadership role with full ownership of the

systems, processes, and standards that keep ETPL's products running reliably, securely,

and at scale.

The DevOps and Reliability Lead will define and operate ETPL's CI/CD pipelines, clo

infrastructure, containerisation strategy, monitoring and alerting framework, and

incidentresponse practice - ensuring that engineering teams can ship with confiden

and that institutional clients experience consistent, high-availability service.

ETPL's products serve operationally sensitive environments - educational institutions

running academic cycles, lending platforms processing financial transactions, a

cooperative governance systems managing member data - where downtime and

performance degradation carry real institutional consequences. The Lead will bring both

the technical depth to architectrobust infrastructure and the operational discipline to

build a reliability culture across the engineering organisation.

Key Responsibilities

Own and evolve ETPL's cloud infrastructure across AWS / Azure - including

compute, networking, storage, database services, and security configurations

ensuring environments are scalable, cost-optimised, and production-ready

● Design, build, and maintain CI/CD pipelines across ETPL's product portfolio,

enabling fast,reliable, and consistent delivery of software from development

through staging to production

● Define and implement infrastructure-as-code (IaC) practices using Terrafor

Pulumi, or equivalent tooling, ensuring all infrastructure is version-controlled,

reproducible, and auditable

● Own ETPL's containerisation and orchestration strategy - managing Docker-based

build standards and Kubernetes cluster operations across product environments

● Build and maintain a comprehensive observability stack - including centralised

logging, metrics collection, distributed tracing, and alerting - using tools such as

Prometheus, Grafana, ELK Stack, Datadog, or equivalent

● Define and own service level objectives (SLOs) and service level indicators (SLI

for ETPL's production systems, working with product and engineering teams to

align reliability targets with institutional client expectations

● Lead incidentresponse across ETPL's production environment - owning the

on-call framework, incident classification, escalation paths, post-incidentreview

and structured follow-through on remediation actions

● Conduct capacity planning and performance modelling for ETPL's product

infrastructure, anticipating growth requirements and ensuring systems can scale

ahead of demand

● Establish and enforce security and compliance standards across the infrastructure

layer- including network security, secrets management, access control,

vulnerability scanning, and data protection practices appropriate to ETPL's

institutional client obligations

● Manage database infrastructure operations - including backups,replication,

failover configuration, and performance tuning - in coordination with produ

engineering teams

ETPL AI - DevOps and Reliability Lead

Page 4

● Drive the adoption of DevOps culture and practices across ETPL's engineering

teams - including developer self-service, deployment ownership, and shared

accountability for production reliability

● Evaluate, select, and govern the use of infrastructure tooling, managed services,

and third-party platform integrations across the engineering organisation

● Lead and develop a small team of DevOps and infrastructure engineers, setting

clear expectations,reviewing work quality, and building capability within the

function

Experience and Profile

● 6 to 9 years of progressive experience in DevOps, site reliability engineering, or

infrastructure engineering roles

● Demonstrated experience owning cloud infrastructure and CI/CD pipelines for

production SaaS or enterprise technology products at meaningful scale

● Proven hands-on experience with containerisation and Kubernetes in a

production environment

● Experience building and operating observability stacks and leading structured

incidentresponse processes

● Experience implementing infrastructure-as-code and bringing discipline to

infrastructure management within a growing engineering organisation

● Prior experience in a multi-product or multi-tenant SaaS environment is strongly

preferred

● Experience managing or mentoring junior DevOps or infrastructure engineers is

expected at this level

Skills and Attributes

● Strong hands-on proficiency across at least one major cloud platform - AWS 

Azure - including compute (EC2 / GKE / AKS), networking (VPC, load balancers,

DNS), and storage services

● Deep expertise in containerisation (Docker) and Kubernetes - including cluster

management, workload configuration, autoscaling,resource governance, a

Helm chart management

● Proficiency in CI/CD tooling - GitHub Actions, Jenkins, PagerDuty - with t

ability to design pipelines that balance speed, safety, and deployment confiden

ETPL AI - DevOps and Reliability Lead

● Strong infrastructure-as-code capability using Terraform, Pulumi, or equivalent,

with a clear understanding of state management, module design, and environment

parity

● Solid observability engineering skills - building and maintaining logging (ELK, Loki),

metrics (Prometheus, Grafana), and tracing (Jaeger, OpenTelemetry) pipelines for

complex distributed systems

● Experience defining SLOs and SLIs, and using error budget thinking to balan

reliability investment with delivery velocity

● Strong incident management capability - able to lead calmly and methodically

under production pressure, drive structured post-incidentreviews, and ensure

remediation actions are followed through

● Working knowledge of infrastructure security - secrets management (Vault, AWS

Secrets Manager), IAM design, network segmentation, vulnerability scanning, and

compliance controls relevant to enterprise SaaS platforms handling sensitive

institutional data

● Proficiency in scripting languages - Bash, Python, or equivalent - for automatio

tooling, and operational workflow developme

● Familiarity with database operations - MySQL, PostgreSQL - including backup

management,replication configuration, failover, and query performan

considerations at the infrastructure level

● Strong analytical capability - able to interpret system metrics, identify reliability

risks, and make well-reasoned infrastructure decisions under uncertainty

● Clear communication skills - able to translate infrastructure concepts and

reliability risks into plain language for product, engineering, and leadership

audiences

● A collaborative, platform-thinking mindset - viewing the engineering teams as

internal customers and orienting infrastructure decisions around enabling their

productivity and delivery confiden

To apply send you resume to [Confidential Information] or WhatsApp on +91-9167464004

More Info

Job Type:
Industry:
Function:
Employment Type: