- Posted 22 hours ago
- Be among the first 10 applicants
Job Description
About ETPL
eReleGo Technologies Private Limited (ETPL) is a Bengaluru-based technology company
building scalable digital platforms across AdTech, EdTech, SaaS ERP, and FinTech
domains. The company focuses on applying advanced technologies such as artifici
intelligence and data-driven systems to solve real operational problems for businesses
and educational institutions.
As ETPL scales into a firm of larger size,revenue, and complexity, it is investi
deliberately in building strong internal capabilities across leadership, operations,
technology, people, and culture. The objective is not justrapid growth, but sustainable,
future-ready growth supported by structured systems, strong talent, and accountable
teams.
Role Overview
We are seeking a DevOps and Reliability Lead with 6 to 9 years of experience to own
ETPL's infrastructure, deployment, and reliability engineering practice across both ETPL
Digital and ETPL AI. This is a seniortechnical leadership role with full ownership of the
systems, processes, and standards that keep ETPL's products running reliably, securely,
and at scale.
The DevOps and Reliability Lead will define and operate ETPL's CI/CD pipelines, clo
infrastructure, containerisation strategy, monitoring and alerting framework, and
incidentresponse practice - ensuring that engineering teams can ship with confiden
and that institutional clients experience consistent, high-availability service.
ETPL's products serve operationally sensitive environments - educational institutions
running academic cycles, lending platforms processing financial transactions, a
cooperative governance systems managing member data - where downtime and
performance degradation carry real institutional consequences. The Lead will bring both
the technical depth to architectrobust infrastructure and the operational discipline to
build a reliability culture across the engineering organisation.
Key Responsibilities
Own and evolve ETPL's cloud infrastructure across AWS / Azure - including
compute, networking, storage, database services, and security configurations
ensuring environments are scalable, cost-optimised, and production-ready
● Design, build, and maintain CI/CD pipelines across ETPL's product portfolio,
enabling fast,reliable, and consistent delivery of software from development
through staging to production
● Define and implement infrastructure-as-code (IaC) practices using Terrafor
Pulumi, or equivalent tooling, ensuring all infrastructure is version-controlled,
reproducible, and auditable
● Own ETPL's containerisation and orchestration strategy - managing Docker-based
build standards and Kubernetes cluster operations across product environments
● Build and maintain a comprehensive observability stack - including centralised
logging, metrics collection, distributed tracing, and alerting - using tools such as
Prometheus, Grafana, ELK Stack, Datadog, or equivalent
● Define and own service level objectives (SLOs) and service level indicators (SLI
for ETPL's production systems, working with product and engineering teams to
align reliability targets with institutional client expectations
● Lead incidentresponse across ETPL's production environment - owning the
on-call framework, incident classification, escalation paths, post-incidentreview
and structured follow-through on remediation actions
● Conduct capacity planning and performance modelling for ETPL's product
infrastructure, anticipating growth requirements and ensuring systems can scale
ahead of demand
● Establish and enforce security and compliance standards across the infrastructure
layer- including network security, secrets management, access control,
vulnerability scanning, and data protection practices appropriate to ETPL's
institutional client obligations
● Manage database infrastructure operations - including backups,replication,
failover configuration, and performance tuning - in coordination with produ
engineering teams
ETPL AI - DevOps and Reliability Lead
Page 4
● Drive the adoption of DevOps culture and practices across ETPL's engineering
teams - including developer self-service, deployment ownership, and shared
accountability for production reliability
● Evaluate, select, and govern the use of infrastructure tooling, managed services,
and third-party platform integrations across the engineering organisation
● Lead and develop a small team of DevOps and infrastructure engineers, setting
clear expectations,reviewing work quality, and building capability within the
function
Experience and Profile
● 6 to 9 years of progressive experience in DevOps, site reliability engineering, or
infrastructure engineering roles
● Demonstrated experience owning cloud infrastructure and CI/CD pipelines for
production SaaS or enterprise technology products at meaningful scale
● Proven hands-on experience with containerisation and Kubernetes in a
production environment
● Experience building and operating observability stacks and leading structured
incidentresponse processes
● Experience implementing infrastructure-as-code and bringing discipline to
infrastructure management within a growing engineering organisation
● Prior experience in a multi-product or multi-tenant SaaS environment is strongly
preferred
● Experience managing or mentoring junior DevOps or infrastructure engineers is
expected at this level
Skills and Attributes
● Strong hands-on proficiency across at least one major cloud platform - AWS
Azure - including compute (EC2 / GKE / AKS), networking (VPC, load balancers,
DNS), and storage services
● Deep expertise in containerisation (Docker) and Kubernetes - including cluster
management, workload configuration, autoscaling,resource governance, a
Helm chart management
● Proficiency in CI/CD tooling - GitHub Actions, Jenkins, PagerDuty - with t
ability to design pipelines that balance speed, safety, and deployment confiden
ETPL AI - DevOps and Reliability Lead
● Strong infrastructure-as-code capability using Terraform, Pulumi, or equivalent,
with a clear understanding of state management, module design, and environment
parity
● Solid observability engineering skills - building and maintaining logging (ELK, Loki),
metrics (Prometheus, Grafana), and tracing (Jaeger, OpenTelemetry) pipelines for
complex distributed systems
● Experience defining SLOs and SLIs, and using error budget thinking to balan
reliability investment with delivery velocity
● Strong incident management capability - able to lead calmly and methodically
under production pressure, drive structured post-incidentreviews, and ensure
remediation actions are followed through
● Working knowledge of infrastructure security - secrets management (Vault, AWS
Secrets Manager), IAM design, network segmentation, vulnerability scanning, and
compliance controls relevant to enterprise SaaS platforms handling sensitive
institutional data
● Proficiency in scripting languages - Bash, Python, or equivalent - for automatio
tooling, and operational workflow developme
● Familiarity with database operations - MySQL, PostgreSQL - including backup
management,replication configuration, failover, and query performan
considerations at the infrastructure level
● Strong analytical capability - able to interpret system metrics, identify reliability
risks, and make well-reasoned infrastructure decisions under uncertainty
● Clear communication skills - able to translate infrastructure concepts and
reliability risks into plain language for product, engineering, and leadership
audiences
● A collaborative, platform-thinking mindset - viewing the engineering teams as
internal customers and orienting infrastructure decisions around enabling their
productivity and delivery confiden
To apply send you resume to [Confidential Information] or WhatsApp on +91-9167464004
