Lead the L1/L2/L3 support team responsible for non-core application stability, availability, EOD/BOD batch operations, DR readiness, and database/application monitoring.
Ensure 99%+ uptime, timely incident resolution, proactive monitoring, and strong collaboration with Engineering, DevOps, Infra, and Business teams.
Strengthen platform reliability through preventive fixes, automation, governance, and continuous improvement.
Key Result Areas
Supporting Actions
Application Uptime & Reliability
Maintain 99% uptime for non-core applications
Daily health checks, proactive alerting
Zero unplanned outages
Batch (EOD/BOD) & Scheduler Management
98% successful EOD/BOD jobs
Publish job status reports
Reduce batch failures by 20% YoY
Incident & Problem Management
95% incidents resolved within SLA
MTTR reduction by 20%
Zero repeat incidents through preventive actions
DR / BCP Readiness
100% DR drill success
Errorfree switchover/fallback
DR documentation updated quarterly
API/Gateway Performance
API gateway success rate >98%
Monitor latency, timeouts, circuit breaker events
Ensure certificate renewal before expiry
Team Leadership & Stakeholder Management
Daily task allocation & monitoring
Weekly RCA reviews, training & capability building