Job Description
The incumbent will be responsible for the architecture, design, implementation and day-to-day operations of the SIA Group's core Data Centre network, keeping mission-critical services highly available, performant, and secure. He/She requires an analytical mindset, good awareness and appreciation of operational practices, and strong expertise in automation and AI-enabled operations in order to reduce toil, improve incident response, and strengthen resilience at scale.
Key Responsibilities
- Service Operations & Reliability
- Provide operational support for core DC networking (routing/switching, segmentation, connectivity)
- Drive Incident, Problem, Change, and Configuration MANAGEMENT to meet service targets and standards
- Serve as technical lead during major incidents, performing deep-dive troubleshooting, RCA, and driving corrective/preventive actions
- Plan and execute complex changes with risk assessment, readiness checks, back-out planning, and post-change validation
- Maintain high-quality configuration and operational documentation; continuously improve runbooks/SOPs
- Operational Automation & AI-enabled Ops
- Design and build automation to improve BAU consistency and efficiency (e.g., config deployment/validation, compliance, health checks, reporting)
- Use Ansible, Python, and APIs to standardise operational workflows.
- Apply AIOps/AI-assisted operations (event correlation, anomaly detection, noise reduction, and predictive alerting) to reduce MTTR
- Scale automation patterns
- Security Operations & Zero Trust
- Operate and improve the security posture of DC network infrastructure
- Support micro-segmentation and Zero Trust-aligned controls within the data centre network
- Proactively identify resiliency and security gaps and drive remediation through engineering changes and automation
- Vendor & Service Delivery Management
- Collaborate with vendors and stakeholders to ensure service availability and timely issue resolution
- Track incidents/problems/tickets to ensure SLA adherence, quality updates, and timely closure
- Participate in service delivery reviews and drive actions that improve outcomes
- Continuous Improvement & Technology Refresh (Ops-led)
- Identify operational pain points and implement improvements across process, tooling, monitoring, and automation
- Participate in POCs/lab validations to improve operability, reliability, and automation
Requirements
- Bachelor's degree in Computer Sciences or related discipline
- 6+ years of hands-on network operations/engineering experience in a multi-vendor environment.
- Strong DC networking fundamentals (routing/switching, troubleshooting, complex change execution).
- Experience with operational practices: incident/problem/change management, RCA, and service reliability.
- Experience implementing process re-engineering and automation to improve processes.
- Strong problem-solving skills with the ability to isolate and resolve issues in complex environments.
- Excellent team player; ability to manage conflicts to achieve common goals.
- Strong communication skills; comfortable engaging technical and non-technical stakeholders.
- Experience in two or more of: core DC networking, security, network automation.
- Working knowledge of Ansible/Python/APIs (or motivation to build this capability quickly).
- Familiarity with segmentation/micro-segmentation and operational security hygiene.
- Certifications such as Cisco/F5/Palo Alto/Check Point (or equivalent).
- Awareness of AIOps/AI trends and ability to apply them pragmatically.
We thank all candidates for your interest in Singapore Airlines, and regret that only shortlisted candidates will be notified.