
Search by job, company or skills
Responsible for availability monitoring, outage detection, and performance optimization of our Azure AI cloud platform
Ensures continuous availability, resilience, and efficiency of the Red Hat OpenShift-powered RE:AI cloud platform.
Manage incident response, root cause analysis, and implement disaster recovery strategies to ensure business continuity
Oversee cybersecurity operations including vulnerability management, threat detection, and access control enforcement
Support security audits, compliance reporting, and ensure alignment with Singtel policies, regulatory frameworks and industry best practices
Collaborate with other developer teams to integrate monitoring, automation, and security best practices into AI/ML workflows
Drive continuous improvement in platform operations through automation, observability, and operational excellence initiatives
Skills for Success:
Bachelor's degree in Computer Science, Engineering, or a related field
4-6 years of experience in cloud administration and/or operations
Deep expertise in Azure operations and monitoring services including Azure Monitor, Log Analytics, Application Insights
Hands-on experience in observability and performance tuning for Red Hat OpenShift clusters, leveraging built-in monitoring and logging tools.
Strong background in incident management, SRE practices, and disaster recovery design
Hands-on experience with cloud security operations: IAM, SIEM/SOAR, vulnerability management, firewalls, endpoint detection
Proficiency in infrastructure-as-code (Terraform, Bicep, ARM) and automation scripting (PowerShell, Python)
Familiarity with AI/ML infrastructure (AKS, GPU VMs, data pipelines, model hosting) and their operational demands
Knowledge of security compliance frameworks (ISO 27001, CIS, NIST)
Excellent problem-solving, communication, and leadership skills, especially in high-pressure incident scenarios
Forward thinking ability to identify possible failure scenarios and formulate effective response plans
Job ID: 151520549
Skills:
Impact Analysis, Openshift, Cybersecurity, Automation Scripting, Log Analysis, Security Compliance, Computer Engineering, Regulatory Frameworks, Disaster Recovery Plans, implementing monitoring tools, Disaster Recovery Management, Hosting Services, Automation Management in Product Development, Outage Management Systems, Team Monitoring, Security Operations