Search by job, company or skills

Senior AI Platform Operations Engineer

5-8 Years
SGD 0.84 - 1.44 LPA
  • Posted 11 days ago
  • Be among the first 10 applicants

Job Description

Make an Impact by:

Responsible for availability monitoring, outage detection, and performance optimization of our Azure AI cloud platform

Ensures continuous availability, resilience, and efficiency of the Red Hat OpenShift-powered RE:AI cloud platform.

Manage incident response, root cause analysis, and implement disaster recovery strategies to ensure business continuity

Oversee cybersecurity operations including vulnerability management, threat detection, and access control enforcement

Support security audits, compliance reporting, and ensure alignment with policies, regulatory frameworks and industry best practices

Collaborate with other developer teams to integrate monitoring, automation, and security best practices into AI/ML workflows

Drive continuous improvement in platform operations through automation, observability, and operational excellence initiatives

Skills for Success:

Bachelor's degree in Computer Science, Engineering, or a related field

4-6 years of experience in cloud administration and/or operations

Deep expertise in Azure operations and monitoring services including Azure Monitor, Log Analytics, Application Insights

Hands-on experience in observability and performance tuning for Red Hat OpenShift clusters, leveraging built-in monitoring and logging tools.

Strong background in incident management, SRE practices, and disaster recovery design

Hands-on experience with cloud security operations: IAM, SIEM/SOAR, vulnerability management, firewalls, endpoint detection

Proficiency in infrastructure-as-code (Terraform, Bicep, ARM) and automation scripting (PowerShell, Python)

Familiarity with AI/ML infrastructure (AKS, GPU VMs, data pipelines, model hosting) and their operational demands

Knowledge of security compliance frameworks (ISO 27001, CIS, NIST)

Excellent problem-solving, communication, and leadership skills, especially in high-pressure incident scenarios

Forward thinking ability to identify possible failure scenarios and formulate effective response plans

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151131381

Similar Jobs

Singapore, Cecil Street

Skills:

policies & proceduresInfrastructure as Code (IaC)OpenshiftData PipelineSecurity ComplianceHigh PressureThreat and Vulnerability ManagementDisaster Recovery Plansimplementing monitoring toolsDisaster Recovery ManagementAutomation Management in Product DevelopmentDeveloping SkillsOutage Management SystemsTeam MonitoringAzure MonitorSecurity Operations