Search by job, company or skills

5-7 Years
SGD 1.08 - 1.56 LPA
Early Applicant
  • Posted 2 days ago
  • Be among the first 10 applicants

Job Description

Responsibilities

  • Design, deploy and maintain backend infrastructure to ensure high availability, scalability and reliability of production systems.
  • Manage cloud infrastructure on AWS and AliCloud, including performance optimization, cost management and operational excellence.
  • Deploy, monitor and maintain Kubernetes clusters, distributed backend systems and supporting infrastructure services.
  • Build and maintain CI/CD pipelines and infrastructure automation using GitHub Actions, Terraform and Ansible.
  • Collaborate with software engineers to support backend service deployment, troubleshooting and production operations.
  • Investigate production incidents, perform root cause analysis and implement preventive improvements.
  • Implement infrastructure security best practices to ensure secure and reliable production environments.
  • Develop internal DevOps platforms and automation tools to improve engineering productivity and operational efficiency.
  • Implement monitoring, observability and alerting solutions to enhance service reliability.
  • Research and integrate AI technologies into infrastructure operations, including intelligent alert analysis, ChatOps and operational automation.


Qualifications:

  • 5+ years of hands-on experience in Kafka and Redis operations in large-scale production environments, be able to cooperate with developers to optimize code
  • Proficient in Python / Go / Java (at least one language) and SQL programming languages
  • Hands-on experience with containerization and orchestration (Docker, Kubernetes)
  • Strong experience with CI/CD tools such as GitHub Actions, Ansible, Terraform etc
  • At least 3 years of experience with AWS cloud platform. GCP, Azure, or Ali Cloud is a plus
  • Excellent problem-solving and troubleshooting skills
  • Strong team collaboration attitude and develop partnership with other teams and business
  • Practical experience building or operating AIOps systems (anomaly detection, alert correlation, automated healing, or RCA)
  • Familiarity with LLM-based DevOps automation (e.g., building chat-based ops assistants or AI-driven observability workflows)
  • Experience using or integrating tools like Dify, Agno, or LangChain into operational workflows

More Info

Job Type:
Industry:
Function:
Employment Type:

Job ID: 152213775

Similar Jobs

Remote

Skills:

Azure DevopsDevSecOpsKubernetesTerraformARM TemplatesService Bus

Remote

Skills:

DevopsAzureDatabricksTerraform

Remote

Skills:

AWSDockerKubernetes

Remote

Skills:

DevopsAppdynamicsAzureLinuxDynatraceSplunkAWSCI/CD

Beware of Scammers

We don’t charge money for job offers