Project Role : Operations Engineer
Project Role Description : Support the operations and/or manage delivery for production systems and services based on operational requirements and service agreement.
Must have skills : Cloud Infrastructure
Good to have skills : NA
Minimum 5 Year(s) Of Experience Is Required
Educational Qualification : 15 years full time education
Summary:
We re looking for a Infrastructure Site Reliability Engineer who combines software development expertise with strong DevOps and cloud infrastructure skills. You ll be responsible for designing, developing, deploying, and maintaining scalable web applications while implementing CI/CD pipelines, automating infrastructure, and ensuring high system reliability and performance.
You ll work closely with developers, QA engineers, and operations teams to streamline deployment processes and enhance development workflows.
Roles & Responsibilities:
IaC Development
Design, develop, and maintain Infrastructure Automation using IaC principals and DevOps Skills.
Develop Predictive automation using Monitoring and Observability tools.
DevOps & Cloud Infrastructure
Build and maintain CI/CD pipelines using tools such as GitHub Actions, GitLab CI, Jenkins, or CircleCI.
Manage infrastructure as code (IaC) using Terraform, CloudFormation, or Ansible, Pulumi and CrossPlane.
Deploy and monitor applications in cloud environments (AWS, Azure, or GCP).
Set up automated monitoring, logging, and alerting (Prometheus, Grafana, ELK Stack, Datadog).
Enhance system reliability, security, and scalability through automation and best practices.
Operations & Collaboration
Participate in code reviews and DevOps strategy discussions.
Troubleshoot and resolve production issues with a focus on uptime and performance.
Implement security best practices and assist with compliance standards.
Mentor junior developers and contribute to improving development workflows.
Professional & Technical Skills:
- DevOps & Automation Tools
CI/CD Pipelines:
Jenkins, GitHub Actions, GitLab CI, CircleCI, Travis CI, ArgoCD
Infrastructure as Code (IaC):
Terraform, Ansible, AWS CloudFormation, Pulumi
Configuration Management:
Chef, Puppet, Ansible, SaltStack
Monitoring & observability
Splunk, DataDog, Prometheus.
Cloud Platforms:
AWS, Azure, Google Cloud Platform (GCP)
Compute & Networking:
EC2, Lambda, VPC, Load Balancers, Route 53, DNS, VPNs
Storage & Databases:
S3, RDS, DynamoDB, PostgreSQL, MySQL, MongoDB
Containerization:
Docker, Podman
Container Orchestration:
Kubernetes (K8s), Helm, OpenShift
- System Administration & Networking
Operating Systems:
Linux (Ubuntu, CentOS, Alpine), Bash scripting, basic Windows Server knowledge
Networking Fundamentals:
TCP/IP, DNS, SSL/TLS, HTTP/HTTPS, VPN, firewalls
Security & Compliance:
IAM, Secrets Management (Vault, AWS Secrets Manager), SSH, encryption, vulnerability scanning
- Observability & Incident Response
Monitoring:
Application and system performance metrics
Alerting & Incident Management:
PagerDuty, Opsgenie, Slack integrations
Logging:
Centralized log management (ELK, Fluentd, Splunk)
Strong problem-solving and debugging skills
Collaboration with developers, QA, and IT teams
Understanding agile methodologies (Scrum, Kanban)
Continuous learning and adapting to new tools and cloud services
Bonus / Emerging Skills
Serverless architectures (AWS Lambda, Azure Functions)
GitOps (Flux, ArgoCD)
AIOps and observability automation
SRE principles (Service Level Indicators/Objectives)
Additional Information:
- The candidate should have minimum 7.5 years of experience in Cloud Infrastructure.
- This position is based at our Bengaluru office.
- A 15 years full time education is required.