

Search by job, company or skills

Company Description
AQUMEN Labs is a trusted Quality Engineering, DevOps, and AI Consulting partner to several industry leaders and high-growth startup companies across India and global markets. We actively engage with the startup ecosystem to help teams build resilient, scalable, and production-ready platforms by applying lessons learned from working with companies that have successfully scaled from early-stage ventures to global enterprises.
Our customers value the engineering-first, automation-driven approach that AQUMEN Labs brings to their technology organizations. By embedding modern quality practices, cloud-native DevOps, and AI-led insights into delivery pipelines, we help teams achieve faster releases, higher reliability, zero-touch deployments, and measurable business outcomes while remaining cost-efficient and operationally lean. Our clientele spans India, the USA, the UK, the Middle East, Australia, and Southeast Asia.
We thrive on solving complex engineering challenges across modern technology stacks, including cloud-native platforms, distributed microservices, CI/CD and GitOps, observability and SRE, AI/ML systems, data platforms, enterprise applications, and large-scale platform implementations. Our expertise extends across consumer-centric and regulated industries such as eCommerce, Payments & FinTech, Banking, Retail, Media, Healthcare, SaaS, and Digital Platforms.
At AQUMEN Labs, we combine quality engineering, DevOps excellence, and AI-driven intelligence to help organizations build software that is not only functional but reliable, scalable, and future-ready.
Role Description
This is a consulting, on-site role for a Site Reliability Engineer located in Bengaluru. The Site Reliability Engineer will be responsible for ensuring the reliability, scalability, and performance of the company's systems and services. Day-to-day responsibilities include designing and implementing reliable infrastructure, monitoring system performance, identifying and resolving incidents, and automating processes to minimise manual intervention. The engineer will collaborate with development and operations teams to optimise system performance while maintaining high availability and security.
Qualifications
Job ID: 139481913
Skills:
Unix, Elk, Prometheus, Bash, Grafana, Datadog, Gcp, Linux, Docker, Terraform, Ansible, Helm, Azure, Kubernetes, Python, AWS, GitOps, GitLab CI, GitHub Actions, OpenTelemetry, oci, ArgoCD
Skills:
PostgreSQL, Itil, Azure, Oracle, Kubernetes, Datadog, AWS, PagerDuty
Skills:
Gcp, Azure, Kubernetes, Logging, AWS, Alerting, tracing, Monitoring
Skills:
Java, Prometheus, Bash, Grafana, Google Cloud, Docker, Elastic Search, Splunk, Azure, Kubernetes, Python, AWS
Skills:
Tcp, UDP, Prometheus, Bash, Http, Grafana, Docker, Splunk, Python, Kubernetes, Open Telemetry, Ip