Search by job, company or skills

Site Reliability Engineer

Site Reliability Engineer

aviato consulting
5-7 Years
Not Disclosed
  • Posted 21 hours ago
  • Be among the first 10 applicants

Job Description


Aviato Consulting is seeking an experienced Site Reliability Engineer to join our growing team. This isn't just another SRE role; it's an opportunity to own critical infrastructure, drive technical strategy, and shape the reliability culture for major Australian and EU clients, all within a supportive, G-inspired environment built on transparency and collaboration.

What's In It For You

  • Learn from the Best: Report directly to and receive mentorship from our Head of SRE, an experienced ex-Google Manager.
  • High-Impact Projects: Take ownership of complex GCP environments for diverse, significant clients across Australia and the EU.
  • Drive Innovation, Not Just Tickets: Architect solutions and implement cutting-edge practices like SLOs, error budgets, and predictive AI anomaly detection to proactively improve systems.
  • A Culture That Works: Founded by ex-Googlers, we foster a transparent, collaborative, and low-bureaucracy environment.

What You'll Do (Your Impact):

  • Own & Architect Reliability: Design, implement, and manage highly available, scalable architectures on Google Cloud Platform (GCP).
  • Master Kubernetes & AI Infrastructure: Architect and manage production-grade Kubernetes clusters (GKE), and support high-performance infrastructure for AI/ML workloads (e.g., GPU/TPU pools).
  • Drive Automation & IaC: Lead robust automation strategies using Terraform, Ansible, and scripting (Python, Go, Bash) for CI/CD pipelines.
  • Elevate Observability with AIOps: Architect monitoring, logging, and alerting using Grafana, Dynatrace, and Sentry, integrating AI-driven insights for automated root-cause analysis.
  • Lead Incident Response: Spearhead incident management, conduct blameless post-mortems, and leverage GenAI to automate runbook generation.
  • Champion SRE Principles: Actively promote SLOs, SLIs, and error budgets, and mentor team members.

What You'll Bring (Your Expertise):

  • Proven SRE Experience: 5+ years of hands-on experience in a Site Reliability Engineering or Cloud Engineering role focusing on production systems.
  • Deep GCP & Kubernetes Knowledge: Demonstrable expertise in core GCP services and managing Kubernetes clusters in production (GKE highly desirable).
  • Infrastructure as Code Mastery: Significant experience using Terraform in complex environments.
  • Automation & Scripting Prowess: Strong proficiency in Python or Go for automating operational tasks.
  • Next-Gen Observability Expertise: Experience with modern monitoring tools, preferably with exposure to AI-driven alerting and logs clustering.
  • Problem-Solving Acumen: Strong analytical skills with experience leading incident response for critical systems.
  • (Desirable): Experience with AI infrastructure (vector databases, LLM inference), or API Management platforms like Apigee.

Technologies We Use (You'll Master):

  • Cloud: Google Cloud Platform (GCP)
  • Containerisation & Orchestration: Kubernetes (GKE), Docker
  • Infrastructure & Automation: Terraform, Ansible
  • Monitoring & AIOps: Grafana, Dynatrace, Sentry, Google Cloud Operations Suite
  • CI/CD: Jenkins, GitHub Actions, Bamboo (or similar)
  • Scripting & AI Tooling: Python, Go, Bash, GitHub Copilot, Gemini/Vertex AI
  • Collaboration: JIRA, Confluence, Slack

Work Timings : 5am to 2 pm IST

Ready to Elevate Your SRE Career

If you're a passionate Senior SRE ready to tackle complex challenges on GCP, work with leading clients, and benefit from exceptional mentorship in a fantastic culture, Aviato is the place for you. Apply now and help us build the future of reliable cloud infrastructure!

More Info

Job Type:
Industry:
Employment Type:

Key Skills

About Company

Similar Jobs

5-7 yrs
Chennai, India
Skills:
Restful Services, Java, Maven, Prometheus, Kafka, Grafana, Azure Sql, Google Cloud, Jenkins, Git, Javascript, Docker, Splunk, Azure, Kubernetes, Azure Cosmos, Mega cache
5-10 yrs
Gurugram, Gurugram, India
Skills:
DAST, Prometheus, Aem, Linux Scripting, Grafana, Datadog, DevSecOps, Terraform, Dynatrace, Kubernetes, AWS, CI CD, SCA, LLM-based solutions, OpenTelemetry, SAST, AIOps
6-8 yrs
Pune, India
Skills:
Terraform, Docker, AWS CloudFormation, Kubernetes, AWS, GitHub Actions, Microservices architectures, CI CD, AWS networking, EKS
5-7 yrs
India
Skills:
Datadog, Gcp, Terraform, Nginx, Haproxy, Python, Kubernetes, AWS, GitOps, HashiCorp Vault, FinOps, Go, Istio, Kargo, Argo CD, AWS Secrets Manager
8-10 yrs
Hyderabad, India
Skills:
data engineering , Java, Prometheus, Node.js, Grafana, Datadog, Sql, Gcp, Incident Management, Dynatrace, Splunk, Azure, Python, AWS, Site Reliability Engineering, Observability, Monitoring