Stack Management : Deploy, configure, and maintain the core observability stack using Prometheus, Grafana, Alertmanager and Loki.
Dashboarding & Visualization : Collaborate with engineering teams to design and build comprehensive Grafana dashboards for application and infrastructure health monitoring.
Alerting Strategy : Configure and fine-tune Alertmanager rules to ensure accurate, actionable alerts while minimizing alert fatigue.
Log Management : Architect and manage centralized logging solutions using Loki to ensure efficient log aggregation and querying.
System Optimization : Monitor the performance of the observability stack itself, optimizing resource usage and scaling infrastructure as needed.
Continuous Improvement : Evaluate and integrate modern observability tools (such as VictoriaMetrics and VictoriaLogs) to enhance system correlation and analysis capabilities.
Experience : 9+ years of hands-on experience in DevOps, SRE or Observability roles.
Core Stack : Deep technical expertise in Prometheus, Grafana, Alertmanager, and Loki (PLG Stack).
Automation : Proficiency in scripting (Python, Bash) and Infrastructure as Code (e.g., Terraform, Ansible).
Infrastructure & OS : Strong working knowledge of Linux/Unix administration.