Search by job, company or skills

Fresher
Not Disclosed
Early Applicant
  • Posted 19 days ago
  • Be among the first 10 applicants

Job Description

Job Description

Senior DevOps Engineer
General Summary

As a Senior DevOps Engineer - Observability Platform , you will be responsible for building and maintaining scalable, reliable infrastructure and deployment pipelines with a strong emphasis on observability - metrics, logs, and traces - across systems running on Kubernetes and Azure. You will work closely with development teams to improve development velocity while ensuring system reliability, security, and performance. This role is critical in providing a standardized, observability platform that gives both internal engineering teams and external, customer-facing services deep, reliable visibility into system health, performance, and reliability

Minimum Qualifications
  • 3+ years of academic or work experience with Programming Language such as C, C++, Java, Python, etc.
  • Infrastructure Management: Design, implement, and maintain cloud-based infrastructure using Infrastructure as Code principles
  • Automation: Develop automation scripts and tools to streamline operations and eliminate manual processes
  • Containerization: Manage containerization strategies and orchestration using Docker and Kubernetes
  • Observability Platform: Design, build, and operate a standardized, self-service metrics, logs, and tracing platform (Grafana, Prometheus, Loki, OpenTelemetry) serving both internal teams and external, customer-facing services running on Kubernetes and Azure
  • Instrumentation & Telemetry: Partner with engineering teams to instrument applications and infrastructure, standardizing telemetry collection with OpenTelemetry
  • SLOs & Alerting: Define and maintain SLIs/SLOs and error budgets, build actionable dashboards, and tune alerting to maximize signal and reduce noise
  • Performance Optimization: Use observability data to analyze and optimize system performance, scalability, and cost-efficiency
  • Documentation: Create and maintain thorough documentation for infrastructure, deployment processes, and operational procedures
  • Incident Response & Escalation: Provide second-tier engineering escalation during business hours and own the telemetry, SLO, and alerting tooling that powers incident detection and reduces MTTD/MTTR front-line 24/7 on-call is owned by the dedicated SRE team, not observability engineers. Lead post-mortem analysis for observability-platform incidents

Requirements

Qualifications


  • 5+ years of experience in DevOps, Observability, or similar roles, including hands-on production experience operating Kubernetes based stack.
  • Advantage - Strong background in tool development with security focus
Technical Skills

  • Cloud Platforms: Extensive hands-on experience with Azure, including its observability services
  • Infrastructure as Code: Proficiency with Terraform, AWS CloudFormation, or similar IaC tools
  • Containerization: Advanced knowledge of Docker and Kubernetes ecosystem
  • Observability Stack: Hands-on experience with Grafana, Prometheus, Loki, Tempo or Jaeger, OpenTelemetry, and Alertmanager experience scaling metrics storage with Thanos, Mimir, or Cortex
  • Programming/Scripting: Strong coding skills in Python, Bash, or Go
Soft Skills
  • Problem-Solving: Excellent analytical and troubleshooting skills.
  • Communication: Strong verbal and written communication skills.
  • Collaboration: Ability to work effectively in a team environment and collaborate with cross-functional teams.
  • Leadership: Proven leadership skills and the ability to mentor junior engineers.

More Info

Job Type:
Employment Type:

Key Skills

Thanos

Loki

Alertmanager

OpenTelemetry

Infrastructure as Code

Mimir

Jaeger

Tempo

Similar Jobs

Bengaluru, India
Skills:
Automated Testing, Technology, Selenium-Python
India
Skills:
Microsoft Azure, Linux Administration, VMware, Ansible, Python Scripting, Windows Server Administration, Kubernetes, Terraform, Server infrastructure, Bicep, Storage Networking, Observability monitoring and alerting tools, DevOps practices, CI/CD tools
Bengaluru, India
Skills:
Git, Bash and or Python scripting, ITIL ServiceNow or Jira, VMware vSphere ESXi, Ubuntu Server administration, LVM multipathing iSCSI and Fibre Channel, Oracle PostgreSQL and middleware support, RHEL 7 8 9 administration, Nagios Zabbix or Prometheus Grafana, Kickstart and automated provisioning, DNS NTP and TCP IP administration, Red Hat Satellite administration
Gurugram, India, Gurugram
Skills:
threat management , Scripting, Linux System Administration, Databases, Docker, Elasticsearch, Sonarqube, Siem, programming, Python, Java, Vulnerability Management, Identity And Access Management, Bash, Iam, Distributed Systems, Kubernetes, Containers orchestration, OpenSearch, application-level threat detection, DevSecOps tools, Log ingestion pipelines, Cybersecurity products, NoSQL databases, Snyk, risk assessment methodologies
India
Skills:
PowerShell, Bash, Dns, Fibre Channel, Group Policy, DHCP, Microsoft Hyper-v, Fcoe, Powercli, Vcenter, Vmware Esxi, Python, Iscsi, virtual networking concepts, Active Directory, VMware vSAN