Search by job, company or skills

Monitoring and Observability Consultant

6-11 Years
Early Applicant
Quick Apply
  • Posted a month ago
  • Be among the first 50 applicants

Job Description

  • Design end-to-end monitoring and observability solutions to provide comprehensive visibility into infrastructure, applications, and networks.
  • Implement monitoring tools and frameworks (e.g., Prometheus, Grafana, OpsRamp, Dynatrace, New Relic) to track key performance indicators and system health metrics.
  • Integration of monitoring and observability solutions with IT Service Management Tools.
  • Develop and deploy dashboards, alerts, and reports to proactively identify and address system performance issues.
  • Architect scalable observability solutions to support hybrid and multi-cloud environments.
  • Collaborate with infrastructure, development, and DevOps teams to ensure seamless integration of monitoring systems into CI/CD pipelines.
  • Continuously optimize monitoring configurations and thresholds to minimize noise and improve incident detection accuracy.
  • Automate alerting, remediation, and reporting processes to enhance operational efficiency.
  • Utilize AIOps and machine learning capabilities for intelligent incident management and predictive analytics.
  • Work closely with business stakeholders to define monitoring requirements and success metrics.
  • Document monitoring architectures, configurations, and operational procedures.

Required Skills:

  • Strong understanding of infrastructure and platform development principles and experience with programming languages such as Python, Ansible, for developing custom scripts.
  • Strong knowledge of monitoring frameworks, logging systems (ELK stack, Fluentd), and tracing tools (Jaeger, Zipkin) along with the OpenSource solutions like Prometheus, Grafana.
  • Extensive experience with monitoring and observability solutions such as OpsRamp, Dynatrace, New Relic, must have worked with ITSM integration (e.g. integration with ServiceNow, BMC remedy, etc.)
  • Working experience with RESTful APIs and understanding of API integration with the monitoring tools.
  • Familiarity with AIOps and machine learning techniques for anomaly detection and incident prediction.
  • Knowledge of ITIL processes and Service Management frameworks.
  • Familiarity with security monitoring and compliance requirements.
  • Excellent analytical and problem-solving skills, ability to debug and troubleshoot complex automation issues

More Info

Job Type:
Industry:
Function:
Employment Type:

About Company

Job ID: 108816707

Beware of Scammers

We don’t charge money for job offers