Job Description
Role Summary
As Observability Engineer / SME, you will be the technical authority for designing, implementing, and operating enterprise monitoring and observability solutions across hybrid IT environments. Working hands-on with platforms such as HPE OpsRamp and SolarWinds (and similar ITOM/observability tools), you will deliver end-to-end monitoring — onboarding, discovery, service mapping, alert correlation, and ITSM-integrated event management — to improve reliability, reduce noise, and accelerate MTTD/MTTR for our clients. You will also mentor L1/L2 teams and serve as the central point of expertise for monitoring delivery.
Key Responsibilities
Design & Implementation
- Design, deploy, configure, customize, and integrate observability solutions using HPE OpsRamp, SolarWinds, and similar platforms (Nagios, Zabbix, ManageEngine OpManager, etc.) across on-prem, cloud, and edge infrastructure.
- Analyze client infrastructure design, architecture, and security standards to define fit-for-purpose monitoring solutions.
- Develop and maintain architecture blueprints, design documents, SOPs, and operational runbooks for monitoring and event-management systems.
- Lead onboarding, discovery, service mapping, alert correlation, and runbook automation for green-field and brown-field engagements.
Operations & Event Management
- Own day-to-day operations of observability platforms — ensuring high availability, performance, capacity, data accuracy, and timely upgrades/patching.
- Manage event-management processes: event ingestion, correlation, noise reduction, deduplication, and routing to the appropriate resolver groups.
- Troubleshoot complex incidents escalated from L1/L2 teams; lead root-cause analysis (RCA) and turn telemetry into actionable reliability insights.
- Establish and continuously improve standards, thresholds, and SLA-based response models; maintain telemetry data hygiene (retention, indexing, access controls).
- Integrate monitoring tools with ITSM platforms (ServiceNow, BMC, Ivanti) and cloud platforms (AWS, Azure, GCP) — including alert/ticketing workflows via MID server/connectors.
Automation, Collaboration & Leadership
- Build automation scripts and self-healing runbooks (Python/Bash/PowerShell) to improve operational efficiency and reduce manual toil.
- Coordinate with service desk, network, system, and application teams to drive proactive incident prevention using observability insights.
- Conduct knowledge transfer and training for L1/L2 teams; develop upskilling roadmaps and mentor junior engineers.
- Deliver executive reports, performance dashboards, and service-improvement recommendations to client and internal stakeholders.
Required Skills & Experience
Monitoring & Observability (mandatory)
- 6–8 years of hands-on experience in IT monitoring tool deployment, implementation, operational support, and report building.
- Strong hands-on expertise with HPE OpsRamp and SolarWinds and their modules; experience with similar tools (ManageEngine OpManager, Nagios, Zabbix, PRTG) is an advantage.
- Working knowledge of broader observability/APM platforms — Dynatrace, AppDynamics, Datadog, Prometheus, Grafana, ELK, Splunk — and OpenTelemetry concepts.
- Proven experience in Managed Services / NOC delivery and ITSM-integrated event management.
Infrastructure, Cloud & Automation
- Strong understanding of servers, networks, virtualization, storage, and hybrid cloud environments.
- Hands-on integration experience with ITSM tools (ServiceNow preferred, BMC, Ivanti) and cloud platforms (AWS, Azure, GCP).
- Scripting for automation and integration — Python, Bash, or PowerShell; familiarity with containerization (Docker, Kubernetes) is a plus.
- Sound understanding of ITIL processes — Incident, Problem, Change, and Configuration Management.
Soft Skills
- Strong analytical and problem-solving ability with high attention to detail.
- Excellent communication, documentation, and reporting skills (design docs, SOPs, RCA, dashboards).
- Proven mentoring and cross-team collaboration; self-motivated and able to act as a technical authority in a fast-paced environment.
Qualifications
- Bachelor's degree in Computer Science, Information Technology, or a related field.
- ITIL Foundation certification (Intermediate/Expert preferred).
- Certifications in HPE OpsRamp/GreenLake/OneView, SolarWinds, or related ITOM/observability and cloud platforms are an advantage; AIOps/AI-related training is a plus.
More Info
Key Skills
ManageEngine OpManager
Ivanti
OpenTelemetry
HPE OpsRamp
