AppOLens CoE Engineer
AppOLens CoE Engineer
ntt data north america- Posted 2 months ago
- Be among the first 10 applicants
Job Description
Job Description
We are seeking a Lead Integration & Observability Specialist to design, implement, and lead enterprise observability and reliability solutions, while supporting cloud-based integration platforms on AWS/Azure. The role focuses on monitoring, automation, and operational readiness of applications, APIs, data pipelines, and messaging systems.
This is a hands-on technical leadership role with mentoring and solution ownership responsibilities. The working environment includes Windows-based servers and .NET-based applications. Prior experience in Windows/.NET environments is preferred but not mandatory. The candidate should be a fast learner and willing to work across different technologies, platforms, and application environments.
Key Responsibilities
We are seeking a Lead Integration & Observability Specialist to design, implement, and lead enterprise observability and reliability solutions, while supporting cloud-based integration platforms on AWS/Azure. The role focuses on monitoring, automation, and operational readiness of applications, APIs, data pipelines, and messaging systems.
This is a hands-on technical leadership role with mentoring and solution ownership responsibilities. The working environment includes Windows-based servers and .NET-based applications. Prior experience in Windows/.NET environments is preferred but not mandatory. The candidate should be a fast learner and willing to work across different technologies, platforms, and application environments.
Key Responsibilities
- Lead the implementation of enterprise observability for applications, APIs, services, batch jobs, and data pipelines.
- Design and standardize monitoring, alerting, logging, metrics, and health checks across distributed systems.
- Integrate observability platforms with incident management and automation tools to support proactive issue detection and remediation.
- Support reliability and availability of integration platforms built on AWS/Azure.
- Perform advanced troubleshooting using logs, metrics, and traces to resolve production issues.
- Define operational readiness standards and non-functional requirements.
- Mentor engineers on observability best practices and platform usage.
- Collaborate with product, support, and operations teams to improve service stability and delivery.
- Work across different application environments, including Windows servers, .NET applications, cloud platforms, and integration/messaging systems.
- 7+ years of overall IT experience.
- 5+ years of relevant experience in Observability / Monitoring / Reliability Engineering.
- Strong hands-on experience with enterprise observability tools, such as IBM Instana, Dynatrace, AppDynamics, Prometheus, or Grafana.
- Expertise in monitoring and alerting design.
- Log management and analysis.
- Metrics and distributed tracing.
- Health checks and SLO/SLI concepts.
- Experience monitoring AWS/Azure workloads.
- Strong troubleshooting and incident analysis skills.
- Experience defining operational and non-functional requirements.
- Technical leadership and mentoring experience.
- Automation and ITSM integration, including ServiceNow workflows and incident automation.
- CI/CD and release management exposure.
- Cloud integration and messaging exposure.
- Experience working in Windows server environments.
- Experience supporting or monitoring .NET-based applications, IIS, Windows services, or related Microsoft technology platforms.
- Ability to quickly learn and adapt to new tools, platforms, and application environments.
- Willingness to work across various technology stacks, including Windows/.NET, cloud platforms, integration tools, messaging systems, and observability platforms.
- Exposure to enterprise production support, application reliability, and operational readiness practices.
More Info
Key Skills
Cloud integration and messaging
ServiceNow workflows
CI CD and release management
Health checks and SLO SLI concepts
Automation and ITSM integration
Metrics and distributed tracing
Log management and analysis
Monitoring and alerting design
Operational and non-functional requirements
Technical leadership and mentoring
Troubleshooting and incident analysis
IBM Instana
