Monitoring Engineer
Monitoring Engineer
movado group, inc- Posted 4 hours ago
- Be among the first 10 applicants
Job Description
The Global Monitoring & Operations Engineer/Analyst is responsible for real-time monitoring, incident detection, triage, and coordination across the enterprise application and infrastructure landscape. This role operates within a 24×7 Global Command Center, ensuring high availability, performance, and reliability of business-critical systems including eCommerce, POS, ERP (SAP), and integrations.
The role plays a critical part in proactive issue detection, rapid incident response, and seamless escalation to L2/L3 teams, supporting a follow-the-sun operational model. Focus is on monitoring, triage, and coordination — not deep engineering troubleshooting or root cause ownership.
Required Skills & Experience
Within 6–12 months, the successful candidate will:
The role plays a critical part in proactive issue detection, rapid incident response, and seamless escalation to L2/L3 teams, supporting a follow-the-sun operational model. Focus is on monitoring, triage, and coordination — not deep engineering troubleshooting or root cause ownership.
- Real-Time Monitoring & Event Management
- Monitor enterprise systems, applications, integrations, and infrastructure using centralized monitoring tools
- Track health, availability, and performance of:
- DTC and B2B eCommerce platforms and websites
- ERP systems (SAP ECC, BW, etc.)
- POS systems and retail operations
- Integration platforms (APIs, EDI, middleware)
- Cloud Infrastructure and Network Infrastructure
- Identify anomalies, alerts, and potential incidents in real time
- Ensure all critical alerts are acknowledged within defined SLAs
- Incident Detection & Triage (L1 Operations)
- Perform first-level triage and validation of alerts
- Classify incidents based on severity, business impact, and urgency
- Execute predefined runbooks and standard operating procedures (SOPs)
- Resolve common and repeatable issues where possible
- Document diagnostic findings clearly to accelerate L2/L3 resolution
- Escalation & Coordination
- Escalate incidents to L2 Application Support or L3 Product/Engineering teams as required
- Ensure proper documentation and context is provided during escalation
- Act as the central coordination point during incidents, ensuring communication across teams
- Incident Management & Communication
- Create and manage incident tickets in ITSM tools (e.g., ServiceNow)
- Maintain accurate and timely updates throughout the incident lifecycle
- Communicate status updates to stakeholders, including business users when required
- Support major incident coordination by providing timely updates, documentation, and bridge call support as directed by the Incident Manager
- Proactive Monitoring & Prevention
- Identify recurring issues and trends
- Recommend improvements to monitoring thresholds, alerts, and automation
- Provide incident timelines, logs, and supporting data to assist problem management and RCA activities led by L2/L3 teams
- Automation & Continuous Improvement
- Collaborate with L2/L3 teams to:
- Enhance monitoring coverage and observability
- Support automation and automate repetitive tasks where possible
- Improve alert quality (reduce noise, eliminate false positives)
- Maintain and update runbooks and knowledge base articles / documentation
- Global Operations Model
- Participate in a 24×7 shift-based model (follow-the-sun across regions)
- Ensure seamless handover between shifts and regions
- Maintain operational discipline and adherence to SLAs and OLAs
Required Skills & Experience
- Experience with monitoring and observability tools (e.g., Dynatrace, Datadog, Splunk, New Relic, AppDynamics, Zabbix, Nagios, PRTG etc.,)
- Understanding of:
- eCommerce, Point of Sale, and B2B Web and platforms
- ERP systems (preferably SAP)
- API integrations and middleware (e.g., Boomi)
- Cloud and Network Infrastructure
- Database system performance indicators
- Familiarity with cloud environments and distributed systems
- Basic knowledge of networking, databases, and system performance
- Ability to interpret logs, dashboards, and monitoring alerts
- Strong incident management and troubleshooting skills
- Experience working in 24×7 operations or NOC environments
- Knowledge of ITIL processes (Incident, Problem, Change Management)
- Ability to follow and execute structured runbooks and escalation matrices
- Strong analytical and problem-solving ability
- Clear and concise communication (verbal and written)
- Ability to perform under pressure during critical incidents
- High attention to detail and operational discipline
- Experience in retail, eCommerce, or global enterprise environments
- Exposure to e Commerce, POS, and SAP ecosystems
- Familiarity with automation/scripting (Python, Shell, etc.)
- ITIL certification (preferred)
- 24×7 shift-based operations (rotational shifts)
- Follow-the-sun support across APAC, EMEA, and AMER
- Handover discipline between shifts is critical
Within 6–12 months, the successful candidate will:
- Consistently meet SLA targets
- Accurately classify and escalate incidents with minimal rework
- Demonstrate strong shift handover discipline
- Reduce false escalations through effective triage
- Contribute to improved alert quality and monitoring coverage
More Info
Key Skills
logs
Database system performance indicators
API integrations
ERP systems
monitoring alerts
cloud environments
monitoring and observability tools


