AI Agent Operations Engineer
AI Agent Operations Engineer
umanist na- Posted 11 days ago
- Be among the first 10 applicants
Job Description
Location: Pune(572)
Experience: 8–11 Years
Notice Period: Immediate–45 Days
Work Type: Client-facing production operations
Interview: 2 Technical Rounds + Client Round + HR
Preference: Local Pune candidates preferred
Role Overview
We are looking for an AI Agent Operations Engineer to lead production incident response, proactive monitoring, service reliability, and operational governance for AI-agent systems.
The role requires strong expertise in incident management, observability, SLOs, runbook operations, AI-agent failure analysis, and stakeholder communication.
Key Responsibilities
Experience: 8–11 Years
Notice Period: Immediate–45 Days
Work Type: Client-facing production operations
Interview: 2 Technical Rounds + Client Round + HR
Preference: Local Pune candidates preferred
Role Overview
We are looking for an AI Agent Operations Engineer to lead production incident response, proactive monitoring, service reliability, and operational governance for AI-agent systems.
The role requires strong expertise in incident management, observability, SLOs, runbook operations, AI-agent failure analysis, and stakeholder communication.
Key Responsibilities
- Act as Incident Commander during production incidents; lead bridges, escalations, communication, and resolution.
- Conduct daily proactive trend reviews to identify degradation before alerts are triggered.
- Coordinate rollback decisions for degraded releases and activate fallback processes when required.
- Tune monitoring safeguards and alert thresholds to improve alert precision and reduce noise.
- Lead post-incident reviews, track corrective actions, and maintain runbooks.
- Analyze logs, traces, and monitoring data to identify production issues and trends.
- Maintain shift-handover standards and mentor team members when required.
- Communicate technical risks, incidents, and solutions clearly to non-technical stakeholders.
- Bachelor's/Master's degree in Computer Science or related field
- relevant 6+ years in production operations
- Experience in customer-facing production environments
- Strong monitoring, observability, log & trace analysis
- Hands-on incident management, SLOs & runbook operations
- Understanding of AI-agent failure modes and routing
- Strong analytical and pragmatic approach
- Excellent communication & stakeholder management
- Minimum 2 years stability per organization
- AI/ML Operations exposure
- Scripting/automation for operations
- LLM / AI-agent tooling
- Management-level presentations
- Candidate will be required to travel to Singapore within 30–45 days of onboarding, subject to paperwork completion.
- Additional SGD 3,500/month will be provided for accommodation and other expenses.
- Strong communication skills are mandatory due to the client-facing nature of the role.
- Immediate to 45 days notice period preferred.
More Info
Key Skills
customer-facing production environments
log trace analysis
SLOs
scripting automation for operations
AI-agent failure analysis
strong analytical and pragmatic approach
runbook operations
LLM AI-agent tooling
observability
