Search Jobs

Search by job, company or skills

AI Ops Engineer

AI Ops Engineer

D4insight Tech
Fresher
  • Posted 12 hours ago
  • Be among the first 10 applicants

Job Description

Location: Abu Dhabi, UAE
Experience: 7+ Years

Role Overview:

We are seeking for AI Ops Engineer to establish the operational backbone for enterprise AI platforms, enabling application teams to release AI products safely, repeatedly, and at scale.

The role is responsible for production release discipline, LLMOps practices, deployment automation, operational governance, cost visibility, and self-service operating standards for AI-native delivery teams.

Key Responsibilities:

AI Release & Deployment

  • Build standard release pipelines and promotion controls for AI applications, agents, platform services, and configuration changes.
  • Ensure repeatable deployments with appropriate governance and release evidence.
  • Implement controlled rollout, canary release, and rollback-readiness practices to reduce production risk.

LLMOps & Operational Governance

  • Embed operating controls for models and AI assets to support auditability and AI lifecycle management.
  • Integrate AI quality checks, operational telemetry, dashboards, runbooks, and production-readiness criteria into delivery processes.
  • Establish reusable operational standards and practices for AI-native teams.

Cost & Platform Visibility

  • Provide visibility into AI workload consumption, including:
    • Model usage
    • Token spend
    • Platform capacity
    • Quota management
    • Optimization opportunities
  • Convert proven operating patterns into reusable templates, release standards, onboarding guidance, and operational playbooks.

Required Skills & Experience:

  • 7+ years of relevant experience.
  • Strong production engineering background with cloud-native, AI, or high-scale API platforms in enterprise environments.
  • Hands-on experience with CI/CD, GitHub Actions or equivalent automation, and environment management.
  • Working knowledge of:
    • LLMOps
    • Telemetry
    • Release governance
    • Production readiness
  • Experience with observability, change control, service reliability, and continuous operational improvement.
  • Ability to collaborate with Platform Engineering, QA/SRE, Cybersecurity, Architecture, Product, and Delivery teams.

🩠Apply Now:

More Info

Key Skills

Service reliability

LLMOps

GitHub Actions

Release governance

High-scale API platforms

CI/CD

Observability

Continuous operational improvement

About Company

Similar Jobs

Bengaluru, India
Skills:
Java, Machine Learning, Artificial Intelligence, Api Development, MLops, Kubernetes, Python, Generative AI, Workflow Automation, Operational Intelligence Analytics, Vector Databases, Cloud Platforms, Intelligent Automation, Agentic AI, Observability Platforms