Search Jobs

Search by job, company or skills

AI Engineer - Cloud & MLops

AI Engineer - Cloud & MLops

talentgigs
  • Posted 8 hours ago
  • Be among the first 10 applicants

Job Description

Agentic AI Cloud Engineering / MLOPS Engineer

Brief Description: This role will be responsible for developing agentic solutions from

prototype to production, combining LLMs, RAG, tool calling, orchestration frameworks,

cloud AI services, and modern software engineering practices.

• 4+ years of experience in AI engineering, software engineering, data engineering,

ML engineering, cloud engineering, or similar technical roles.

• Hands-on experience building GenAI applications, AI agents, RAG-based

solutions, enterprise search, copilots, or LLM-powered workflow

automation.

• Strong programming skills in Python, with experience building APIs, backend

services, automation scripts, and reusable AI components.

• Strong understanding of LLMs, including prompt engineering, context

engineering, model selection, temperature/top-p settings, context windows,

embeddings, token usage, latency, and cost trade-offs.

• Practical experience with RAG architecture, including vector databases,

embedding models, retrieval strategies, metadata filtering, document

processing, grounding, and citation-based answers.

• Hands-on experience with multi-agent orchestration patterns, including

supervisor-agent architectures, planner-executor workflows, routing agents,

tool-using agents, evaluator agents, and human-in-the-loop agent flows.

• Experience implementing tool-calling capabilities, allowing agents to interact

with databases, APIs, business applications, documents, and external services.

• Understanding of agent memory design, including session memory, long-term

memory, vector-based memory, user context, conversation history, and

governed memory retention.

• Experience implementing LLM and agent evaluation frameworks, including

accuracy testing, grounding validation, hallucination detection, retrieval quality

assessment, regression testing, adversarial testing, and user feedback

integration.

• Understanding of model governance and responsible AI, including approved

model usage, model selection criteria, evaluation evidence, security controls,

auditability, and lifecycle management.

• Experience implementing guardrails for AI agents, including policy-based

controls, restricted tool usage, approval gates, fallback flows, escalation paths,

human-in-the-loop checkpoints, and kill-switch mechanisms.

• Experience with observability and tracing for agentic systems, including

execution traces, tool-call monitoring, prompt/response metadata, token usage,

latency, error handling, fallback analysis, and production debugging of multistep workflows.

• Familiarity with agent development frameworks such as LangChain,

LangGraph, LlamaIndex, Semantic Kernel, CrewAI, AutoGen, or similar.

• Experience with cloud-native AI and agentic platforms such as AWS Bedrock

Agents, AWS Agent Core, AWS SageMaker, Azure OpenAI, Azure AI Agent

Service, Azure AI Foundry, Semantic Kernel, or equivalent technologies.

• Understanding of enterprise data concepts, including structured data,

unstructured data, semantic layers, data catalogues, metadata, data

quality, and governed access.

• Experience with REST APIs, microservices, authentication, secrets

management, logging, and cloud-native application patterns.

• Strong understanding of security and responsible AI principles, including rolebased access, data privacy, prompt injection risks, hallucination control,

content filtering, auditability, and safe agent execution.

• Ability to work with business stakeholders to understand use cases and translate

them into practical AI agent capabilities.

• Strong communication skills and ability to collaborate with architects, data

engineers, platform engineers, product owners, and business SMEs.

Nice to have:

• Experience with agent observability platforms or tracing tools for LLM

applications, including LangSmith, Arize Phoenix, OpenTelemetry-based tracing,

MLflow tracing, Databricks MLflow, cloud-native monitoring, or equivalent

solutions.

• Experience designing human-in-the-loop AI systems, including approval

workflows, exception management, escalation logic, user feedback capture, and

controlled autonomy.

• Experience with model risk management, responsible AI, AI governance

frameworks, prompt governance, model catalogues, evaluation reports, and

audit-ready documentation.

• Experience designing tool registries, plugin architectures, MCP-based

integrations, OpenAPI-based tools, schema-driven API invocation, and

reusable agent capabilities.

• Experience with advanced multi-agent topologies, including supervisor agents,

planner-executor agents, critic/evaluator agents, router agents, task-specific

specialist agents, and autonomous workflow coordination.

Experience designing tool registries and schema-driven integrations, using OpenAPI,

JSON Schema, structured outputs, function-calling definitions, API contracts, and

validation layers

More Info

Job Type:
Industry:
Employment Type:

Key Skills

LLM and agent evaluation frameworks

model governance

LLMs

multi-agent orchestration

agent memory design

enterprise data concepts

tool-calling capabilities

agent development frameworks

GenAI applications

observability and tracing

guardrails for AI agents

About Company

Similar Jobs

4-6 yrs
Hyderabad, India
Skills:
Apis, Microservices, Rest Apis, Python, LLM and agent evaluation frameworks, model governance, LLMs, multi-agent orchestration, agent memory design, tool-calling capabilities, agent development frameworks, GenAI applications, observability and tracing, guardrails for AI agents