- Posted 3 days ago
- Be among the first 10 applicants
Job Description
Agentic AI Cloud Engineering / MLOPS Engineer
Brief Description: This role will be responsible for developing agentic solutions from
prototype to production, combining LLMs, RAG, tool calling, orchestration frameworks,
cloud AI services, and modern software engineering practices.
• 4+ years of experience in AI engineering, software engineering, data engineering,
ML engineering, cloud engineering, or similar technical roles.
• Hands-on experience building GenAI applications, AI agents, RAG-based
solutions, enterprise search, copilots, or LLM-powered workflow
automation.
• Strong programming skills in Python, with experience building APIs, backend
services, automation scripts, and reusable AI components.
• Strong understanding of LLMs, including prompt engineering, context
engineering, model selection, temperature/top-p settings, context windows,
embeddings, token usage, latency, and cost trade-offs.
• Practical experience with RAG architecture, including vector databases,
embedding models, retrieval strategies, metadata filtering, document
processing, grounding, and citation-based answers.
• Hands-on experience with multi-agent orchestration patterns, including
supervisor-agent architectures, planner-executor workflows, routing agents,
tool-using agents, evaluator agents, and human-in-the-loop agent flows.
• Experience implementing tool-calling capabilities, allowing agents to interact
with databases, APIs, business applications, documents, and external services.
• Understanding of agent memory design, including session memory, long-term
memory, vector-based memory, user context, conversation history, and
governed memory retention.
• Experience implementing LLM and agent evaluation frameworks, including
accuracy testing, grounding validation, hallucination detection, retrieval quality
assessment, regression testing, adversarial testing, and user feedback
integration.
• Understanding of model governance and responsible AI, including approved
model usage, model selection criteria, evaluation evidence, security controls,
auditability, and lifecycle management.
• Experience implementing guardrails for AI agents, including policy-based
controls, restricted tool usage, approval gates, fallback flows, escalation paths,
human-in-the-loop checkpoints, and kill-switch mechanisms.
• Experience with observability and tracing for agentic systems, including
execution traces, tool-call monitoring, prompt/response metadata, token usage,
latency, error handling, fallback analysis, and production debugging of multistep workflows.
• Familiarity with agent development frameworks such as LangChain,
LangGraph, LlamaIndex, Semantic Kernel, CrewAI, AutoGen, or similar.
• Experience with cloud-native AI and agentic platforms such as AWS Bedrock
Agents, AWS Agent Core, AWS SageMaker, Azure OpenAI, Azure AI Agent
Service, Azure AI Foundry, Semantic Kernel, or equivalent technologies.
• Understanding of enterprise data concepts, including structured data,
unstructured data, semantic layers, data catalogues, metadata, data
quality, and governed access.
• Experience with REST APIs, microservices, authentication, secrets
management, logging, and cloud-native application patterns.
• Strong understanding of security and responsible AI principles, including rolebased access, data privacy, prompt injection risks, hallucination control,
content filtering, auditability, and safe agent execution.
• Ability to work with business stakeholders to understand use cases and translate
them into practical AI agent capabilities.
• Strong communication skills and ability to collaborate with architects, data
engineers, platform engineers, product owners, and business SMEs.
Nice to have:
• Experience with agent observability platforms or tracing tools for LLM
applications, including LangSmith, Arize Phoenix, OpenTelemetry-based tracing,
MLflow tracing, Databricks MLflow, cloud-native monitoring, or equivalent
solutions.
• Experience designing human-in-the-loop AI systems, including approval
workflows, exception management, escalation logic, user feedback capture, and
controlled autonomy.
• Experience with model risk management, responsible AI, AI governance
frameworks, prompt governance, model catalogues, evaluation reports, and
audit-ready documentation.
• Experience designing tool registries, plugin architectures, MCP-based
integrations, OpenAPI-based tools, schema-driven API invocation, and
reusable agent capabilities.
• Experience with advanced multi-agent topologies, including supervisor agents,
planner-executor agents, critic/evaluator agents, router agents, task-specific
specialist agents, and autonomous workflow coordination.
Experience designing tool registries and schema-driven integrations, using OpenAPI,
JSON Schema, structured outputs, function-calling definitions, API contracts, and
validation layers
More Info
Key Skills
LLM and agent evaluation frameworks
model governance
LLMs
multi-agent orchestration
agent memory design
tool-calling capabilities
agent development frameworks
GenAI applications
observability and tracing
guardrails for AI agents

