Search by job, company or skills

Senior Machine Learning Engineer

5-7 Years
  • Posted a day ago
  • Be among the first 10 applicants

Job Description

Role: Senior AI/ML Engineer – Agentic AI & LLM Systems

Function: Artificial Intelligence / Machine Learning Engineering

Location: India (Remote)

Type: Full-time

Industry: Information Technology & Services, Management Consulting

About Company

The company is a digital engineering firm founded in 2020, headquartered in Tampa, Florida. It specializes in AI-driven digital transformation for enterprises.

With 450+ professionals across seven global offices, it operates in 25+ countries. The company has completed 55+ engagements spanning platform engineering, data analytics, quality engineering, and supply chain transformation.

In 2023, it expanded its product engineering capabilities through an acquisition. The culture is ownership-driven, fast-paced, and built for engineers who want their work to ship and matter.

Position Overview

This role sits at the technical core of the company's AI practice. The engineer will own the architecture and hands-on development of enterprise-scale agentic AI systems, advanced RAG pipelines, and multi-agent orchestration platforms built to handle millions of daily interactions. The role also involves driving LLM selection, cost optimization, and observability standards, while mentoring a team of AI/ML engineers. The work spans the full stack — from vector retrieval and knowledge graphs to containerized cloud deployment and CI/CD for AI systems.

Role & Responsibilities

  • Architect end-to-end agentic AI platforms supporting millions of daily interactions with sub-100ms latency, including multi-agent orchestration frameworks coordinating 15+ specialized agents with state persistence, memory management, and error recovery
  • Design and implement advanced RAG systems combining vector search, semantic search, knowledge graphs (Neo4j), and hybrid retrieval strategies using Pinecone, Weaviate, Milvus, Chroma, and FAISS
  • Build production agentic AI applications using LangChain, LangGraph, CrewAI, AutoGen, and Azure ADK with ReAct patterns, Chain-of-Thought reasoning, and advanced function calling
  • Develop and maintain MCP (Model Context Protocol) servers for tool standardization and agent-to-system communication; implement prompt optimization, semantic caching, and token-level cost reduction strategies
  • Establish evaluation frameworks using RAGAS and custom metrics; implement observability via LangSmith and OpenTelemetry for hallucination tracking, retrieval quality, and model performance monitoring in production
  • Deploy and manage containerized AI workloads on Kubernetes and Docker across AWS, Azure, or GCP; design CI/CD pipelines with automated testing and canary releases for AI model deployment
  • Mentor junior and mid-level AI engineers, lead architectural reviews, and define technical standards for AI system reliability, observability, and cost governance

Must Have Criteria

  • 5+ years of backend software development with Python required; Java, Node.js, Go, or Rust acceptable as a secondary language — with strong REST API and microservices experience
  • 2+ years of hands-on production experience building LLM and agentic AI systems at scale using LangChain and LangGraph
  • Production experience with at least one agentic AI framework: CrewAI, AutoGen, or Azure ADK
  • Hands-on experience with vector databases (Pinecone, Weaviate, Milvus, Chroma, or FAISS) and RAG architecture design including hybrid retrieval strategies
  • Experience working with multiple LLM providers — OpenAI GPT-4, Anthropic Claude, Meta Llama, or Google Gemini — including model routing and trade-off evaluation
  • Advanced Docker and Kubernetes expertise with production deployment experience on AWS, Azure, or GCP
  • Experience with LLM observability platforms (LangSmith or equivalent) and OpenTelemetry-based distributed tracing

Nice to Have

  • Experience developing Model Context Protocol (MCP) servers and familiarity with emerging AI tooling standards
  • Advanced LLM knowledge: fine-tuning, quantization, or knowledge distillation at scale
  • Experience building knowledge graphs and graph-based retrieval systems using Neo4j or GraphQL
  • Expertise in multimodal AI systems (text, image, or audio processing pipelines)
  • Contributions to open-source AI/ML projects, published research, or AWS/Azure/GCP cloud certifications
  • Track record of MLOps implementation and cost optimization at enterprise scale

What We Offer

  • Technical ownership of cutting-edge agentic AI and LLM systems with direct impact on millions of end users
  • Access to the latest LLM models, AI research, and a professional development budget covering conferences and certifications
  • Clear career trajectory toward Staff Engineer or Technical Leadership roles within a fast-growing global AI practice
  • Collaborative engineering culture with a strong emphasis on ownership, speed, and delivering measurable client outcomes

About Company

Job ID: 151545677

Similar Jobs

Chennai, India

Skills:

Machine LearningKafkaTensorflowSklearnData ScienceDockerFlaskPythonJavaScalaRedshiftSqlRedisRabbitmqJenkinsPytorchSparkMongoDBFastAPIKubernetesAirflowSagemakerMLFlowMilvus

India

Skills:

NetworkingDistributed SystemsIncident ManagementKubernetesPythonML-specific failuresDiffusersTorchObservability

Bengaluru, India

Skills:

AWSComputer VisionJavaPythonAzureGcpenterprise-grade productsGocloud platformsapplied ML researchLLM fine-tuningreinforcement learninglarge-scale distributed systems

Bengaluru, India

Skills:

Machine LearningCloud ArchitectureSqlDeep LearningMLopsDatabricksPythonLLMsRAG systemsGraph ModellingMLFlowAI features for model inferenceEvaluation frameworksIdentity Access Graph

Bengaluru, India

Skills:

TensorflowPytorchKerasPythonAWSLangChainLangGraphAutogenTensorFlow ServingHugging FaceSpacy