Search Jobs

Search by job, company or skills

AI Engineer (LLM & Agent Systems)

AI Engineer (LLM & Agent Systems)

speegile consulting
Early Applicant
  • Posted 3 days ago
  • Be among the first 20 applicants

Job Description

About the Role

We are looking for an experienced AI Engineer to design, build, and deploy LLM-driven agentic systems end-to-end — from data and model fine-tuning to retrieval, evaluation, guardrails, and production deployment. You should be equally comfortable reasoning about how a transformer works internally and shipping a reliable, observable agent system in production. This role owns the AI/ML depth of our stack, with support from backend and frontend engineers for integration.

Experience: 3+ Years

Qualification: B.E. / B.Tech / M.Sc. / M.Tech / MCA — Computer Science / IT or related field

Location: Preference Mumbai / Relocation to Mumbai / Remote

Employment Type: Full-Time

Notice Period: Max 30 Days

Must-Have Skills

•      3+ years of experience building Python-based ML/AI systems in production

•      Strong, practical understanding of how transformer, encoder-decoder, and deep learning models work internally — not just API-level usage

•      Hands-on experience fine-tuning LLMs (LoRA / QLoRA / full fine-tuning) and working with large-scale datasets end-to-end

•      Foundational Knowledge of LLMs and Can Lay the foundation for building a LLM model from scratch.

•      Solid grasp of the full LLM landscape — RAG, vectored and vector-less retrieval, evaluation, observability, and guardrails

•      Experience with voice/speech models (ASR/TTS & Translation) in Fine Tuning the models can even lay the foundation to build a voice model from scratch 

•      Knowledge graphs, Redis, Advanced RAG , Self corrective RAG

•      Experience with cloud AI deployments (AWS / GCP / Azure)

•      Hands-on experience with LLM agent frameworks and vector databases (ChromaDB, Weaviate, pgvector)

•      Strong knowledge of PyTorch or TensorFlow and Scikit-Learn

•      Production experience with FastAPI, Docker, and MLOps

Good-to-Have Skills

•      Advanced prompt optimization and agent evaluation techniques

•      Experience building AI observability tools from scratch

•      Exposure to multilingual or low-resource language models

•      Expert-level usage of agentic coding IDEs (Cursor, Windsurf, Claude Code)

Soft Skills

•      Strong problem-solving and system design mindset

•      Ability to work across AI, backend, and front-end teams

•      Comfortable mentoring freshers/junior engineers on AI fundamentals

•      Clear communication and documentation skills

•      Passion for building production-grade AI systems

What You'll Work On

LLM Agents & Prompt Engineering

•      Design and implement LLM agents using LangGraph, PydanticAI, and Google ADK

•      Build tool-augmented reasoning pipelines: RAG, Chain-of-Thought, ReAct, and planner–executor architectures

•      Develop robust, tested prompt strategies to improve reliability and reduce hallucination

Model Fundamentals & Fine-Tuning

•      Deep working knowledge of transformer architecture, Mamba architecture— encoder-only, decoder-only, and encoder–decoder models — and how attention, embeddings, and positional encoding actually work under the hood

•      Fine-tune LLMs using LoRA, QLoRA, and full fine-tuning depending on the use case and compute budget. If the fine tuning doesn't get the desired results, a LLM model will be built from scratch

•      Prepare, clean, and process large-scale datasets for pre-training/fine-tuning — deduplication, tokenization, sampling, and quality filtering at scale

•      Understanding of core deep learning model families (CNNs, RNNs/LSTMs, transformers) and when to use each

•      Work with voice/speech models — ASR (speech-to-text) and TTS (text-to-speech) — and understand how they integrate into conversational AI pipelines

Retrieval, RAG & Knowledge Systems

•      Implement both vectored (embedding-based) and vector-less (keyword/graph/hybrid) retrieval strategies, choosing the right approach per use case

•      Integrate vector databases (FAISS, Pinecone, pgvector, ChromaDB, Weaviate) and knowledge graphs (Neo4j)

•      Design chunking, embedding, and re-ranking strategies for high-precision retrieval

•      Implement Prompt Engineering, Context Engineering, Loop Engineering & Efficient Low token retrieval

Evaluation, Observability & Guardrails

•      Build offline and online evaluation harnesses for agent and LLM outputs (accuracy, groundedness, latency, cost)

•      Implement guardrails for safety, PII redaction, and scope control (e.g. NeMo Guardrails, Guardrails AI, or custom rule/LLM-based filters)

•      Build tooling for trace analysis, state debugging, and hallucination detection

•      Set up observability dashboards for LLM/agent systems (e.g. LangSmith, Arize, custom logging pipelines)

•      Benchmark agent orchestration frameworks for performance, cost, and reliability

Backend & MCP Integration

•      Build scalable APIs using FastAPI (sync & async execution)

•      Implement Model Context Protocol (MCP) for secure tool and data access

•      Manage agent state, context routing, and plugin-based workflows

MLOps & Deployment

•      Deploy and monitor models in cloud environments (AWS / GCP / Azure)

•      Work with model serving frameworks (e.g. vLLM, TGI) and apply quantization for efficient inference

•      Implement logging, observability dashboards, and automated recovery workflows

Front-End Collaboration

•      Build or collaborate on UI using React, TypeScript, or Next.js

•      Create seamless UI–API bridges for agent interactions and basic dashboards

More Info

Job Type:
Industry:
Employment Type:

Key Skills

cloud AI deployments

pgvector

LLM landscape

vector databases

voice speech models

Scikit-Learn

Python-based ML AI systems

ChromaDB

fine-tuning LLMs

large-scale datasets

LoRA

RAG

LLM agent frameworks

Weaviate

QLoRA

About Company