

Search by job, company or skills

Role Overview:
We're looking for an AI Engineer for one of our Tier-1 IT clients with hands-on experience in building, fine-tuning, and optimizing LLM-based applications. The ideal candidate will have solid expertise in RAG (Retrieval-Augmented Generation) architectures, parameter-efficient fine-tuning (e.g., LoRA), and model quantization techniques for deployment efficiency.
Key Responsibilities:
Design, implement, and optimize end-to-end LLM-based solutions for real-world applications.
Develop and maintain RAG pipelines integrating vector databases, embeddings, and retrieval techniques.
Fine-tune pre-trained language models using LoRA or similar methods.
Apply quantization and optimization strategies to deploy models efficiently on constrained environments.
Collaborate with data scientists, software engineers, and product teams to integrate AI features into production systems.
Monitor, evaluate, and continuously improve model performance and reliability.
Required Skills:
3–5 years of experience in AI/ML development or applied NLP.
Proficient in Python and frameworks such as PyTorch or TensorFlow.
Strong understanding of LLM architectures (e.g., GPT, Llama, Falcon, Mistral).
Experience with RAG frameworks (LangChain, LlamaIndex, or custom retrieval setups).
Hands-on knowledge of LoRA, PEFT, and model quantization (GPTQ, AWQ, or similar).
Familiarity with vector databases like FAISS, Pinecone, or ChromaDB.
Good understanding of prompt engineering and evaluation techniques.
Cloud deployment experience (AWS, Azure, or GCP) is an advantage.
Preferred Skills:
Exposure to opensource models and fine-tuning pipelines.
Experience integrating AI models into web or enterprise products.
Knowledge of containerization and MLOps (Docker, Kubernetes, MLflow).
Job ID: 141163497
Skills:
Retrieval-Augmented Generation (RAG), Infrastructure as Code (IaC), Amazon Web Services, Google Cloud Platform, Terraform, Azure, LangChain, Regula test cases, RAG vector databases, Cloud infrastructure deployment pipelines, Vector databases, Rego policies, LangGraph, LLM APIs, Open Policy Agent
Skills:
PowerShell, Bash, Python, Azure Container Registry ACR, Azure Data Lake Storage ADLS Gen2, Azure Key Vault, Azure DevOps Pipelines and GitHub Actions, Docker and Helm, Azure SQL Database, Azure Kubernetes Service AKS, Azure Data Factory and Databricks, Azure Monitor and Log Analytics, Azure App Service Virtual Machines VNet, YAML-based CI CD automation
Skills:
Apis, automated testing practices, GenAI, AI services, application components, secure coding practices, prompt engineering frameworks, LLM integrations, RAG pipelines, inference services, dependency management
Skills:
Github, FastAPI, Python, LangChain, pgvector, AI-native development tools, retrieval-augmented generation, vector databases, Pinecone, Claude, prompt engineering, LangGraph, LLM APIs, OpenAI, Weaviate
Skills:
Vue, React, Opencv, FastAPI, Rest Apis, Python, TensorRT, NVIDIA workstation configurations, GRPC, ONNX Runtime, 3D vision point cloud processing, Streamlit