AI / Machine Learning Engineer (LLMs & RAG Pipelines)
AI / Machine Learning Engineer (LLMs & RAG Pipelines)
hyrezy tech solutions- Posted an hour ago
- Be among the first 10 applicants
Job Description
Location: Bangalore & Gurgaon - Hybrid (India)
Employment Type: Full-Time
About Us
We are actively building the next generation of intelligent product ecosystems. Our team doesn't just talk about artificial intelligence—we deploy production-grade LLM pipelines, vector search workflows, and scalable AI infrastructure into real-world applications every single week. If you are passionate about generative AI, low-latency inference, and turning cutting-edge research into reliable software, you will find your home here.
The Role
We are seeking an innovative AI / Machine Learning Engineer to take charge of our AI infrastructure. You will design, fine-tune, and scale artificial intelligence systems, integrating advanced LLM providers and vector databases into our core product architecture. You will work closely with our core engineering team to ensure our AI features are fast, cost-effective, and deeply integrated.
Key Responsibilities
Employment Type: Full-Time
About Us
We are actively building the next generation of intelligent product ecosystems. Our team doesn't just talk about artificial intelligence—we deploy production-grade LLM pipelines, vector search workflows, and scalable AI infrastructure into real-world applications every single week. If you are passionate about generative AI, low-latency inference, and turning cutting-edge research into reliable software, you will find your home here.
The Role
We are seeking an innovative AI / Machine Learning Engineer to take charge of our AI infrastructure. You will design, fine-tune, and scale artificial intelligence systems, integrating advanced LLM providers and vector databases into our core product architecture. You will work closely with our core engineering team to ensure our AI features are fast, cost-effective, and deeply integrated.
Key Responsibilities
- RAG Implementation: Architect and scale Retrieval-Augmented Generation (RAG) frameworks using vector databases like Qdrant or Pinecone.
- LLM & Prompt Optimization: Optimize prompt engineering workflows, manage token limits, and evaluate performance across various LLM providers (OpenAI, Anthropic, open-source models).
- Inference Engineering: Build and maintain low-latency, scalable AI inference endpoints using Python and FastAPI.
- Model Fine-Tuning: Fine-tune open-source models on domain-specific datasets to improve accuracy and business relevance.
- Pipeline Monitoring: Monitor AI system accuracy, latency, and cost efficiency, implementing caching strategies and fallback mechanisms.
- Experience: 3+ years of professional software/ML engineering experience with a strong focus on production-grade AI systems.
- Python Mastery: Deep expertise in Python, asynchronous programming, and modern frameworks like FastAPI.
- Vector & Database Skills: Hands-on experience with vector databases (Qdrant, Pinecone) and relational databases (MySQL/PostgreSQL).
- LLM Ecosystem Knowledge: Proven experience working with LangChain, LlamaIndex, Hugging Face transformers, and major LLM APIs.
- System Thinking: Strong grasp of API design, Docker containerization, and cloud deployment (AWS/GCP).
- Competitive Compensation: Top-tier salary bracket paired with performance bonuses and equity options.
- Cutting-Edge Stack: Freedom to experiment with and deploy the latest AI models and frameworks.
- Learning & Development: Dedicated stipend for AI conferences, research papers, and advanced certifications.
- Flexibility & Wellness: Remote-first culture and comprehensive health insurance for you and your family.
- Initial Screening: 30-minute introductory chat with our recruitment team.
- Technical Assessment: Practical take-home coding challenge or LLM integration task.
- Deep Dive & System Design: Architecture discussion focused on RAG pipelines, scaling, and past ML projects.
- Final Chat & Offer: Closing alignment with leadership and official offer rollout.
