DATA SCIENTIST
Location: Pune, India | Type: Full-Time | Experience: 4-8 Years
ROLE SUMMARY
Enterprise Data Scientist designing, building, and productionizing analytics and GenAI solutions at scale. Develop, deploy, and operate machine learning and applied GenAI models—including LLM-based insight generation, summarization, and decision augmentation—using large-scale structured and semi-structured data. Strong emphasis on scalability, reliability, governance, and enterprise-ready implementations.
KEY RESPONSIBILITIES
- Translate complex business problems into data science and machine learning solutions that drive measurable outcomes across enterprise use cases
- Perform advanced data exploration, feature engineering, model development, and evaluation on large-scale structured and semi-structured datasets
- Build and deploy predictive, prescriptive, and descriptive models ensuring interpretability, robustness, and alignment with business objectives
- Partner with business, product, and analytics teams to validate assumptions, define success metrics, and deliver actionable insights
- Apply GenAI techniques to augment data science workflows—LLM-based insight generation, summarization, classification, and decision support
- Design and implement Retrieval-Augmented Generation (RAG) solutions to ground LLM outputs in enterprise data and analytical results
- Collaborate on GenAI-enabled analytical applications (e.g., conversational analytics, insight assistants) with focus on accuracy, relevance, and explainability
- Productionize data science and GenAI models using enterprise-grade MLOps/LLMOps practices—versioning, deployment, monitoring, and retraining
- Build scalable, secure, reliable analytical pipelines in collaboration with Data Engineering and Cloud teams
- Monitor model performance, data drift, and GenAI output quality; drive continuous improvements based on real-world usage
- Define and track model and GenAI performance metrics (accuracy, stability, bias, latency, business impact)
- Run experiments and controlled rollouts to optimize models, GenAI prompts, and retrieval strategies
- Ensure solutions meet enterprise requirements for governance, security, compliance, and responsible AI
REQUIRED EXPERIENCE
- 4-8 years in Data Science / AI Engineering
- 3+ years building and deploying machine learning models (supervised, unsupervised, time-series), covering feature engineering, model evaluation, and performance optimization
- 2+ years working with NLP or language-based systems, including text classification, information extraction, and semantic modeling
- 1+ years delivering GenAI or conversational AI solutions in production, with focus on applied LLM use cases, RAG, and enterprise deployment
CORE TECHNICAL SKILLS
- Strong foundation in statistics, machine learning, and applied data science
- Advanced proficiency in Python with hands-on production experience
- SQL expertise for data querying, transformation, and analytical pipeline development
- Apache Spark / PySpark for distributed data processing at scale
- Databricks ecosystem (Databricks SQL, MLflow, Feature Store, Jobs)
- ML frameworks: PyTorch and/or TensorFlow for model development and experimentation
- LangChain and LangGraph for operationalizing LLM-based analytical workflows, RAG, and prompt design
- MLOps/LLMOps practices—model and prompt versioning, deployment, monitoring, retraining strategies
- Production-level experience tracking model quality, data drift, and GenAI output reliability
- Data quality, explainability, responsible AI, and enterprise governance fundamentals
PREFERRED QUALIFICATIONS
- Vector databases and embedding models (FAISS, Pinecone, Weaviate, ChromaDB) for RAG
- Cloud platforms expertise (AWS SageMaker, Azure ML, or GCP Vertex AI)
- Experience evaluating and benchmarking GenAI outputs using quantitative and qualitative metrics
- Exposure to model serving and inference optimization for production systems
- Knowledge of data governance frameworks and compliance requirements
- Bachelor's or Master's in Computer Science, Statistics, Mathematics, Data Science, Engineering, or related field
ABOUT INFOVISION
InfoVision is a technology and talent solutions partner delivering enterprise-grade analytics, AI/ML, and digital transformation services. We enable organizations to harness data and artificial intelligence to drive measurable business impact through scalable, secure, and responsible AI solutions.