Search Jobs

Search by job, company or skills

Lead Data Scientist

Lead Data Scientist

Info Origin Inc.
Early Applicant
  • Posted 4 days ago
  • Be among the first 10 applicants

Job Description

Job Title: Lead Data Scientist

Location: Noida

Employment Type: Full-Time

Experience Required: 5+ Years

About the Role

We're seeking a seasoned Lead Data Scientist to lead NLP/ML research efforts and architect end-to-end solutions. You will be instrumental in driving innovation, mentoring junior team members, and working closely with stakeholders to align AI capabilities with strategic goals.

Key Responsibilities

• Lead the development and deployment of production-grade NLP and ML models.

• Work with large language models (LLMs), RoBERTa, GPT APIs, and hybrid architectures.

• Guide and review the work of junior scientists and interns.

• Develop pipelines for training, evaluation, and continuous model monitoring.

• Collaborate with cross-functional teams to scope new projects and define success metrics.

• Write technical papers, PoCs, and support patentable innovations when applicable.

Required Skills

• 5+ Years of relevant experience along with prior exposure to team leadership responsibilities.

• 5+ Years of deep knowledge of modern NLP architectures (transformers, LLMs, embeddings).

• Design, build, and deploy Generative AI models (e.g., LLMs, Diffusion Models) into production systems.

• Strong understanding of Neural Networks, Transformer models (e.g., BERT, GPT), and deep learning frameworks (e.g., TensorFlow, PyTorch).

• Proficiency in Python and ML libraries (e.g., Scikit-learn, Hugging Face, Keras, OpenCV).

• Proficiency in data pipelines, MLOps, model versioning, and evaluation frameworks.

• Ability to distill complex problems and explain them clearly to technical and non-technical audiences.

Preferred Qualifications

• Published work in conferences/journals or experience contributing to open-source projects.

• Familiarity with vector databases (Pinecone, FAISS), LangChain, or prompt engineering.

• Experience in managing research interns or small AI/ML teams.

Education

Master's or PhD in Computer Science, Machine Learning, or a related field.

Required Skills

5+ Years of deep knowledge of modern NLP architectures (transformers, LLMs, embeddings).

Design, build, and deploy Generative AI models (e.g., LLMs, Diffusion Models) into production systems.

Strong understanding of Neural Networks, Transformer models (e.g., BERT, GPT), and deep learning frameworks (e.g., TensorFlow, PyTorch).Proficiency in Python and ML libraries (e.g., Scikit-learn, Hugging Face, Keras, OpenCV).Proficiency in data pipelines, MLOps, model versioning, and evaluation frameworks.Ability to distill complex problems and explain them clearly to technical and non-technical audiences.

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

embeddings

data pipelines

model versioning

Diffusion Models

Hugging Face

Scikit-learn

RoBERTa

GPT

BERT

evaluation frameworks

Generative AI models

About Company

Similar Jobs

5-7 yrs
Gurugram, Gurugram, India
Skills:
Data Science, Pyspark, Databricks, Azure, Python, Sql, Deep Learning, AWS, Supervised and Unsupervised ML, Statistics, GenAI concepts
5-7 yrs
Delhi, India
Skills:
snowflake , Machine Learning, Cassandra, PostgreSQL, Nosql, Nlp, Terraform, Docker, MySQL, Elasticsearch, Oracle, Python, AWS, SQL Server, HBase, Redis, Sql, Deep Learning, Hive, Gcp, Amazon Redshift, Spark, MongoDB, Azure, Airflow, MLflow, Beam, GitHub Actions, Teradata, Luigy, Nifi
7-9 yrs
Gurugram, Gurugram, India
Skills:
object detection , Machine Learning, Pyspark, Sql, Deep Learning, Docker, Kubernetes, Ocr, Python, Computer Vision, LLMs, Geospatial models, Facial recognition, Cloud environments, Statistical Modeling, ML pipelines, Real-time data processing frameworks, RAG architectures, Time-series forecasting
3-8 yrs
Gurugram, India, Gurugram
Skills:
SAS, Sql, Spark, Python, R
8-10 yrs
Gurugram, India
Skills:
Text Analytics, Tensorflow, Numpy, Nlp, Nltk, Pytorch, Excel, Python, Power Bi, Sql, Pandas, Spark, Databricks, knowledge graphs, pgvector, scikit-learn, vector databases, evaluation frameworks, Streamlit, RAG pipelines, SpaCy, LLMs, prompt engineering, Agentic AI, FAISS, feature stores, knowledge databases, embeddings