Job Title: AI Data Engineer
Duration: 6 months contract starting with possibility for extension
100% Remote (India)
Fulltime contract 8 hours per day/40 hour per week
Introduction: We are seeking an experienced AI Data Engineer to join our dynamic team and play a pivotal role in building and maintaining the AI-ready data foundation that powers agent workflows, analytics, and batch disposition decisioning. This role is critical in enabling data-driven decision-making across the organization by ensuring seamless data ingestion, integration, storage, governance, and enablement across enterprise systems and cloud platforms. The ideal candidate will have a strong background in designing scalable data solutions, leveraging cutting-edge technologies to support AI and analytics initiatives.
Roles and Responsibilities:
- Develop and maintain robust data pipelines integrating multiple enterprise systems such as SAP, gLIMS, Veeva, MODA, AMPS, and Batch Tracker to ensure reliable data flow and accessibility.
- Design, implement, and optimize Snowflake data models and persistent storage solutions tailored for AI and advanced analytics use cases.
- Build and manage API and event-driven integrations to support real-time agent workflows and processing requirements.
- Implement comprehensive data quality management, data lineage tracking, governance frameworks, and continuous monitoring to maintain data integrity and compliance.
- Enable AI agents and Retrieval-Augmented Generation (RAG) solutions by curating trusted, accessible, and well-governed data assets.
- Ensure all data solutions are secure, scalable, and compliant with organizational policies and industry regulations.
- Collaborate with cross-functional teams including data scientists, analysts, and IT to align data engineering efforts with business objectives.
Qualifications:
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
- 6+ years of hands-on experience in data engineering, with a focus on AI and analytics data infrastructure.
- Proven expertise in Data Modelling, data pipeline development, and ETL processes.
- Strong programming skills in Python for data processing and automation tasks.
- Proficiency in writing and optimizing complex SQL queries.
- Extensive experience working with cloud platforms, especially AWS, including services related to data storage, processing, and security.
- Hands-on experience with Snowflake for data warehousing and scalable data storage solutions.
- Familiarity with API development and event-driven architecture to support real-time data workflows.
- Knowledge of RAG (Retrieval-Augmented Generation) techniques and their data requirements is highly desirable.
- Strong understanding of data governance, data quality frameworks, and compliance standards.
- Excellent problem-solving skills, attention to detail, and ability to work collaboratively in a fast-paced environment.
Tools and Technologies:
- Python
- SQL
- ETL frameworks and tools
- AWS (e.g., S3, Lambda, Glue, Redshift)
- Snowflake
- API development and integration tools
- Data governance and monitoring tools
- Familiarity with RAG implementation frameworks