Senior AI Data Engineer
jobline resources pte. ltd.- Posted 6 hours ago
- Be among the first 10 applicants
Job Description
Responsibilities
. Design, build, optimise, and maintain batch and streaming data ingestion pipelines using platforms such as Databricks and Kafka, ensuring scalability, reliability, observability, and alignment with enterprise data architecture standards.
. Perform data transformation and cleansing using PySpark or SQL based on business and technical requirements
. Monitor and troubleshoot data workflows to ensure data quality and pipeline reliability
. Provide technical guidance to engineers and delivery partners on data platform patterns, reusable components, code quality, deployment readiness, and production support practices.
. Lead integration of data from diverse source systems including files, APIs, databases, and streaming platforms, working with source-system owners and consuming teams to define fit-for-purpose ingestion patterns and delivery timelines.
. Help maintain metadata and pipeline documentation for transparency and traceability
. Own production readiness for assigned data and AI platform components, including observability, incident triage, root-cause analysis, release coordination, and continuous improvement of operational runbooks.
. Participate in integrating pipelines with tools such as Microsoft Fabric, Databricks, Delta Lake, and other platform components
. Build and maintain knowledge base and RAG solution on variety of hosting platforms
. Implement and operate knowledge base storage, lifecycle management and embedding/vectorization
. Contribute to automation efforts using version control and CI/CD workflows
. Apply data governance, security, access control, and operational risk policies during solution design and implementation, ensuring pipelines and knowledge platforms meet enterprise compliance requirements.
Requirements
. Bachelor's degree in Computer Science, Engineering, or a related field
. 5-8 years of experience in data engineering, data platform engineering, or cloud-scale analytics solution delivery, with demonstrated ownership of production pipelines and platform components.
. Proven ability to independently design, build, optimise, and operate production-grade batch or streaming data pipelines, including orchestration, observability, error handling, performance tuning, and operational support.
. Hands-on experience with Python and SQL for data transformation and validation
. Familiarity with Apache Spark (especially PySpark) and large-scale data processing concepts
. Experience with implementing knowledge base and RAG solutions for agentic AI use cases
. Self-starter with strong problem-solving skills and a keen attention to detail
. Able to work independently and lead technical discussions with engineers, architects, product owners, source-system teams, and business stakeholders to translate requirements into secure and maintainable platform solutions.
. Strong documentation and communication skills
. Strong understanding of enterprise data architecture, cloud security, access control, CI/CD, release management, and production operations for data and AI platform solutions.
Licence no: 12C6060



