Search by job, company or skills

Data Engineer - Data & AI Integration Specialist

Data Engineer - Data & AI Integration Specialist

difinity digital
  • Posted 6 days ago
  • Be among the first 10 applicants

Job Description

Job Title: Data Engineer - Data & AI Integration Specialist

Location: Kochi

Employment Type: Full-Time - Permanent

Experience Level: Mid-Senior Level

About the Role

We are seeking a skilled Data Engineer with expertise in enterprise data integration, open-source technologies, cloud platforms, and modern AI engineering workflows. In this role, you will take end-to-end ownership of data pipeline design, ETL/ELT execution, data warehouse management, data governance, and AI data infrastructure. You will also serve as a technical bridge to clients by capturing business requirements, translating them into technical solutions, creating architecture diagrams, and supporting AI-driven analytics and application integrations.

Key Responsibilities

  • Architecture & Client Engagement: Communicate directly with clients to gather business requirements, analyze technical constraints, and draft comprehensive architecture diagrams for data and AI workloads.
  • ETL & Data Pipeline Design: Own, design, build, and maintain scalable ETL/ELT pipelines across cloud-based and open-source environments.
  • AI & Generative AI Data Infrastructure: Ingest, structure& unstructured and optimize data for vector databases, embedding generation, LLMs, and internal AI application integration (e.g., chatbots, intelligent search).
  • Warehouse Management & Governance: Oversee data warehouse administration and enforce strict data governance, quality checks, and validation frameworks.
  • Access Control & Security: Implement and manage Role-Based Access Control (RBAC) across all databases, schemas, tables, and AI datasets.
  • Data Ingestion: Extract, transform, and ingest data from relational/non-relational databases, REST APIs, web services, and flat files (CSV, JSON, XML, Excel).
  • Core Development: Write optimized, high-performance code using Python, PySpark, and SQL to process large-scale datasets.

Required Qualifications

  • 5+ years of hands-on experience in data engineering, data architecture, or enterprise data integration.
  • Mandatory expertise in Python, SQL, and PySpark.
  • Proven experience working with cloud-based data services (AWS, GCP, Azure, etc.) alongside open-source data technologies.
  • Practical experience building data pipelines that prepare and serve data for AI/ML workloads, vector databases, or embedding generation.
  • Demonstrated ability to build solution architecture diagrams and present technical concepts clearly to clients and stakeholders.
  • Solid experience in data warehouse management, dimensional modeling (star/snowflake schema), and configuring Role-Based Access Control (RBAC).
  • Strong background in end-to-end ETL/ELT workflow design and data governance enforcement.
  • Excellent client-facing communication skills with a focus on active requirement gathering.
  • Experience extracting and integrating data from major enterprise ERP systems (e.g., SAP, Salesforce, Oracle, Microsoft Dynamics).

Preferred Qualifications (Pluses)

  • Experience or strong familiarity with the Microsoft Azure ecosystem (Azure Data Factory, Azure Synapse, ADLS).
  • Understanding of Microsoft Fabric (Lakehouse/Warehouse paradigms).
  • Hands-on experience building or integrating AI tools, chatbots, or LLM applications.
  • Understanding of Power BI integration and key financial reports/business analytics.
  • Familiarity with AI-assisted software development and code-generation tools (e.g., GitHub Copilot, Cursor).
  • Basic knowledge of MLOps practices, model deployment, and serverless workflow orchestration (Azure Functions, Logic Apps).

More Info

Job Type:
Industry:
Employment Type:

Key Skills

cloud platforms

vector databases

data ingestion

embedding generation

About Company