Search by job, company or skills

AI Data Architect / Senior Data Engineer

Early Applicant
  • Posted 5 days ago
  • Be among the first 20 applicants

Job Description

The Role

As an AI Data Architect, you will own the data pipeline that powers our agentic AI platform. You will design the ingestion, transformation, contextualization, enrichment, validation, and semantic modelling layers that connect a wide range of structured and unstructured enterprise data sources into an AI-ready data corpus.

This is a senior individual contributor role with real ownership over the data foundation of the venture. The role requires hands-on experience in enterprise data engineering, schema discovery, data profiling, data quality automation, ontology implementation, knowledge graph integration, and cloud-scale pipeline engineering.

Key Responsibilities

  • Design production-grade enterprise connectors and ETL/ELT pipelines for both structured enterprise systems such as ERP, CRM, OSS/BSS, billing, finance, HR, and unstructured sources such as emails, documents, logs, transcripts, and media files.
  • Build ingestion and transformation pipelines using Python, SQL, PySpark, Apache Spark, Airflow, dbt, Dagster, Flink, or equivalent technologies.
  • Create frameworks for data labelling, contextualization, harmonization, enrichment, and classification workflows to configure AI agents.
  • Architect integration with knowledge graphs and vector databases for hybrid search, semantic retrieval, contextual reasoning, and AI-ready data access.
  • Build and maintain Ontology/knowledge graph pipelines using Neo4j, RDF/OWL, Apache Jena, Stardog, GraphDB, or equivalent technologies.
  • Implement graph validation frameworks such as SHACL or ShEx to programmatically enforce data integrity rules over enterprise knowledge graphs.
  • Implement data quality automation using frameworks such as Great Expectations, AWS Glue DataBrew, dbt tests, custom validation pipelines, or equivalent tools.
  • Define data profiling routines for real-world enterprise data, including missing keys, duplicate entities, inconsistent encoding, changing column meanings, incomplete master data, and conflicting source records.
  • Experience implementing semantic guardrails, jailbreak protection, data exfiltration prevention, and toxic output mitigation. RabbitMQ / Apache Kafka (Agent Message Queuing).
  • Implement privacy and compliance controls, including masking, anonymization, access control, PII handling, GDPR compliance, and Indian Digital Personal Data Protection Act / DPDP Act alignment.
  • Partner with AI/ML architects to ensure pipeline outputs match agent input contracts, retrieval requirements, ontology models, and downstream AI consumption patterns.
  • Mentor junior data engineers, lead design reviews, and help establish engineering practices for a high-quality, product-grade data platform.

Must-Have Qualifications

  • 14+ years of experience in data engineering, data architecture, platform engineering, or enterprise-scale data solution delivery.
  • Strong production experience on modern data platforms such as Databricks, Snowflake, BigQuery, cloud data lakes, lakehouses, or equivalent enterprise data platforms.
  • Deep working knowledge of Python, SQL, PySpark, Apache Spark, and modern data pipeline development practices.
  • Hands-on experience with both structured and unstructured data ingestion at enterprise scale.
  • Strong experience in building pipelines for enterprise sources such as ERP, CRM, OSS/BSS, billing systems, finance systems, ServiceNow, Salesforce, SAP, Oracle, and legacy databases.
  • Workingknowledge of vector databases such as Pinecone, Weaviate, pgvector, Milvus, Chroma, or equivalent technologies.
  • Hands-on knowledge of knowledge graphs, graph data modelling, graph querying, and enterprise graph implementation using Neo4j, Cypher, RDF, OWL, or equivalent technologies.
  • Experience with semantic data models, ontologies, industry data standards, or domain-specific enterprise taxonomies.

Good to Have

  • Exposure to telecom, BFSI, manufacturing, or other complex enterprise domains.
  • Experience with OSS/BSS, ERP, CRM, billing, order management, product catalog, service inventory, or network inventory systems.
  • Experience with RDF triple stores such as Apache Jena, Stardog, GraphDB, Amazon Neptune, or equivalent technologies.
  • Experience with data catalogues, metadata management tools, lineage platforms, or governance platforms.
  • Contributions to open-source data tooling, graph tooling, ontology tooling, or data quality frameworks.

Why This Role Is Exciting

You will architect the data foundation from day one. Your designs will shape how the platform discovers, understands, contextualizes, validates, and prepares enterprise data for AI consumption.

This role offers the opportunity to build the core data backbone for a venture-backed Infosys platform, influence early product architecture, work closely with AI/ML architects, and create a scalable foundation that can support multiple industries and enterprise clients over time.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151364713