Search by job, company or skills

Data Engineer

  • Posted 23 days ago
  • Over 50 applicants have applied

Job Description

Role Summary

Design, build and operate scalable enterprise data platforms that power AI applications and business workflows. Own ingestion, transformation, storage, indexing and delivery of structured and unstructured data from enterprise systems and AI providers. Ensure AI-generated signals, metadata and telemetry are reliable, searchable and available for real-time workflows and analytics.

Key Responsibilities

  • Build scalable ETL/ELT and streaming pipelines from enterprise systems and third-party AI platforms.
  • Develop API connectors to ingest AI-generated metadata, events and telemetry.
  • Normalize, validate and enrich data into common enterprise schemas.
  • Build metadata repositories, data lakes, vector indexes and search infrastructure.
  • Optimize storage and retrieval for enterprise AI applications.
  • Partner with Solution Architects, AI Engineers and Backend Engineers.
  • Implement data quality, lineage, governance and monitoring.

AI & Data Platform Responsibilities

  • Ingest outputs from LLMs, video AI, speech AI, vision AI and enterprise systems.
  • Build vector indexing and semantic retrieval pipelines.
  • Manage embeddings, metadata stores and searchable knowledge repositories.
  • Support AI analytics, telemetry and operational reporting.

Required Technical Skills

  • Python, SQL Spark/Airflow or equivalent is critical .
  • Kafka/RabbitMQ or similar streaming platforms.
  • PostgreSQL, MongoDB, Redis.
  • Vector databases (Pinecone, Weaviate, Milvus or equivalent).
  • REST APIs, JSON, Parquet.
  • Docker, Kubernetes fundamentals and AWS/Azure/GCP.

AI-Driven Engineering Practices (Mandatory)

  • Experience using AI coding assistants such as GitHub Copilot, Cursor, Windsurf, Claude Code or equivalent.
  • Use AI for ETL/ELT development, SQL generation, API integration, documentation, testing and refactoring.
  • Review and validate AI-generated code for correctness, security, maintainability and performance.
  • Use AI to troubleshoot pipeline failures, optimize performance and automate repetitive engineering tasks.
  • Knowledge of prompt engineering and AI-assisted software development lifecycle.

Preferred Experience

  • Enterprise AI platforms, workflow automation, media technology, OTT, digital asset management or enterprise SaaS.

Success Measures

  • Reliable, scalable, low-latency data pipelines.
  • High-quality searchable metadata and vector indexes.
  • Near real-time availability of AI-generated signals.
  • Strong governance, monitoring and operational reliability.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 150840597

Similar Jobs

Mumbai, India

Skills:

Apache AirflowApache FlinkPostgresql SqlPython

Mumbai, India

Skills:

Spark SQLSpark CoreAzure Data FactorySpark StreamingAzure Synapse AnalyticsPysparkAzure DatabricksPythonSqlAzure Data Lake Storage ADLS Gen2

Mumbai

Skills:

catalog S3AvroLambdaEncryptionKinesisMySQLPythonJavaEmrOracle DbKmsIamAirflowSecrets Managerdata de serializationEventBridgedata lakesJSON-LDAI-assisted software development toolsSQL-based technologiesParquetStep FunctionsAWS security controlsMSKGlue ETLIcebergLake FormationAWS cloud technologiesGlueAthena

Remote

Skills:

Data GovernancePythonSqlJavaSalesforceFine tuning

Mumbai, India

Skills:

graph databases SparqlHadoopKafkaTigerGraphRdfGcpNeo4jSparkAzureAWSowlCypherAmazon NeptuneSemantic ModellingGremlinOntology