
Search by job, company or skills
About the Role
We are looking for an AI Data & Knowledge Engineer witha strong foundationin data engineering and an interest in AI-powered data systems. This role will help build andmaintainscalable backend data pipelines, ETL/ELT workflows, and structured knowledge assets that power enterprise AI applications. The ideal candidate will contribute to big data processing, data integration, retrieval pipelines, and knowledge graph foundations that enable reliable GenAI use cases.
Key Responsibilities
Build andmaintainbackend data pipelines to ingest, transform, and serve structured and unstructured data for AI applications.
Support ETL/ELT workflows across enterprise source systems, data platforms, and downstream AI services.
Assistin developing RAG pipelines, vector indexing workflows, and knowledge graph assets under guidance from senior engineers.
Contribute to datamodeling, ontology creation, metadata tagging, and semantic enrichment of enterprise data.
Support MCP server integration and AI data enablement tasks for domain use cases.
Perform data quality checks, validation, reconciliation, and documentation for pipelines and data contracts.
Helpoptimizepipeline performance, reliability, and scalability across batch and near-real-time workloads.
Participate in code reviews, sprint ceremonies, and engineering discussions to build platform and domain knowledge.
Required Qualifications
Bachelor's degree in Computer Science, Data Engineering, Information Systems, or a related field.
0-7 years of experience in data engineering, software engineering, or a related technical role.
Working knowledge of Python and SQLand/or Java.
Basic understanding of ETL/ELT pipelines, data transformation, and data integration concepts.
Exposure to big data or distributed processing tools such as Spark, Databricks, or similar platforms is preferred.
Basic familiarity with data lakes, warehouses, orlakehousearchitectures.
Understanding of data quality, metadata, and governance concepts.
Exposure to RAG, vector databases, semantic search, or knowledge graph concepts is a plus.
Familiarity with orchestration tools such as Airflow or similar platforms is beneficial.
Across the globe, institutional investors rely on us to help them manage risk, respond to challenges, and drive performance and profitability. We keep our clients at the heart of everything we do, and smart, engaged employees are essential to our continued success.
We are committed to fostering an environment where every employee feels valued and empowered to reach their full potential. As an essential partner in our shared success, you'll benefit from inclusive development opportunities, flexible work-life support, paid volunteer days, and vibrant employee networks that keep you connected to what matters most. Join us in shaping the future.
As an Equal Opportunity Employer, we consider all qualified applicants for all positions without regard to race, creed, color, religion, national origin, ancestry, ethnicity, age, disability, genetic information, sex, sexual orientation, gender identity or expression, citizenship, marital status, domestic partnership or civil union status, familial status, military and veteran status, and other characteristics protected by applicable law.
Discover more information on jobs at
Read our
At State Street, we partner with institutional investors all over the world to provide comprehensive financial services, including investment management, investment research and trading, and investment servicing.
Whether you are an asset manager, asset owner, alternative asset manager, insurance company, pension fund or official institution, you can rely on us to be focused on your challenges. We are committed to doing what it takes to help you perform better — now and in the future.
Job ID: 152100583
Skills:
Java, Data Transformation, Sql, ELT, Data Quality, Spark, Databricks, Data Integration, Big Data, Python, Etl, Airflow, RAG vector databases, distributed processing tools, knowledge graph concepts, data lakes, governance concepts, metadata, semantic search