Search by job, company or skills

Data Engineer

5-10 Years
  • Posted 7 hours ago
  • Be among the first 10 applicants

Job Description

About Us

We are an early-state AI startup building a structured research data platform for the investment research domain. We serve professional investment teams focused on industry fundamental research. With our first batch of seed clients (mainstream hedge funds) already onboarded, we are a lean and highly efficient team.

We are looking for an early core engineer to work alongside the founding team to build this data platform from 0 to 1: from data ingestion and structured parsing, to query services for analysts. This is an end-to-end role where you will have complete ownership.

Core Responsibilities

1. Data Ingestion

  • Build and maintain multi-source data ingestion and pipelines.
  • Ensure pipelines are robust, stable, and production-scheduled.

2. Document Parsing

  • Extract structured data from formats like PDF, HTML, and XBRL.
  • Handle scanned PDFs, complex tables, and multilingual text (Chinese/English).

3. Data Warehouse Design

  • Design and maintain schema-based tables.
  • Leverage modern cloud data warehouses (Snowflake, BigQuery, or Postgres).

4. Pipeline Orchestration & Monitoring

  • Build production-grade ETL/ELT using Airflow, Prefect, or Dagster.
  • Implement data quality checks, alerting, and data lineage tracking.

5. LLM Application Engineering

  • Implement OpenAI/Anthropic APIs and open-source models.
  • Productionize structured outputs and Retrieval-Augmented Generation (RAG).
  • Set up basic LLM evaluation and quality monitoring frameworks.

6. Backend Service Layer

  • Build high-performance query APIs using FastAPI or similar frameworks.
  • Power downstream investment analyst tools.

7. Cloud Infrastructure

  • Utilize cloud ecosystems (GCP or AWS).
  • Use managed services to minimize operational and DevOps overhead.

Key Requirements

1. Engineering Experience

  • 5–10 years of production-level software engineering experience.
  • Proven track record of building and deploying complete online systems.

2. Technical Stack

  • Strong proficiency in Python and SQL.
  • Solid experience in data modeling, writing tests, and performance tuning.

3. Document Processing

  • Hands-on experience parsing messy, real-world documents (PDF, HTML, XBRL).

4. AI & LLM Expertise

  • 2–3 years of hands-on experience deploying production LLM applications.
  • Proven experience in launching real-world AI features.

5. Data Architecture

  • Mastery of at least one core orchestration framework (Airflow, Prefect, Dagster).
  • Proficiency in schema design for at least one modern cloud data warehouse.

6. Soft Skills & Communication

  • Strong communication skills to align with research and founding teams.
  • Ability to articulate technical designs clearly in discussions.
  • Professional working proficiency in English.

7. Strong Preferred

  • Previous experience as a Founding Engineer.
  • Experience building end-to-end data systems at early-stage AI startups.

Preferred Qualifications (Bonus Points)

  • Experience building data infrastructure tailored for NLP/LLM workloads.
  • Hands-on experience with dbt (Data Build Tool).
  • Design experience in multi-tenancy and data access control.
  • Basic familiarity with graph databases (e.g., Neo4j).

What We Offer

  • Direct Impact: Your work forms the bedrock of our platform, directly impacting investment research daily.
  • Competitive Compensation: An attractive salary and equity package.
  • Elite Collaboration: Work directly with seasoned investment professionals (analysts, PMs, and quants).
  • High-Execution Culture: A rigorous, fast-paced, and execution-oriented work environment.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 153602583

Similar Jobs

Singapore

Skills:

JavaData ModellingPysparkKafkaSqlELTDockerTerraformKubernetesPythonEtlAirflowcdcFlinkCI CDIceberglakehouse architectureNiFiCloudera Machine Learningobservability tooling

Singapore

Skills:

Apache SparkData IntegrationData ModellingData TransformationEtl Developmentelastic search/open searchpython/scala

Singapore, Cecil Street

Skills:

snowflake HadoopMicrosoft Power BiTableauData IntegrationGoogle CloudHiveHueData ArchitectureData LakeDatabricksAzureTalendPythonAWSEtlAlibaba Cloud

Singapore, Beach Road

Skills:

data warehouses snowflake JavaData Warehouse ConceptsScalaDimensional ModelingSqlData LakePythonAirflowdbt-coreBig Data processing frameworksSCDslake housedata meshmodern data warehouse

Singapore

Skills:

BigQueryAws RedshiftHadoopPysparkKafkaAWS AthenaSqlNosqlGcpMySQLSparkSnsData LakePythonAWSAws S3EtlApache IcebergParquetAirflowAWS MSKFirehose

Beware of Scammers

We don’t charge money for job offers