Job Description
The Data Platform Architect is the lead designer and technical authority for our modern data ecosystem. Architect is responsible for ensuring the platform provides a high-performance, cost-effective, and secure foundation that evolves alongside business needs. This is a design-and-direction role where you set the standards, review engineering output, and serve as the primary bridge between complex technical systems and business value.
Work Mode : Remote
What the Architect will own:
- Platform Design & Research: Lead architectural decisions regarding compute engine selection, open-table format implementation, and tiered storage design.
- Agnostic Infrastructure: Architect a decoupled data environment that ensures interoperability across multiple engines and prevents proprietary vendor lock-in.
- Governance & Compliance: Design and oversee the implementation of automated data governance, including PII discovery, row/column-level security, and auditability.
- Standards & Frameworks: Define the Definition of Done for data pipelines, establishing coding standards, CI/CD patterns, and technical documentation requirements.
- Architectural Decision Records (ADR): Maintain a version-controlled repository of all consequential technical decisions, documenting the rationale, trade-offs, and long-term implications.
- Performance & FinOps: Monitor and optimize platform performance and spend, ensuring sub-second query speeds for massive user bases while maintaining a lean cloud footprint.
- Technical Stewardship: Conduct deep-dive code and design reviews for all data models and orchestration workflows; mentor and unblock senior engineering staff.
Technical Skills & Experience:
- 8+ Years in Data Engineering / Architecture: Proven experience delivering production-grade Lakehouse environments for high-concurrency (1,000+ user) organizations.
- Modern Data Stack Fluency: Extensive experience with Databricks (Lakehouse/Unity Catalog) and high-performance warehouses like Amazon Redshift or Snowflake.
- Open-Table Formats: Deep hands-on expertise with Apache Iceberg or Delta Lake, including optimization strategies for partitioning and schema evolution.
- Transformation & Modeling: Mastery of dbt (Core) for complex SQL-based modeling and PySpark or Python for sophisticated data processing.
- High-Efficiency Compute: Familiarity with vectorized/embedded engines like DuckDB for specialized or cost-sensitive processing tasks.
- Orchestration Mastery: Advanced experience with Apache Airflow, specifically in designing resilient, dependency-aware DAGs in resource-constrained environments.
- Cloud Ecosystems: Expert-level knowledge of AWS (S3, EC2, IAM) or equivalent services in Azure/GCP, with a focus on storage-compute separation.
- Experience with open-source catalog implementations like Apache Polaris.
- Knowledge of Data Ops principles and automated data quality testing frameworks.
- Experience translating technical debt and architectural roadmaps for C-level executives.
- Background in managing fixed-resource infrastructure (e.g., EC2/VM-based processing) vs. elastic serverless models.
- Kubernetes (K8s). Deep understanding of Pods, Deployments, Services, ConfigMaps, and Secrets management.
How We Partner To Protect You: TaskUs will neither solicit money from you during your application process nor require any form of payment in order to proceed with your application. Kindly ensure that you are always in communication with only authorized recruiters of TaskUs.
DEI: In TaskUs we believe that innovation and higher performance are brought by people from all walks of life. We welcome applicants of different backgrounds, demographics, and circumstances. Inclusive and equitable practices are our responsibility as a business. TaskUs is committed to providing equal access to opportunities. If you need reasonable accommodations in any part of the hiring process, please let us know.
We invite you to explore all TaskUs career opportunities and apply through the provided URLhttps://www.taskus.com/careers/.
More Info
Key Skills
Apache Iceberg
Data Engineering Architecture
dbt Core
Databricks Lakehouse
Apache Polaris
Unity Catalog
DuckDB
Data Ops principles
Delta Lake



