Role: Senior Data Architect – AWS Cloud & Enterprise Data Solutions
Function: Data Architecture / Cloud Engineering
Location: Bangalore
Type: Full-time
Industry: Information Technology & Services, Management Consulting
About Company
The company is a digital engineering firm founded in 2020, headquartered in Tampa, Florida. It specializes in AI-driven digital transformation and platform engineering for enterprises.
Over 450 professionals work across seven global offices, including Bangalore, Trivandrum, Toronto, Dallas, Belgrade, Johannesburg, and Bogota. The company has completed 55+ projects across 25+ countries.
In 2023, it expanded its product engineering capabilities through the acquisition of another firm. Its culture is fast-paced, ownership-driven, and deeply client-centric.
Position Overview
This role leads the design and delivery of enterprise-scale data platforms spanning data warehousing, data lakes, real-time streaming, and ML pipelines on AWS. The architect sets data infrastructure standards, drives governance and security frameworks, and makes strategic technology decisions that directly shape how the organization and its clients derive value from data. This is a hands-on leadership role with direct influence over architecture review boards, engineering team mentorship, and platform roadmap.
Role & Responsibilities
- Design and own enterprise data architectures spanning data warehouses (Redshift), data lakes (S3 medallion architecture), real-time streaming (Kinesis, MSK), and ML pipelines (SageMaker), aligned to business scalability requirements
- Architect ETL/ELT pipelines using AWS Glue (PySpark/Scala), Lambda, and Step Functions — incorporating CDC patterns, incremental load strategies, and data quality validation frameworks
- Design and optimize Amazon Redshift clusters including distribution/sort key strategies, compression, concurrency scaling, materialized views, and query performance tuning
- Implement data lake governance using AWS Lake Formation with fine-grained access control, data cataloging, Parquet/ORC format optimization, and lifecycle management
- Design real-time streaming architectures using Amazon Kinesis (Data Streams, Firehose, Analytics) and Apache Flink for millisecond-latency event-driven pipelines
- Define and enforce data security standards — KMS encryption, IAM/ABAC policies, PII masking, audit logging — ensuring compliance with GDPR, HIPAA, and SOC 2
- Lead architecture review boards, establish data platform standards and patterns, and mentor data engineers and analytics engineers across the organization
- Implement Infrastructure as Code using CloudFormation or Terraform and design CI/CD pipelines for automated data infrastructure deployments
Must Have Criteria
- 7+ years of AWS cloud experience with 4+ years focused specifically on data architecture (Redshift, S3, Glue, RDS, DynamoDB, Athena, Lambda, EMR, Kinesis)
- Expert-level Amazon Redshift knowledge: cluster design, node types, distribution/sort keys, compression, concurrency scaling, and query optimization via EXPLAIN plans
- Hands-on AWS Glue expertise: job development in PySpark and/or Scala, crawlers, schema discovery, data catalog management, and CDC pattern implementation
- Proven experience designing medallion (bronze/silver/gold) data lake architectures on S3 with AWS Lake Formation for governance and access control
- Strong proficiency in Python (including Boto3/AWS SDK) and expert-level SQL for complex query writing, performance tuning, and data manipulation
- Experience implementing data security frameworks: IAM policies, KMS encryption, VPC security, data masking/tokenization, and regulatory compliance (GDPR, HIPAA, SOC 2)
- Infrastructure as Code experience with AWS CloudFormation and/or Terraform, including CI/CD pipeline design for data infrastructure deployments
Nice to Have
- AWS Certified Data Analytics – Specialty, AWS Certified Solutions Architect – Professional, or AWS Certified Database – Specialty certification
- Experience with Apache Kafka or AWS MSK, Delta Lake or Apache Iceberg, and/or Apache Flink for advanced streaming and ACID-compliant data lake patterns
- Hands-on experience with Apache Airflow or similar orchestration tools, and/or modern BI platforms such as Tableau, Power BI, or Looker
- Domain expertise in finance, healthcare, retail, or telecommunications, with a track record of leading enterprise data migration or modernization programs
- Experience with AWS SageMaker for ML pipeline integration, or AWS DMS for large-scale database migration projects
What We Offer
- Architectural ownership of enterprise-scale data platforms with direct impact on client outcomes across 25+ countries
- Hands-on access to the full AWS data services stack — Redshift, Glue, Lake Formation, Kinesis, SageMaker — on real, large-scale engagements