Company: krtrimaIQ Cognitive Solutions
Job type: Full-time
Location: Remote
Experience : 8+ years of relevant experience in data engineering, cloud data platforms, or related fields. Candidates should demonstrate a proven track record of designing and implementing production-grade data pipelines on Google Cloud Platform, optimizing complex data solutions, and working effectively in collaborative technical environments.
About the company:
KrtrimaIQ is a forward-thinking technology company specializing in cloud-based data solutions and intelligent data engineering practices. We empower organizations to harness the full potential of their data through scalable, reliable, and innovative platforms built on modern cloud infrastructure. Our team is composed of experienced data engineers, architects, and technical leaders who are passionate about solving complex data challenges. We work with enterprise clients across diverse industries, delivering production-grade data pipelines and analytics solutions that drive business intelligence and operational excellence.
Responsibilities
- Design and develop scalable ETL/ELT pipelines using Dataflow (Java/Apache Beam), and Cloud Composer (Airflow) to reliably process and transform data at scale
- Build and optimize large-scale data warehousing solutions in BigQuery, including partitioning, clustering, and query performance tuning
- Develop and maintain SQL-based transformation workflows using Dataform, applying sound data modeling practices for analytics-ready datasets
- Optimize data pipeline performance and cost-effectiveness across GCP services, including efficient use of BigQuery slots, Dataflow autoscaling
- Architect and implement data solutions using GCP-native services such as Cloud Storage, Pub/Sub, Cloud Functions, Cloud Composer, and Dataform, following cloud engineering best practices
- Troubleshoot and resolve complex data pipeline issues across the GCP data stack, implementing systematic and repeatable solutions
- Collaborate with analytics teams, data scientists, and business stakeholders to understand requirements and deliver solutions on GCP
- Establish and maintain data quality standards, implementing validation and monitoring mechanisms using tools like Cloud Monitoring, Cloud Logging, and Dataplex
- Leverage AI-assisted development tools to accelerate pipeline development, improve code quality, and streamline debugging
- Lead technical design discussions and mentor junior team members on GCP data engineering best practices
- Participate in code reviews and drive continuous improvement of data engineering processes, including CI/CD for GCP-based pipelines
- Mentor other data engineers on the team
Skills
- Design and Development of large scale data engineering pipelines on GCP with its native data engineering skills
- Expert proficiency with BigQuery for large-scale data warehousing, analytics, and query optimization
- Advanced hands-on experience with Java/Google Dataflow (Apache Beam) for building and managing streaming and batch data processing pipelines
- Strong Hands-on experience with Google Dataform for SQL-based data transformation and modeling within BigQuery
- Working knowledge of Cloud Composer (Airflow) for orchestration and workflow scheduling
- Familiarity with Pub/Sub for real-time messaging and event-driven pipelines, and Cloud Storage for data lake architecture
- Advanced Python and Java programming for data pipeline development and automation
- Strong SQL skills for complex data transformations and query optimization within BigQuery
- Strong SQL experience with advanced features like window functions, advanced joins, stored procedures, CTEs, Temp Tables, Conditional Pivoting etc.
- Familiarity with data modeling concepts (dimensional modeling, star/snowflake schemas) for analytics-ready data design
- Experience with AI-assisted development (e.g., Copilot, Claude, Gemini Code Assist) to accelerate pipeline development and code quality
- Proficiency with CI/CD practices (e.g., Cloud Build, GitHub Actions) for automated testing and deployment of GCP data solutions
- Working knowledge of containerization (Docker, GKE/Cloud Run) for packaging and deploying data applications
- Exposure to IAM, VPC Service Controls, and data governance practices within GCP is a plus