

Search by job, company or skills

This job is no longer accepting applications
Company Description
Blend is a premier AI services provider, committed to creating meaningful impact for its clients through the power of data science, AI, technology, and people. We help organisations solve complex business challenges by combining deep domain understanding with modern data and AI capabilities. Our teams work across strategy, analytics, engineering, and product delivery to create scalable, high-value solutions that improve decision-making, efficiency, and growth
Job Description
You'd be joining the software engineering team behind a production-grade, event-driven platform at the intersection of cloud infrastructure and AI. The system processes complex, multi-step workflows in near real-time with AI capabilities woven throughout — running on Azure, managed as a Python Nx monorepo
This isn't a role that values depth in one discipline above all else — it's a role that values breadth. We're looking for a mid-level engineer who has worked meaningfully across software engineering, data engineering, and DevOps, and who can use that perspective to understand how a complex distributed system behaves end-to-end. Your primary focus will be building and owning the observability layer for the platform: Grafana dashboards that give the engineering team clear visibility into system health, performance, and behaviour. You'll work collaboratively with the team to discover what needs to be measured, build the tooling to surface it, and grow into the team's subject matter expert on observability.
Responsibilities
• Design and build Grafana dashboards that monitor the platform across multiple dimensions: service health, throughput, latency, queue depths, error rates, and alerting for impending or ongoing issues
• Instrument and surface metrics from across the system — including AI-specific signals such as LLM token consumption and inference timing
• Work closely with the engineering team to identify the data points that matter and translate them into useful, actionable views
• Connect observability tooling to Azure Monitor and Open Telemetry data sources
• Set up and maintain alerting to give the support team early warning of degradation or failure
• Act as the team's SME on observability — helping engineers understand what the dashboards reveal, coaching them to build their own dashboard components, and growing collective capability across the team
• Contribute to the broader platform engineering effort across CI/CD, deployment tooling, and cloud infrastructure as needed
Qualifications
Job ID: 151911507
Skills:
snowflake , composer , Maven, ELT, Docker, Terraform, Flask, Nexus, Python, Apache Spark, Jenkins, Git, Bitbucket, FastAPI, Bamboo, Etl, Flink, Beam, Pub Sub, Cloud Run, GitHub Actions, Dash, Cloud SQL, Streamlit, Cloud Functions, GCP Deployment Manager, GCS
Skills:
automation, Linux, shell scripting, Kubernetes, Alerts, Python, GitHub Copilot, GitHub Actions, logs, GitOps, agentic engineering tools, observability metrics, Claude Code
Skills:
Bash, Artifactory, Saltstack, Jenkins, Gcp, Linux, Terraform, Docker, Ansible, Splunk, Kubernetes, Python, AWS, Harness
Skills:
.NET, react.js , Kafka, Spring Boot, Nosql, Azure Active Directory, Docker, Cosmos DB, Python, Azure DevOps, Express.Js, Oauth2, Jwt, Node.js, Sql, Redis, Jenkins, SignalR, FastAPI, Azure SQL Database, GRPC, LangGraph, WebSockets, GitHub Actions, LangChain, Azure App Services, Azure Blob Storage, Azure Kubernetes Service
Skills:
Power Bi, Sql, Azure Data Factory, Azure, Python, Azure DevOps, Airflow, Data Pipelines, Great Expectations, Azure Key Vault, CI CD, GitHub Actions, R, Data Quality Framework, rbac, Microsoft Fabric, SoDA