Search by job, company or skills

Sr Technical Consultant - Azure Cloud, DevOps, Sql & Site Reliability Engineering(SRE)

Early Applicant
  • Posted 23 hours ago
  • Be among the first 10 applicants

Job Description

Overview:

  • The BYX Platform provides shared cloud provisioning, compute, storage, observability and runtime capabilities for enterprise applications and implementation teams.
  • We are seeking an astute Cloud Platform with a strong technical foundation and hands-on experience in Microsoft Azure, Kubernetes, containers, observability, networking and production automation.

Our current technical environment:

  • Cloud and Compute: Microsoft Azure, Kubernetes, node pools, container workloads, Azure Container Registry, KEDA, Dask, autoscaling, CPU, memory and GPU capacity.
  • Observability: Logs, metrics and traces collectors and exporters Elastic/Elasticsearch, Logstash, Kibana, dashboards, alerts and correlation identifiers.
  • Data and Storage: MongoDB/Atlas, SQL, Redis, NFS, managed storage, replicas, regional capacity and service quotas.
  • Networking and Security: DNS, CIDR, private endpoints, firewall rules, allowlists, TLS/certificates, SPNs, API keys, tokens, image-pull credentials and Git-hosted secrets.
  • Operations and Automation: Linux, Git, CI/CD deployment workflows, infrastructure automation and scripting using Bash, Python or PowerShell.

What you'll do:

  • Review and act on incidents, service requests, infrastructure requests and provisioning failures logged by implementation teams and platform users.
  • Own L2/L3 cloud-platform issues from initial triage through recovery, validation, communication, root-cause analysis and closure.
  • Troubleshoot Kubernetes pods, deployments, replicas, services, events, health checks, node pools, scheduling and resource constraints.
  • Diagnose container image-pull, startup, shutdown, registry-authentication, rollout and workload-reconciliation failures.
  • Investigate CPU, memory, GPU, quota, region, placement and capacity issues affecting platform workloads.
  • Troubleshoot workload and event-driven autoscaling using Kubernetes metrics and technologies such as KEDA.
  • Support Azure resource provisioning, provider operations, resource lifecycle workflows and reconciliation between desired and actual state.
  • Diagnose ACR, private endpoint, firewall, allowlist, CIDR, DNS, TLS and runtime-connectivity problems.
  • Trace logs, metrics and distributed telemetry from the workload through collectors, exporters and observability ingestion pipelines.
  • Support Elastic/Logstash/Kibana ingestion, index mappings, access, dashboards, alerts and environment filters.
  • Investigate operational issues involving MongoDB/Atlas, SQL, Redis, NFS and related managed data or storage services.
  • Support certificates, service principals, API keys, image-pull credentials and infrastructure credential rotation.
  • Support regional releases, environment configuration, disaster-recovery workflows and post-deployment validation.
  • Develop automation to improve platform reliability, reduce manual provisioning and shorten incident recovery time.
  • Maintain runbooks, dashboards, alerts, known-error records and diagnostic procedures for use by platform support teams.
  • Participate in capacity planning, incident reviews, change reviews, release readiness and an agreed production on-call rotation.

What we are looking for:

  • Minimum 5-10 years of relevant work experience in Azure cloud infrastructure, DevOps, site reliability engineering or platform operations.
  • This role includes Rotational Shifts(Night shifts of 2 months in Year)
  • Hands-on production experience with Microsoft Azure services and operational troubleshooting
  • Strong Kubernetes troubleshooting skills covering workloads, events, services, networking, scheduling, scaling and resource management.
  • Experience with a container registry Azure Container Registry experience is highly relevant.
  • Practical observability experience across logs, metrics, traces, dashboards and alerts.
  • Experience with Elastic Stack components-Elasticsearch, Logstash and Kibana-or a closely comparable platform.
  • Working knowledge of DNS, TLS/certificates, CIDR, firewalls, proxies, load balancing and private networking.
  • Strong Linux administration, application log analysis and production incident-troubleshooting skills.
  • Scripting experience with Bash, Python or PowerShell for diagnostics and operational automation.
  • Experience troubleshooting CI/CD deployments, configuration changes and failed rollouts.
  • Working knowledge of Git, service identities and secrets-management fundamentals.
  • Working knowledge of MongoDB/Atlas, Redis, SQL or NFS operations is preferred.
  • Experience with infrastructure as code such as Terraform, Bicep or ARM templates and Kubernetes packaging such as Helm is preferred.
  • Experience with incident response, Root Cause Analysis, post-incident reviews and controlled production changes.
  • Strong collaboration and communication skills with the ability to work across application, security, networking and vendor teams.
  • Willingness to participate in a scheduled production on-call rotation and planned out-of-hours changes when required.

Our Values


If you want to know the heart of a company, take a look at their values. Ours unite us. They are what drive our success - and the success of our customers. Does your heart beat like ours Find out here:

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status.

About Company

Job ID: 153777759

Similar Jobs

Bengaluru, India

Skills:

JenkinsHadoopTerraformAnsibleOpenshiftSparkKafkaKubernetesGitOps

Bengaluru, India

Skills:

DASTKibanaPrometheusGrafanaPython ScriptingTerraformDockerSplunkKubernetesLinux BashDeployment Managercontainer image scanningCloud Monitoring APIDevSecOps practicesGitHub ActionsSBOM generationOpenTelemetrySASTGCP Billing APIsGCP services

Bengaluru, India

Skills:

Salesforce DevelopmentCloud TechnologiesJsonSqlELTGitData IntegrationAPEXData ModellingSoqlEtlSalesforce Data360Container OrchestrationSOAP APIsSandbox managementSecuritySalesforce CLIData unificationCI/CDIdentity resolutionChange set deployments

Bengaluru

Skills:

JenkinsGitMavenDockerKubernetesAEM DevOpsAEM as a Cloud ServiceAEM Cloud ManagerCI/CDAWS/Azure/GCP

Early Applicant
Bengaluru, India

Skills:

Elk StackOpenshiftDevopsGrafanaLinuxWindows ServerTerraformHelmUnixSplunkAzure CloudGithubDynatracePowerShellPrometheusKubernetesPythonDockerInfrastructure OperationsAzure MonitorCI CDShell Bash scriptingLog AnalyticsContainer PlatformsObservabilityInfrastructure AutomationPlatform Engineering

Beware of Scammers

We don’t charge money for job offers