Search Jobs

Search by job, company or skills

Platform Engineer

Platform Engineer

recrew ai
  • Posted 2 hours ago
  • Be among the first 10 applicants

Job Description

Role: Senior AI Platform Engineer

Function: Platform Engineering / AI-ML / MLOps

Location: Bangalore

Type: Full-time

Industry: Artificial Intelligence, Critical Infrastructure

About Company

A research-first AI company incubated at the Indian Institute of Science (IISc). The company is building advanced AI for the planning and operations of critical networks.

Its core technology is a World Model for critical networks — frontier AI, not another LLM wrapper. The founding team previously built and scaled a deep-tech company in the private 5G and cellular connectivity space.

Position Overview

The company's autonomous agents diagnose complex problems and plan actions across high-stakes infrastructure networks — and this role owns the platform that makes them run. As Senior Platform Engineer, you will build and own the core backend, data systems, and MLOps infrastructure that enable agents to operate reliably at scale, learn from every case, and remain fully observable and auditable. This is a deep-work IC role with direct reporting to the founders and outsized ownership over platform architecture from the ground up.

Role & Responsibilities

  • Own the agent serving and runtime path end-to-end, including model versioning, deployment pipelines, and rollback mechanisms using LangChain/LangGraph.
  • Build and maintain the tool and data integration layer connecting agents to external data sources, APIs, and domain-specific infrastructure systems.
  • Design and operate the learning and feedback pipeline — capturing agent outcomes, labeling signals, and closing the loop for continuous improvement.
  • Enforce platform governance: budget controls, rate limits, and approval gates for agent actions on critical infrastructure.
  • Build audit-grade tracing, observability, and logging systems (using Langfuse or equivalent) that meet real-world accountability requirements for autonomous action systems.
  • Own platform reliability, latency, and uptime against defined SLOs — including designing for retries, timeouts, and graceful degradation in distributed environments.
  • Collaborate closely with agent, data, and domain teams to ensure platform abstractions support evolving research and product requirements.

Must Have Criteria

  • 4+ years of backend engineering experience building and operating production-grade, scaled software systems end-to-end.
  • Hands-on experience operating production ML/AI platforms — model versioning, reproducible jobs, and safe rollout/rollback in live environments.
  • Strong data engineering skills: streaming pipelines, schedulers, SQL, object storage, and data provenance/quality practices.
  • Proficiency in async Python and FastAPI for building high-performance serving and API layers.
  • Experience with containerized, distributed deployments using Docker and Kubernetes on major cloud providers.
  • Demonstrated ability to build audit-grade logging and observability systems for systems that take real-world actions.
  • Distributed-systems fundamentals: retries, timeouts, circuit breakers, and graceful degradation under failure.

Nice to Have

  • Prior experience with LangChain, LangGraph, or similar agent orchestration frameworks.
  • Familiarity with Langfuse or other LLM tracing/observability tooling.
  • Experience serving open-source LLMs on-premises (local model serving infrastructure).
  • Master's degree in Computer Science, Engineering, or a related field.
  • Background in network infrastructure, telecom, or critical systems domains.

What We Offer

  • Founding-team-level ownership over platform architecture at a seed-stage, IISc-incubated AI company.
  • Direct collaboration with founders and exposure to frontier AI research at the intersection of World Models, LLMs, and critical infrastructure.
  • Deep-work engineering culture — small team, high autonomy, minimal process overhead.
  • Opportunity to shape engineering culture, tooling choices, and platform direction from zero.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

production ML AI platforms

reproducible jobs

model versioning

async Python

timeouts

quality practices

distributed-systems fundamentals

safe rollout rollback

data provenance

audit-grade logging

retries

schedulers

observability systems

streaming pipelines

graceful degradation

About Company

Similar Jobs

6-10 yrs
Bengaluru, India
Skills:
GitDockerPostgreSQLRest ApisAzureKubernetesPythonMicroservicesDevOps practicesLangfuse
4-6 yrs
Bengaluru, India
Skills:
TerraformDockerBashPythonObservability and monitoring toolsGoAPI Gateway technologiesEvent-driven architectures
5-7 yrs
Bengaluru, India
Skills:
PostgreSQLPrometheusGrafanaBash ScriptingDockerTerraformMicrosoft AzureAzure DevOpsPowerShellElk StackRedisRest ApisHelmAzure Key VaultARM Bicep templatesAzure RBACMicrosoft Defender for CloudAzure Cost ManagementApplication InsightsAzure CLIAzure PolicyCIS BenchmarksAzure MonitorAzure networking
7-9 yrs
Bengaluru, India
Skills:
JavaAws LambdaNode.jsMicroservicesTerraformDockerAmazon ConnectPythonKubernetesApi GatewayServerless ArchitectureDevOps tools and practicesCloud-native architecture principlesAmazon LexAWS Well-Architected Framework
2-4 yrs
Bengaluru, India
Skills:
react.js Amazon Web ServicesJava Programming LanguageAutomationAgile MethodologyMicroservicesDevopsSoftware DevelopmentSoftware EngineeringMicrosoft AzureScalabilityKubernetesAWSFull Stack DevelopmentPython Programming LanguageApplication Programming Interface APIDocker SoftwareJavaScript Programming LanguageSQL Programming LanguageCSA certifiedAngular Web Framework