Search by job, company or skills

Senior Site Reliability Engineer

Early Applicant
  • Posted 11 days ago
  • Be among the first 20 applicants

Job Description

About TrueFoundry


Every production AI system, whether it's powering customer support, writing code, analyzing financial data, or diagnosing medical conditions, needs the same foundational infrastructure.A way to route between models. A way to manage tools and integrate them securely. A way to orchestrate agents and enforce governance. A unified compute layer to run it all

.That infrastructure layer is being built right now

.We are looking for a Senior SRE/DevOps Engineer to join the team

.The Problem We're Solvin

gCompanies are moving beyond simple chatbots to production agentic systems. These systems route between OpenAI, Anthropic, Google, and self-hosted models. They integrate dozens of tools via protocols like MCP. They orchestrate multi-agent workflows where agents coordinate with other agents

.The infrastructure to support this doesn't exist yet. You can't just duct-tape together a few API calls and call it production-ready

.You need a control plane that handles

  • :Intelligent routing with observability, cost policies, and fallback logi
  • cCentralized tool and MCP server management with security and lifecycle control
  • sAgent orchestration with governance and guardrail
  • sA unified compute layer to run self-hosted models, custom tools, and agent

sAI Gateway is the control plane: five composable components (Prompts, LLM Gateway, MCP Gateway, Guardrails, Agent Gateway) that handle routing, orchestration, and governance

.We're Series A, backed by Intel Capital and Sequoia. Companies like CVS, Mastercard, Siemens, Paytm, Synopsys, and Zscaler run production AI workloads on our platform

.Roles / Responsibilities

  • :Write Terraform modules for deploying different components of infrastructure in AWS, like Kubernetes, RDS, Prometheus, Grafana, and Static Websit
  • eThe SRE will work closely with TrueFoundry customers, gaining a deep understanding of the TrueFoundry platform to ensure smooth deployments, reliable operations, and best practices adoption. This role will also involve training and onboarding new customers, assisting them in implementing TrueFoundry effectively, and helping drive platform adoption and operational excellence across customer teams
  • .Configure networking and autoscaling. continuous deployment, security, and multiple environment
  • sMake sure the infrastructure is SOC2, ISO 27001, and HIPAA complian
  • tAutomate all the steps to provide a seamless experience to developers

.Requirement

s*** Experience with Golang or Python is a must.*

  • *4-8 years of work experience writing clean production cod
  • eWell-versed in maintaining infrastructure as code (Terraform, CloudFormation, etc). High proficiency with Terraform / Terragrunt is absolutely critica
  • lExperience in setting up CI/CD pipelines from scratc
  • hExperience with ETL pipelines, Bigdata infr
  • aUnderstanding of common security issue

sPerks of Working at TrueFoundr

  • yJoin a fast-growing Series A, Bay Area-based startup building cutting-edge AI infrastructure
  • .Comprehensive health insurance for you and your family
  • .Flexible hybrid work, 3 days a week in the office (Tuesday, Wednesday and Thursday), with flexibility around the schedule
  • .Lunch and snacks are on us whenever you come into the office
  • .Work from another TrueFoundry office for up to 3 months each year, whether that's our London or Bay Area office, if you'd like a change of scenery

.

More Info

Job Type:
Industry:
Function:
Employment Type:

About Company

Job ID: 151772365

Similar Jobs

Bengaluru, India

Skills:

KubernetesGolangSlasNetworkingDatadogMicroservicesDockerroutingVpnshell scriptingAWSDatabasesSSLSqlWeb TechnologiesCloud InfrastructureHttpContainersTerraformAzureSLIsNo-SQLKey Valueweb socketserror budgetsSLOs

Bengaluru, India

Skills:

AWS EKSGolangElkShell ScriptsGrafanaZabbixJenkinsTerraformAnsibleNetworking ProtocolsPythonMonitoring stacksGitLab CITICKSIPArgoCD

Bengaluru, India

Skills:

Incident ManagementDistributed SystemsForecastingCloud cost optimizationAWS billing analysisFinOpsReliability engineeringInfrastructure-as-codeOperational ExcellenceBudget TrackingProduction SupportLinux systems knowledgeTaggingcost allocation

Bengaluru, India

Skills:

YamlDevopsmicrosoft power automateMicrosoft AzureKubernetesInfrastructure as CodeAzure DevOps ServicesKusto Query LanguageScripting and Automation

Bengaluru, India

Skills:

JavaUnixElkNodejsGrafanaAutosysShellCloudwatchLinuxRestful ServicesPythonAWSLoggingAlertingEFKPagerDutySLOsMonitoringOracle FCMRE

Beware of Scammers

We don’t charge money for job offers