

Search by job, company or skills

About TrueFoundry
Every production AI system, whether it's powering customer support, writing code, analyzing financial data, or diagnosing medical conditions, needs the same foundational infrastructure.A way to route between models. A way to manage tools and integrate them securely. A way to orchestrate agents and enforce governance. A unified compute layer to run it all
.That infrastructure layer is being built right now
.We are looking for a Senior SRE/DevOps Engineer to join the team
.The Problem We're Solvin
gCompanies are moving beyond simple chatbots to production agentic systems. These systems route between OpenAI, Anthropic, Google, and self-hosted models. They integrate dozens of tools via protocols like MCP. They orchestrate multi-agent workflows where agents coordinate with other agents
.The infrastructure to support this doesn't exist yet. You can't just duct-tape together a few API calls and call it production-ready
.You need a control plane that handles
sAI Gateway is the control plane: five composable components (Prompts, LLM Gateway, MCP Gateway, Guardrails, Agent Gateway) that handle routing, orchestration, and governance
.We're Series A, backed by Intel Capital and Sequoia. Companies like CVS, Mastercard, Siemens, Paytm, Synopsys, and Zscaler run production AI workloads on our platform
.Roles / Responsibilities
.Requirement
s*** Experience with Golang or Python is a must.*
sPerks of Working at TrueFoundr
Job ID: 151772365
Skills:
Kubernetes, Golang, Slas, Networking, Datadog, Microservices, Docker, routing, Vpn, shell scripting, AWS, Databases, SSL, Sql, Web Technologies, Cloud Infrastructure, Http, Containers, Terraform, Azure, SLIs, No-SQL, Key Value, web sockets, error budgets, SLOs
Skills:
AWS EKS, Golang, Elk, Shell Scripts, Grafana, Zabbix, Jenkins, Terraform, Ansible, Networking Protocols, Python, Monitoring stacks, GitLab CI, TICK, SIP, ArgoCD
Skills:
Incident Management, Distributed Systems, Forecasting, Cloud cost optimization, AWS billing analysis, FinOps, Reliability engineering, Infrastructure-as-code, Operational Excellence, Budget Tracking, Production Support, Linux systems knowledge, Tagging, cost allocation
Skills:
Yaml, Devops, microsoft power automate, Microsoft Azure, Kubernetes, Infrastructure as Code, Azure DevOps Services, Kusto Query Language, Scripting and Automation
Skills:
Java, Unix, Elk, Nodejs, Grafana, Autosys, Shell, Cloudwatch, Linux, Restful Services, Python, AWS, Logging, Alerting, EFK, PagerDuty, SLOs, Monitoring, Oracle FCMRE