
Search by job, company or skills
AWS · TypeScript · Reliability · Security · Scale
Reliability, scale, security, and cost are all yours.Otto builds an AI computer.
Our hardware sits on a customer's desk running embedded Linux, maintains a live connection to our cloud, and receives signed over-the-air updates from us.
Every part of that path runs on AWS. Our infrastructure is infrastructure-as-code, and our stack is heavily TypeScript.
You will design and operate the platform behind Otto: cloud infrastructure, CI/CD, observability, fleet reliability, security, capacity planning, and cost.
You will also write code, review PRs, deploy to production, respond to incidents, and participate in on-call.
This is a small team. There is no infrastructure layer between you and the system.
Infrastructure as Code
Compute
Data
Delivery
Observability
Nearby Stack
Some of our web surfaces run on Vercel. Experience with it is helpful, but not required.
What You'll OwnProduction Infrastructure on AWSDesign, build, and operate our production infrastructure across Lambda, Fargate, RDS, ElastiCache, DynamoDB, S3, CloudFront, API Gateway, VPC, and IAM.
You should be comfortable owning the system end to end rather than relying on a separate platform team.
Infrastructure as CodeEvery resource lives in our CDK tree.
Development and production environments should be reproducible from the same code, with infrastructure changes reviewed through the same PR process as application code.
CI/CDBuild and maintain GitHub Actions pipelines for:
You should know when each deployment strategy is appropriate and why.
ObservabilityOwn metrics, logs, traces, dashboards, and alerts across the platform.
The goal is simple: page a human only when a human is actually needed.
SLIs, SLOs, and Error BudgetsDefine reliability targets, track them, and help the product and engineering teams make informed decisions when reliability budgets are being consumed.
Incident ResponseLead incidents when they happen.
Run blameless postmortems and turn findings into real reliability work rather than documentation that gets forgotten.
OTA Updates to the Otto FleetOur releases eventually reach physical devices sitting in customers homes and offices.
You will help own:
A bad release can reach hardware we cannot physically access, so deployment discipline matters.
AWS CostOwn infrastructure efficiency and visibility.
This includes:
Help establish and maintain a strong production security posture, including:
Operate PostgreSQL, Redis, and DynamoDB in production.
That includes:
If something repetitive can be automated, automate it.
Build internal tooling in TypeScript or Bash that allows engineers to self-serve instead of relying on manual infrastructure work.
MentoringHelp raise the operational and reliability bar across the entire engineering team.
Must HaveThese are not familiarity with requirements.
We are looking for someone who has owned these systems in production and understands where they fail.
AWS CDKDeep, current experience with AWS CDK v2 in TypeScript, including:
You should know what to do when the synth is clean but the deployment still isn't.
AWS in ProductionStrong hands-on experience with:
We are looking for direct ownership, not experience where another platform team handled the difficult parts.
AWS LambdaYou should understand:
Experience actually operating PostgreSQL in production, including:
Strong understanding of DynamoDB data modeling, including:
You should write real production TypeScript, not just infrastructure glue.
Our infrastructure, backend, internal tooling, and device runtime are all heavily Node.js and TypeScript.
You will regularly read and review code outside of a traditional infrastructure role.
Linux and NetworkingStrong understanding of:
You should be comfortable debugging problems like a persistent socket that only drops for customers on one ISP.
CI/CD and Incident ResponseYou should have experience:
These are not required, but they count for a lot.
Otto is a small team with a short path from decision to production.
You will work in the same repositories as the rest of the engineering organization and regularly read code outside your lane.
Our systems span cloud infrastructure, APIs, data infrastructure, web applications, and physical devices running in the field.
We value clear technical writing.
Design documents, postmortems, architecture decisions, and the sentence in a PR explaining why something is changing all matter.
If a deployment changes something persisted on an Otto device that has been running in someone's home for six months, we want that risk understood and documented before the release goes out.
On-call is shared and real.
In return, fixing the system that woke you up is treated as the work — not as an interruption from the work.
Job ID: 153799117
Skills:
Mariadb, Cloud Services, Computer Networking, Linux Unix administration, Open source technologies, Container orchestration technology, Galera cluster operations, Private cloud environments, MySQL databases, Percona
Skills:
scalability testing , Prometheus, Grafana, Datadog, Microservices, Docker, Distributed Systems, Python, Azure DevOps, AWS, Performance Tuning, PowerShell, Bash, Gcp, Splunk, Rest Apis, Kubernetes, Go, API monitoring, dependency management, GitHub Actions, observability, reliability best practices, Monitoring, production support operations, Troubleshooting, Performance Analysis
Skills:
S3, PowerShell, Bash, Vpc, AWS CloudWatch, Datadog, Jenkins, Ec2, Docker, Terraform, ECS, AWS CloudFormation, Route 53, Python, AWS, Fargate, AWS WAF, AWS GuardDuty
Skills:
PostgreSQL, Itil, Azure, Oracle, Kubernetes, Datadog, AWS, PagerDuty
Skills:
Gcp, Azure, Kubernetes, Logging, AWS, Alerting, tracing, Monitoring