Search by job, company or skills

Founding Engineer, AI Platform Reliability & Scale

Founding Engineer, AI Platform Reliability & Scale

Stealth Startup
4-6 Years
Not Disclosed
  • Posted 20 hours ago
  • Be among the first 10 applicants

Job Description

Founding Engineer , AI Platform Reliability and Scale

Company: Early-Stage Silicon Valley AI Startup (Ex-Apple & Ex-Meta Founders)

Headquarters : San Francisco Bay area, California.

Role Type: Full-Time (Founding Stage)

About Us

We're an early-stage Silicon Valley AI startup founded by ex-Apple and ex-Meta engineers. Our platform allows anyone to describe an application and have AI build, host, and run it—with zero code to maintain.

We are live with real users, growing fast, and looking for our next Founding Engineer to make the platform more reliable, faster, and cheaper to run as we scale. If you love finding the root cause of complex systems issues and ensuring they never happen again, let's talk.

What You'll Work On

  • Reliability: Make AI-generated apps work correctly the first time and keep them running smoothly.
  • Scale: Scale our infrastructure, data layer, and pipelines to match rapid user growth.
  • Performance & Cost: Drive down AI generation latency and cost without compromising quality.
  • Observability: Build deep observability systems that pinpoint what is breaking and why.
  • Continuous Learning: Turn every system failure into a learning loop for the platform.

How You Think

  • Systems First: You trace a problem across multiple layers to its actual root cause and fix it there.
  • Prevent, Don't Patch: You prefer making a failure impossible over adding temporary band-aids.
  • Data Before Opinions: You measure the problem, implement the fix, and prove the metrics moved.
  • Total Ownership: You identify what matters most, tackle unassigned problems, and drive them to production.
  • AI-Native: You use Claude and coding agents daily, reviewing their output with a critical engineering eye.

What We're Looking For

  • 4+ years of experience building and running backend or distributed systems in production.
  • End-to-End Ownership: You have owned a production system end-to-end, including leading incident response and post-mortems.
  • AI Experience: Built or worked on LLM-powered or agentic systems deployed to real users.
  • Tech Stack: Strong Python proficiency, alongside solid fundamentals in databases, message queues, and cloud infrastructure.
  • Observability: Hands-on comfort with logs, metrics, traces, and alerting.
  • Pragmatic Judgment: Clear thinking about design patterns, trade-offs, and knowing when not to over-engineer.
  • Hands-on Coding: You still write code every day and genuinely enjoy it.

Bonus Points:

  • Experience with workflow engines, LLM evals, or model fine-tuning.
  • Background as an early engineer at a scaling startup or a previous founder.

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

Traces

Logs

Alerting

Message Queues

About Company

Similar Jobs

4-6 yrs
Bengaluru, India
Skills:
Integrations, Apis, concurrency, Cloud Operations, Authentication, operational tooling, billing integrations, asynchronous jobs, incident diagnosis, retries, data models, Relational Databases, performance observability, AI coding tools, queues, Access Control, usage tracking
7-9 yrs
Hyderabad, India
Skills:
Gcp, Docker, PostgreSQL, Prometheus, FastAPI, Grafana, Azure, Kubernetes, Redis, Python, AWS
5-9 yrs
Ahmedabad, India
Skills:
Kafka, Gcp, Sqs, Azure, Python, AWS, LLMs, AI Agents, Agentic Workflows, LangGraph, RAG, Google ADK, LlamaIndex
3-5 yrs
Kolkata, India
Skills:
Sql, ELT, Etl, AWS, Python, Azure, Gcp, data processing systems, LLMs, LlamaIndex, AI agents, data pipelines, backend software development, LangChain, LangGraph
3-5 yrs
India
Skills:
Python, Flutter, Node.js, React, Next.js