Search by job, company or skills

AI Forward - Relability & Platform Engineer / SRE (Fully Remote)

4-6 Years
  • Posted 4 hours ago
  • Be among the first 10 applicants

Job Description

We are hiring a senior scalability engineer to help prepare a modern microservices backend for high-concurrency production usage. This role is focused on backend performance, reliability, capacity planning, async workflow design, database scalability, and observability.

This is not a generic DevOps-only role. We are looking for someone who can reason deeply about how backend systems behave under load: database pressure, queue depth, retries, rate limits, stuck jobs, async workflows, API latency, and failure recovery.

Responsibilities

  • Analyze backend architecture and identify scalability bottlenecks across APIs, services, database, cache, queues, storage, and third-party integrations.
  • Design and execute load tests for high-concurrency user flows.
  • Tune Postgres queries, indexes, connection pools, pagination, and transaction patterns.
  • Design reliable queue/background-job systems with retries, idempotency, dead-letter queues, and reconciliation jobs.
  • Define observability for scale: metrics, dashboards, request IDs, error tracking, alerts, and stuck-state monitoring.
  • Improve reliability of webhook-driven and async processing workflows.
  • Partner with backend engineers to implement code-level and architecture-level scalability improvements.
  • Recommend infrastructure and autoscaling changes based on measured bottlenecks.

Required Skills

  • 4+ years of backend, platform, SRE, performance engineering, or scalability engineering experience.
  • In-depth system design knowledge is of paramount importance.
  • Strong understanding of high-concurrency backend systems and distributed systems failure modes.
  • Deep Postgres experience: query plans, indexing, connection pools, slow queries, transaction design, and pagination.
  • Experience with Kafka, Redis, queues, background workers, retry/backoff, DLQs, and idempotent processing.
  • Hands-on load testing experience with k6, Artillery, Locust, JMeter, or similar.
  • Strong observability experience with logs, metrics, dashboards, alerting, and error tracking.
  • Cloud/container familiarity with AWS, Docker, ECS/Fargate, Kubernetes, or similar.
  • Ability to work with backend application code and collaborate with engineering teams.

Nice To Have

  • Production Node.js or NestJS experience.
  • TypeORM or ORM performance tuning experience.
  • DevOps/platform experience deploying and operating applications at scale.
  • Experience with object storage, upload-heavy systems, media processing, or webhook-heavy workflows.
  • OpenTelemetry or distributed tracing experience.
  • Terraform, Pulumi, AWS CDK, or infrastructure-as-code experience.

What Success Looks Like

  • We know what breaks first at 10x and 100x traffic.
  • Critical backend flows have load-test baselines and dashboards.
  • Database and queue bottlenecks are identified and mitigated.
  • Async workflows are retryable, idempotent, and observable.
  • The team has clear production scalability priorities before launch.

Employment Details

  • Role: Senior Scalability / Platform Engineer
  • Location: Remote only
  • Seniority: Senior / Lead-level preferred

More Info

Job Type:
Industry:
Function:
Employment Type:

About Company

Job ID: 153584993

Beware of Scammers

We don’t charge money for job offers