We're looking for a coding-strong Platform Engineer Lead who can build internal platform capabilities (tooling/services/standards) and also drive reliability practices (telemetry, alert hygiene, incident readiness). You'll collaborate with client engineering teams, review/shape technical designs (RFCs), mentor engineers, and deliver platform outcomes that improve developer velocity and operational confidence.
Responsibilities
- Build and evolve internal infrastructure platforms that improve developer experience and delivery flow.
- Drive observability/telemetry practices (e. g., Prometheus/OpenTelemetry) and enable adoption patterns across services.
- Improve incident management maturity: drills, response loops, learning loops, and alerting hygiene.
- Review and influence technical designs/RFCs with reliability, operability, and scalability thinking.
- Work across cloud-native ecosystems (Kubernetes/Nomad, etc. ) depending on project needs.
- Mentor engineers through pairing/mob sessions; contribute to strong engineering practices and documentation.
Requirements
- Strong software engineering foundation: you've built platform tooling/services in code (not only pipelines/IaC).
- Experience building/operating reliable platforms or systems at a meaningful scale.
- Strong grasp of observability (Prometheus / OpenTelemetry) and practical instrumentation patterns.
- Experience with incident response, drills, and improving alert quality and infrastructure.
- Cloud and infrastructure networking, Linux, networking, and exposure to orchestration environments.
- IaC experience with Pulumi/CloudFormation/CDK. CloudFormation communication and writing (design notes, runbooks, platform docs).
This job was posted by Lavanya Priyadarshini from MVPRockets.