AI Infrastructure Engineer (Backend Platform)
Experience: 5–8 Years
Employment Type: Full-time
About the Role
We are looking for an experienced AI Infrastructure Engineer to build and scale the backend systems that power production-grade AI applications. You will work on high-performance, distributed systems that orchestrate AI workloads, optimize inference performance, and deliver reliable, low-latency conversational experiences for users at scale.
This role is ideal for engineers who have built scalable backend platforms, integrated Large Language Models (LLMs) into production systems, and have hands-on experience with asynchronous architectures, real-time communication, and distributed systems.
Key Responsibilities
AI Infrastructure & Orchestration
- Design and implement asynchronous, event-driven AI orchestration systems.
- Build and optimize multi-agent workflows for complex AI interactions.
- Own end-to-end latency from user request to AI response.
- Develop resilient AI inference pipelines with retries, circuit breakers, and graceful fallback mechanisms.
- Implement intelligent request routing and load balancing across multiple AI models and providers.
- Build scalable services for AI conversation orchestration and migrate monolithic components into distributed microservices where required.
Backend Platform Engineering
- Design and develop highly scalable backend services capable of handling high concurrent traffic.
- Build real-time communication infrastructure using technologies such as WebSockets or Server-Sent Events (SSE).
- Optimize backend performance through caching, asynchronous processing, and efficient data retrieval.
- Implement distributed messaging using Kafka, RabbitMQ, or similar event-streaming technologies.
AI Platform & Performance
- Integrate production-grade LLM APIs and model serving platforms.
- Optimize inference performance through batching, response caching, prompt optimization, and context management.
- Design conversation state management for multi-turn AI interactions.
- Implement intelligent fallback strategies across multiple AI providers.
Reliability & Observability
- Build monitoring and observability solutions for latency, throughput, error rates, and AI performance metrics.
- Monitor production systems using tools such as Grafana, Prometheus, Datadog, or OpenTelemetry.
- Troubleshoot production issues related to AI inference, distributed systems, and backend scalability.
Required Skills & Experience
Experience
- 3–5 years of experience building scalable backend systems.
- Proven experience developing production systems serving high concurrent user traffic (10,000+ concurrent users preferred).
- Experience designing distributed systems and microservices architectures.
Backend Technologies
- Strong programming skills in one or more of the following:
- Python
- Go
- Java
- Node.js
- Experience with backend frameworks such as FastAPI, Spring Boot, Express/NestJS, or equivalent.
- Strong understanding of asynchronous programming concepts.
Distributed Systems
- Hands-on experience with event-driven architectures.
- Experience using Kafka, RabbitMQ, AWS SQS, Google Pub/Sub, NATS, or similar messaging platforms.
- Strong understanding of distributed system design principles.
AI & LLM Integration
- Experience integrating Large Language Models into production applications.
- Hands-on experience with one or more of the following:
- OpenAI
- Anthropic Claude
- Google Gemini
- Azure OpenAI
- vLLM
- Triton Inference Server
- TensorFlow Serving
- Experience implementing retry logic, rate limiting, fallback strategies, and AI provider failover.
- Understanding of prompt optimization, inference latency, batching, and conversation context management.
Caching & Performance
- Strong experience with Redis for caching, session management, and rate limiting.
- Experience optimizing backend latency and throughput.
Real-Time Systems
- Experience building streaming applications using:
- WebSockets
- Server-Sent Events (SSE)
- Event streaming platforms
Observability
Experience with one or more of:
- Grafana
- Prometheus
- Datadog
- OpenTelemetry
Good to Have
- Experience building multi-agent AI systems using LangChain, LangGraph, CrewAI, or custom orchestration frameworks.
- Experience implementing Retrieval-Augmented Generation (RAG) pipelines.
- Experience migrating monolithic applications to microservices.
- Exposure to Kubernetes and Docker.
- Experience working in cloud environments such as AWS, Azure, or GCP.
- Background in fintech, payments, or other high-scale consumer platforms.
- Startup experience with ownership of end-to-end product features.
What We're Looking For
- Strong problem-solving and system design skills.
- Passion for building scalable, resilient backend platforms.
- Ability to work independently in a fast-paced environment.
- Strong communication and collaboration skills.
- Ownership mindset with a focus on performance, reliability, and continuous improvement.
Why Join Us
- Build next-generation AI infrastructure powering real-world applications.
- Solve complex distributed systems and scalability challenges.
- Work with cutting-edge AI technologies and modern cloud-native architectures.
- Take ownership of mission-critical systems with significant business impact.
- Collaborate with a high-performing engineering team focused on innovation and technical excellence.