ROLE SUMMARY
We are looking for an experienced Senior AI System Architect who can design scalable, secure, and production-grade AI-powered platforms from the ground up. This role goes beyond traditional cloud architecture — the candidate must be equally fluent in modern AI system design, including LLM infrastructure, Retrieval-Augmented Generation (RAG) pipelines, agentic workflow orchestration, and vector search systems.
The primary focus will be architecting — an enterprise-grade AI-powered customer support system — along with other AI product initiatives. You will own end-to-end architecture decisions across the application, AI inference, data, integration, cloud infrastructure, and operational reliability layers, while guiding engineering teams to build and evolve these systems.
KEY RESPONSIBILITIES
- Architect end-to-end AI-powered systems including LLM inference pipelines, RAG architectures, multi-agent orchestration, and conversational AI backends for production scale.
- Design vector search and knowledge management infrastructure — selecting vector databases (OpenSearch, Pinecone, Weaviate, pgvector), chunking strategies, and embedding pipelines.
- Define LLM serving patterns: model routing, fallback chains, streaming responses, prompt caching, token budgeting, and latency optimization.
- Architect cloud-native solutions on AWS using services such as EC2, ECS/EKS, Lambda, S3, RDS, DynamoDB, API Gateway, SQS/SNS, Amazon Bedrock, SageMaker, CloudWatch, IAM, and VPC.
- Own architecture across application, infrastructure, data, integration, and deployment layers; build design documents, trade-off analyses, and technical roadmaps.
- Define standards for microservices, CI/CD, observability, resiliency, disaster recovery, and AI governance (prompt injection prevention, PII handling, content filtering, audit logging).
- Evaluate technology choices and make recommendations based on scalability, performance, security, cost, and maintainability; drive modernization and cloud migration initiatives.
- Guide and mentor engineering teams through implementation; review architectures, ensure standards are followed, and troubleshoot complex production issues.
- Work closely with product, DevOps, security, and business stakeholders to translate requirements into technical designs, and communicate decisions to both technical and non-technical audiences.
REQUIRED SKILLS
- Proven experience designing production GenAI systems: RAG pipelines, LLM orchestration layers, agentic workflows, and vector search infrastructure.
- Hands-on AWS architecture experience across core and AI/ML services: Amazon Bedrock (Knowledge Bases, Agents), SageMaker inference endpoints, Lambda-based ML serving, and standard AWS compute, storage, networking, and security services.
- Strong understanding of distributed systems, microservices, REST/gRPC APIs, and event-driven architecture.
- Experience with Kubernetes, Docker, CI/CD pipelines, and infrastructure-as-code (Terraform or CloudFormation).
- Familiarity with vector databases (OpenSearch, Pinecone, Weaviate, FAISS, or pgvector) and LLM infrastructure patterns: streaming, batching, context window management, and cost optimization.
- Knowledge of AI-specific security and compliance: PII/PHI handling in AI pipelines, prompt injection prevention, content filtering, guardrails, and compliance logging.
- Strong knowledge of SQL and NoSQL databases, cloud security (IAM, VPC, encryption, secrets management), and production observability.
- Excellent communication skills — ability to articulate architectural decisions to both engineering teams and non-technical stakeholders.
PREFERRED SKILLS
- AWS Solutions Architect certification (Associate or Professional); AWS ML Specialty or AI Practitioner is a strong plus.
- Hands-on experience with LLM frameworks such as LangChain, LangGraph, or LlamaIndex; familiarity with agentic patterns and MCP (Model Context Protocol).
- Prior experience designing conversational AI, customer support AI, or enterprise virtual assistant platforms.
- Exposure to MLOps practices: model versioning, LLM response A/B testing, experiment tracking, and CI/CD for AI models.
- Experience with high-availability, multi-region, or large-scale enterprise systems; DevOps/SRE practices; performance and cloud cost optimization.
- Working knowledge of Python, Go, Java, or similar backend technologies sufficient to guide and review implementation.