Search by job, company or skills

Pre-Sales - Generative & Agentic AI Solutions Architect

Pre-Sales - Generative & Agentic AI Solutions Architect

host360
6-10 Years
Not Disclosed
  • Posted 5 hours ago
  • Be among the first 10 applicants

Job Description

Job Title: Pre-Sales - Generative & Agentic AI Solutions Architect

Company: Host360

Location: Onsite (Andheri East)

Employment Type: Full Time

Experience: 6–10 Years

Function: Solutions Architecture / Technical Pre-Sales

Focus: Generative AI, Agentic AI, RAG, LLM Inference & GPU-Accelerated Computing

CTC – As per Industry Standards

Role Overview

We are looking for a technically strong and customer-focused Generative & Agentic AI Solutions Architect – Pre-Sales to design, validate, and position enterprise AI solutions across Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), Agentic AI, and GPU-accelerated computing.

You will work closely with customers, Sales, Business Development, Engineering, and technology partners to understand business and technical requirements, develop end-to-end architectures, size AI infrastructure, build and validate PoCs, and guide customers from initial discovery through technical closure and production adoption.

The role requires a strong combination of AI technical depth, solution architecture, hands-on problem solving, and customer-facing pre-sales capability.

Key Responsibilities

  • Engage directly with customers to understand their business challenges, AI use cases, data landscape, workloads, performance requirements, and existing infrastructure.
  • Architect end-to-end Generative AI, LLM, RAG, and Agentic AI solutions aligned with customer requirements.
  • Lead technical discovery sessions, architecture workshops, whiteboarding sessions, and solution-design discussions.
  • Design RAG architectures covering data ingestion, chunking, embeddings, vector search, retrieval, reranking, context management, and enterprise data integration.
  • Design Agentic AI architectures involving tool/function calling, orchestration, memory, planning, APIs, guardrails, human-in-the-loop, and enterprise system integration.
  • Design and size scalable LLM inference platforms and GPU-accelerated infrastructure based on model, workload, concurrency, context length, latency, throughput, and availability requirements.
  • Analyze end-to-end AI workloads and recommend approaches to improve latency, throughput, scalability, GPU utilization, reliability, and cost efficiency.
  • Evaluate and recommend appropriate models, inference engines, AI frameworks, vector databases, and deployment architectures.
  • Build or lead PoCs, technical demonstrations, benchmarks, and reference solutions to validate proposed architectures and customer use cases.
  • Support Sales and Business Development teams throughout the technical sales cycle, including customer presentations, solution positioning, demonstrations, technical qualification, and objection handling.
  • Translate customer requirements into solution architectures, infrastructure sizing, BOMs, technical proposals, RFP/RFI responses, and implementation scope/SOWs.
  • Identify technical risks, dependencies, assumptions, security requirements, and integration challenges before solution commitment.
  • Ensure proposed solutions are technically feasible, scalable, secure, supportable, and commercially practical.
  • Collaborate with Engineering and Delivery teams to ensure effective transition from pre-sales to implementation and production deployment.
  • Collaborate with NVIDIA and other technology partners for architecture validation, GPU sizing, performance optimization, product feedback, and complex customer requirements.
  • Provide technical leadership and guidance on best practices for production-grade Generative AI and Agentic AI deployments.
  • Work with Enterprise, Government, Research, and Public Sector customers on strategic AI initiatives.
  • Stay current with developments in LLMs, Agentic AI, RAG, inference engines, GPU technologies, and distributed AI systems.

Must-Have Skills & Experience

  • 6–10 years of overall experience across Solution Architecture, Technical Pre-Sales, AI/ML, Cloud, Data Platforms, Infrastructure, or related technologies.
  • Strong customer-facing experience in a Solution Architect, Pre-Sales Architect, AI Architect, or Technical Consultant role.
  • Strong understanding of Generative AI, LLMs, RAG, and Agentic AI architectures.
  • Hands-on experience designing or delivering LLM, RAG, or Agentic AI solutions in production or production-like environments.
  • Ability to translate business requirements into end-to-end technical architectures and solution designs.
  • Practical understanding of RAG components including embeddings, vector databases, semantic/hybrid search, retrieval, reranking, context management, and evaluation.
  • Understanding of Agentic AI concepts including tool/function calling, orchestration, memory/state, planning, multi-step workflows, guardrails, and human-in-the-loop patterns.
  • Strong understanding of LLM inference concepts including:
  1. Latency and throughput
  2. Concurrency and batching
  3. GPU memory and utilization
  4. KV cache
  5. Quantization
  6. Multi-GPU deployment
  7. Scalability and cost optimization
  • Experience with Python, Linux, Docker, and Kubernetes.
  • Experience with AI/ML ecosystems such as PyTorch and Hugging Face.
  • Experience or exposure to LLM serving technologies such as NVIDIA NIM, TensorRT-LLM, Triton Inference Server, vLLM, or equivalent.
  • Working understanding of GPU infrastructure, model sizing, networking, storage, security, and observability requirements for AI workloads.
  • Experience conducting customer discovery, architecture workshops, technical presentations, demonstrations, and PoCs.
  • Experience supporting technical pre-sales activities including solution sizing, RFP/RFI responses, technical proposals, and technical closure.
  • Strong written and verbal communication skills with the ability to engage developers, architects, infrastructure teams, business stakeholders, and senior customer leadership.

Preferred / Good-to-Have Skills

  • Experience with NVIDIA AI Enterprise, NIM, NeMo, TensorRT-LLM, Triton Inference Server, Dynamo, or other NVIDIA AI technologies.
  • Experience designing or operating GPU clusters and distributed AI infrastructure.
  • Understanding of advanced inference architectures including distributed/disaggregated serving, intelligent routing and scheduling, KV-cache optimization, heterogeneous inference, and model parallelism.
  • Experience with LangGraph, LlamaIndex, LangChain, Semantic Kernel, or similar Agentic AI frameworks.
  • Experience with model customization, LoRA/PEFT, fine-tuning, quantization, and model evaluation.
  • Experience with Kubernetes GPU scheduling and production AI platform architecture.
  • Experience designing AI solutions across on-premises, private cloud, and public cloud environments.
  • Understanding of AI security, governance, observability, responsible AI, and guardrail frameworks.
  • Experience with GPU/infrastructure sizing, BOM preparation, capacity planning, and performance benchmarking.
  • Experience with high-performance networking and storage for AI workloads.
  • Experience working with Government, Research, Public Sector, CSP, or large enterprise customers.

The ideal candidate combines strong customer engagement and pre-sales capability with sufficient hands-on technical depth to design, validate, defend, and optimize enterprise AI solutions.

More Info

Key Skills

TensorRT-LLM

Triton Inference Server

Generative AI

Hugging Face

vLLM

Agentic AI

GPU-Accelerated Computing

NVIDIA NIM

About Company