Search by job, company or skills

Principal AI Architect

Principal AI Architect

Trianz
8-10 Years
Not Disclosed
  • Posted 3 hours ago
  • Be among the first 10 applicants

Job Description

ABOUT THIS ROLE

You will own the complete architecture for private and sovereign AI deployment at Trianz. This is not a cloud-API-consumption role. You will design the system that runs large language models inside customer environments -- choosing the right models, designing the serving topology, setting CPU and GPU routing policies, and ensuring the entire stack is secure, sovereign, and provider-agnostic. This is a pure IC role with architectural authority over how AI runs in production.

WHAT YOU WILL DO

  • Design the complete private LLM serving architecture: model selection, serving framework (vLLM,TensorRT-LLM, Triton), and runtime topology.
  • Define intelligent CPU vs GPU routing policies: which model sizes and prompt types route to which compute tier based on latency, cost, and throughput targets.
  • Architect multi-cloud model serving: AWS, Azure, GCP -- provider-agnostic, no managed AI lock-in.
  • Design on-premises serving for enterprise customers: RedHat OpenShift, VMware -- containerised model serving on customer hardware.
  • Architect per-tenant model isolation, data-residency compliance, and air-gapped sovereign deployment patterns.
  • Define the release architecture for model versions: rollout, staged deployment, rollback, and promotion gates.
  • Design the LLM governance framework: model behaviour monitoring, inference audit logging, guardrail architecture.
  • Design the DevSecOps pipeline architecture and self-service deployment automation standards.
  • Evaluate open-source models (Llama, Mistral, Qwen, Phi) against closed models for specific enterprise use cases.
  • Produce architecture sign-off documents and review all AI system designs before implementation.

MUST HAVE

  • Deployed LLMs to production in a real enterprise environment -- not just API consumption.
  • Hands-on with vLLM, TensorRT-LLM, or Triton in a production serving context
  • Designed CPU cluster inference (Intel Xeon / AMD EPYC) for open-source models.
  • Kubernetes at production scale (EKS, AKS, GKE, or OpenShift) -- not just local k8s.
  • Designed multi-cloud, provider-agnostic AI architectures.
  • Experience with model quantization (GPTQ, AWQ, GGUF) and routing trade-offs
  • 8+ years in AI/software architecture with at least 3 years in production LLM systems.

GOOD TO HAVE

  • Experience with RedHat OpenShift for on-premises model serving.
  • Familiarity with NVIDIA Dynamo, SGLang, or custom inference schedulers.
  • GPU FinOps -- cost modelling for mixed CPU/GPU inference fleets.
  • Sovereign AI or regulated-industry deployment experience.
  • Intel OpenVINO or AMD ROCm for CPU-optimised inference.

Company Overview

Trianz is an applied AI solutions company that accelerates customer business transformation through AI powered Transformation Services as a Software Model. With 25+ years of transforming enterprises, we've evolved to a product-led, platform-driven organization serving global enterprises across Financial Services, Insurance, Healthcare, Hi-Tech, Manufacturing, and other industries.

With global presence across 4 continents, our platform portfolio under the unified Concierto brand delivers end-to-end transformations including solutions for Migrate, Manage, Maximize, Modernize, Insights & Agentic AI, and SecOps - delivered through strategic partnerships with leading hyperscalers.

We're building the premier innovation-led organization in the digital transformation space through AI-first methodologies and data-driven excellence - RevolutionAIzing Transformations.

More Info

Key Skills

GPTQ

SGLang

Multi-cloud provider-agnostic AI architectures

AMD ROCm

GGUF

RedHat OpenShift

Custom inference schedulers

Model quantization

AWQ

Intel OpenVINO

TensorRT-LLM

LLM serving architecture

NVIDIA Dynamo

GKE

GPU FinOps

AMD EPYC

AKS

vLLM

EKS

Intel Xeon

CPU cluster inference

Routing trade-offs

About Company