Search by job, company or skills

Vision AI Solution Architect

Vision AI Solution Architect

Tata Consultancy Services
8-12 Years
Not Disclosed
Early Applicant
  • Posted 7 days ago
  • Be among the first 10 applicants

Job Description

Vision Solution Architect

Computer Vision, VLMs & Multimodal GenAI Architecture

Experience Level: 8-12 years

Employment Type: Full-Time

Location: Pune, Chennai, Hyderabad, Kolkata, Delhi, Bangalore

About the Role

We are looking for a Vision Solution Architect to design and lead end-to-end computer vision and multimodal AI solutions — spanning classical vision models, CLIP-style embedding models, Vision-Language Models (VLMs), and Vision RAG architectures. The ideal candidate can architect solutions across cloud hyperscalers and edge deployments, guiding fine-tuning, integration, and productionization of vision and multimodal GenAI systems for real-world business use cases.

Key Responsibilities

  • Architect end-to-end vision and multimodal GenAI solutions, from use-case discovery and model selection through deployment and monitoring.
  • Design solutions leveraging CLIP-style embedding models, VLMs (e.g., LLaVA, Gemini Vision, GPT-4V/Claude Vision, Qwen-VL), and traditional CV models (detection, segmentation, OCR).
  • Design and implement Vision RAG pipelines — multimodal embedding generation, vector indexing/retrieval, and grounding of VLM outputs against visual and textual knowledge bases.
  • Lead fine-tuning and domain adaptation of vision and vision-language models (LoRA/QLoRA, adapter-based tuning, contrastive fine-tuning) for domain-specific accuracy.
  • Architect solutions across hyperscaler vision/AI services — AWS (Rekognition, Bedrock multimodal models, SageMaker), GCP (Vertex AI Vision, Gemini multimodal APIs), and Azure (AI Vision, Azure AI Foundry) — selecting the right managed vs. self-hosted approach.
  • Design edge deployment architectures for vision models on constrained devices (NVIDIA Jetson, edge TPUs, mobile/embedded accelerators), including model compression and quantization.
  • Select and optimize model formats, runtimes, and accelerators (ONNX, TensorRT, OpenVINO, CoreML) for target hardware and latency/throughput requirements.
  • Define evaluation frameworks for vision/multimodal model quality — embedding retrieval accuracy, grounding fidelity, hallucination, and task-specific benchmarks.
  • Collaborate with data science, MLOps, and client-facing teams to translate business requirements into scalable, cost-efficient vision architectures.
  • Document reference architectures, best practices, and reusable accelerators for vision and multimodal GenAI engagements.

Required Skills & Experience

  • Strong hands-on experience with CLIP and other contrastive embedding models, and Vision-Language Models (VLMs) for tasks such as captioning, visual QA, and grounding.
  • Practical experience designing Vision RAG systems, including multimodal embeddings, vector databases, and retrieval-augmented generation patterns.
  • Experience fine-tuning and adapting vision/VLM models using parameter-efficient techniques (LoRA/QLoRA) and classical CV model training/transfer learning.
  • Working knowledge of hyperscaler AI/vision services across at least two of AWS, GCP, and Azure, and their trade-offs for vision workloads.
  • Familiarity with edge AI deployment — hardware accelerators (NVIDIA Jetson, Coral/edge TPU), model compression, quantization, and runtime optimization.
  • Proficiency in Python and vision/ML frameworks (PyTorch, Hugging Face Transformers, OpenCV) and model format/runtime tooling (ONNX, TensorRT, OpenVINO).
  • Understanding of containerization and orchestration (Docker, Kubernetes) for scalable vision model serving.
  • Strong architectural and communication skills, with the ability to translate business needs into technical vision/GenAI solution designs.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

Hugging Face Transformers

TensorRT

NVIDIA Jetson

CLIP

ONNX

Coral edge TPU

LoRA

OpenVINO

Vision RAG systems

QLoRA