Search Jobs

Search by job, company or skills

AI Model Governance Specialist (Risk)

AI Model Governance Specialist (Risk)

Icici Securities
  • Posted an hour ago
  • Be among the first 10 applicants

Job Description

Job Summary

The role will be responsible for independently evaluating AI/LLM systems, identifying potential risk vectors, designing robust evaluation frameworks, and establishing controls to support responsible and secure enterprise AI deployment.

Roles & Responsibilities

  • Conduct RAG & GenAI risk assessments across end-to-end RAG pipelines, including knowledge-base ingestion, chunking, embeddings, vector retrieval, context grounding, and generation.
  • Identify AI/LLM risks such as data leakage, prompt injection, context contamination, retrieval failures, and hallucinations.
  • Design and execute LLM evaluation and benchmarking frameworks to assess hallucination rates, toxicity, robustness, model calibration, fairness, and bias.
  • Curate golden/reference datasets and design evaluation test cases with reference answers.
  • Implement and validate LLM-as-a-Judge evaluation pipelines.
  • Establish AI governance and risk mitigation controls covering data preprocessing, pre-deployment validation gates, and post-deployment observability.
  • Independently inspect, evaluate, and validate AI/LLM model outputs.
  • Conduct technical audits of AI/LLM systems and assess model performance against defined evaluation criteria.
  • Support the development of enterprise AI control frameworks covering data provenance, model sign-off, rate-limiting, and continuous drift monitoring.
  • Evaluate automated AI assessment and hallucination detection tooling, including Ragas, TruLens, and DeepEval.

Qualifications

  • B.Tech / M.Tech in Computer Science, Quantitative disciplines, or MCA.
  • Strong quantitative and/or Computer Science foundation.
  • Thorough understanding of AI/ML taxonomy and modern Generative AI workflows.
  • Deep understanding of RAG architecture and associated risks, including vector search, embeddings, relevance scoring, and context grounding/grounding boundaries.
  • Hands-on familiarity with reference-based testing, LLM-as-a-Judge frameworks, red-teaming fundamentals, and semantic similarity evaluation.

Experience & Skills

9 to 12 years of relevant experience

Required Technical Skills:

  • Strong expertise in Generative AI / GenAI and Large Language Models (LLMs).
  • Strong understanding of Retrieval-Augmented Generation (RAG) pipelines.
  • Experience in LLM evaluation, model validation, AI risk assessment, and AI governance.
  • Understanding of AI risk vectors including prompt injection, hallucination, data leakage, and retrieval failures.
  • Familiarity with enterprise LLM orchestration and deployment platforms such as Amazon Bedrock, Azure OpenAI Service, and Vertex AI.

Preferred / Good-to-Have Skills:

  • Working knowledge of Python for querying APIs, parsing JSON outputs, and independently auditing model evaluation scripts.
  • Practical experience with NLP evaluation metrics such as ROUGE, BLEU, BERTScore, and semantic distance metrics.
  • Ability to construct enterprise AI control frameworks covering data provenance, model sign-off, rate-limiting, and continuous drift monitoring.
  • Familiarity with automated hallucination detection and AI evaluation tools such as Ragas, TruLens, and DeepEval.

More Info

Job Type:
Industry:
Function:
Employment Type:

Key Skills

data preprocessing

Generative AI

post-deployment observability

LLM-as-a-Judge frameworks

semantic similarity evaluation

reference-based testing

NLP evaluation metrics

pre-deployment validation

automated hallucination detection

AI risk assessment

AI evaluation tools

AI governance

LLM evaluation

About Company