AI Model Governance Specialist (Risk)
AI Model Governance Specialist (Risk)
Icici Securities9-12 Years
- Posted an hour ago
- Be among the first 10 applicants
Job Description
Job Summary
The role will be responsible for independently evaluating AI/LLM systems, identifying potential risk vectors, designing robust evaluation frameworks, and establishing controls to support responsible and secure enterprise AI deployment.
Roles & Responsibilities
- Conduct RAG & GenAI risk assessments across end-to-end RAG pipelines, including knowledge-base ingestion, chunking, embeddings, vector retrieval, context grounding, and generation.
- Identify AI/LLM risks such as data leakage, prompt injection, context contamination, retrieval failures, and hallucinations.
- Design and execute LLM evaluation and benchmarking frameworks to assess hallucination rates, toxicity, robustness, model calibration, fairness, and bias.
- Curate golden/reference datasets and design evaluation test cases with reference answers.
- Implement and validate LLM-as-a-Judge evaluation pipelines.
- Establish AI governance and risk mitigation controls covering data preprocessing, pre-deployment validation gates, and post-deployment observability.
- Independently inspect, evaluate, and validate AI/LLM model outputs.
- Conduct technical audits of AI/LLM systems and assess model performance against defined evaluation criteria.
- Support the development of enterprise AI control frameworks covering data provenance, model sign-off, rate-limiting, and continuous drift monitoring.
- Evaluate automated AI assessment and hallucination detection tooling, including Ragas, TruLens, and DeepEval.
Qualifications
- B.Tech / M.Tech in Computer Science, Quantitative disciplines, or MCA.
- Strong quantitative and/or Computer Science foundation.
- Thorough understanding of AI/ML taxonomy and modern Generative AI workflows.
- Deep understanding of RAG architecture and associated risks, including vector search, embeddings, relevance scoring, and context grounding/grounding boundaries.
- Hands-on familiarity with reference-based testing, LLM-as-a-Judge frameworks, red-teaming fundamentals, and semantic similarity evaluation.
Experience & Skills
9 to 12 years of relevant experience
Required Technical Skills:
- Strong expertise in Generative AI / GenAI and Large Language Models (LLMs).
- Strong understanding of Retrieval-Augmented Generation (RAG) pipelines.
- Experience in LLM evaluation, model validation, AI risk assessment, and AI governance.
- Understanding of AI risk vectors including prompt injection, hallucination, data leakage, and retrieval failures.
- Familiarity with enterprise LLM orchestration and deployment platforms such as Amazon Bedrock, Azure OpenAI Service, and Vertex AI.
Preferred / Good-to-Have Skills:
- Working knowledge of Python for querying APIs, parsing JSON outputs, and independently auditing model evaluation scripts.
- Practical experience with NLP evaluation metrics such as ROUGE, BLEU, BERTScore, and semantic distance metrics.
- Ability to construct enterprise AI control frameworks covering data provenance, model sign-off, rate-limiting, and continuous drift monitoring.
- Familiarity with automated hallucination detection and AI evaluation tools such as Ragas, TruLens, and DeepEval.
More Info
Key Skills
data preprocessing
Generative AI
post-deployment observability
LLM-as-a-Judge frameworks
semantic similarity evaluation
reference-based testing
NLP evaluation metrics
pre-deployment validation
automated hallucination detection
AI risk assessment
AI evaluation tools
AI governance
LLM evaluation
