AI Model & Dataset Evaluation
- Define and own evaluation frameworks for ML, Computer Vision, LLM, RAG, and Agentic AI systems.
- Establish statistically rigorous model validation methodologies.
- Assess dataset quality, provenance, coverage, and labeling integrity.
- Conduct significance testing before approving model improvements.
AI Security & Red Teaming
- Lead adversarial testing and red-team exercises.
- Evaluate resilience against data poisoning, evasion attacks, prompt injection, jailbreaks, and tool abuse.
- Develop security benchmarks and release-gating criteria for AI deployments.
- Manage vulnerability identification and remediation tracking.
AIOps & Production Reliability
- Design drift monitoring and operational observability frameworks.
- Build shadow deployment and canary rollout strategies.
- Architect high-throughput, low-latency AI inference pipelines.
- Develop confidence calibration and model reliability monitoring.
Governance & Compliance
- Ensure reproducibility, auditability, and certification readiness.
- Maintain AI lifecycle documentation and evaluation records.
- Enforce secure data handling and access-control practices.
Technical Leadership
- Mentor ML engineers, security researchers, and AI practitioners.
- Establish engineering standards for secure AI development.
- Advise leadership on AI risk, deployment readiness, and technology selection.
Required Skills
- 6 - 10 years of experience in Machine Learning, AI Engineering, Security Engineering, or related domains.
- Strong expertise in:
- Statistical testing and model validation
- Classification and detection evaluation
- Adversarial Machine Learning
- LLM, RAG, and Agentic AI evaluation
- MLOps and production AI systems
- Drift detection and monitoring
- Secure deployment practices in high-security environments