AI/ML Engineer Data, Evaluation & Model Improvement
secninjaz technologies llp- Posted 6 hours ago
- Be among the first 10 applicants
Job Description
AI/ML Engineer — Data, Evaluation & Model Improvement
Focus: LLM evaluation, fine-tuning and continuous product improvement
Company: SecNinjaz Technologies LLP
Experience: Around 2 years of relevant hands-on experience
Employment: Full-time, on-site
Locations: Netaji Subhash Place, New Delhi,
About the opportunity
At SecNinjaz, we are building an AI-powered Vulnerability Assessment and Penetration Testing (VAPT) product that helps organisations identify, investigate and validate security weaknesses in authorised environments.
We are looking for an AI/ML Engineer who can build trustworthy datasets, evaluate model and agent performance, and turn measured weaknesses into practical improvements through better data, retrieval and fine-tuning.
You will own the engineering work that helps us answer three questions: Where does the system fail What should we improve Does the change perform better on unseen cases
Strong applied AI/ML skills are essential. Cybersecurity or VAPT experience is a plus. Our security specialists will help define domain labels, review evidence and validate security findings.
What you will work on
- Data pipelines: Transform approved agent traces, tool outputs, expert feedback and verified outcomes into structured, versioned datasets.
- Data quality: Implement cleaning, deduplication, annotation workflows, provenance tracking and sensitive-data handling.
- Evaluation datasets: Create training, validation and held-out test sets, preventing related cases or near-duplicate examples from leaking across splits.
- Model and agent evaluation: Build repeatable evaluations for task completion, tool-use accuracy, evidence quality, false positives, latency and compute cost.
- Failure analysis: Investigate whether failures originate from the model, prompts, retrieval, tools, data or workflow design, and recommend the appropriate improvement.
- Fine-tuning: Run reproducible supervised fine-tuning experiments, including approaches such as LoRA or QLoRA where suitable.
- Retrieval improvement: Evaluate embeddings, rerankers and retrieval pipelines to improve the relevance of the information supplied to agents.
- Experiment tracking: Record dataset versions, configurations, model checkpoints, metrics and failure cases so another engineer can reproduce the results.
- Product integration: Work with the runtime engineer to integrate accepted models and improvements with compatibility tests, regression checks and rollback support.
- Feedback pipelines: Turn reviewed product failures into candidate dataset, evaluation or model improvements before promoting them into a release.
What we are looking for
- Strong Python skills and practical experience building data processing or machine learning pipelines.
- Working knowledge of PyTorch and the Hugging Face ecosystem, or comparable tools.
- A hands-on LLM fine-tuning or post-training project covering data preparation, a baseline, held-out evaluation and failure analysis.
- Understanding of training versus inference, overfitting, generalisation, loss functions and evaluation metrics.
- Familiarity with dataset quality, annotation, deduplication and data leakage prevention.
- Understanding of LLM prompting, structured outputs, retrieval and tool use.
- Ability to work with Linux, Git, GPU environments and reproducible experiment configurations.
- Ability to explain why a change improved results—or why it should be rejected.
A well-executed research project, substantial personal project or implemented coursework can demonstrate relevant skills. Be ready to explain your contribution and the limitations of your results.
Good to have
- Familiarity with cybersecurity, VAPT workflows, vulnerability reports or security datasets.
- Experience with preference optimisation, DPO, reinforcement learning or distillation.
- Exposure to PEFT, TRL, experiment tracking and dataset versioning tools.
- Experience with model quantisation, GPU memory optimisation or distributed training.
- Experience evaluating RAG systems, embeddings, rerankers or agent trajectories.
- Familiarity with private model deployment and inference serving.
You do not need experience with every technology listed.
What you will gain
- Hands-on access to SecNinjaz's on-premises NVIDIA H200 GPU infrastructure for planned product training, fine-tuning, evaluation and inference workloads.
- An opportunity to deepen your expertise through real experiments and measurable product outcomes.
- Collaboration with cybersecurity specialists who can help turn domain expertise into high-quality datasets and evaluations.
- Ownership of the improvement process, from identifying a failure to evaluating and integrating a better solution.
Your initial contribution
Establish a reviewed dataset and evaluation baseline, then run a bounded model-improvement experiment. Compare the candidate against the baseline on unseen cases and make an evidence-based recommendation on adoption.
How to apply
Apply through LinkedIn with your CV and a relevant project link,
- We are especially interested in how you measured improvement, prevented data leakage and investigated failures.
Or
email your CV to [Confidential Information] with a short technical write-up explaining your dataset, fine-tuning approach, evaluation results and personal contribution.and completed and ongoing AI/agentic AI projects, your specific contributions, current CTC, expected CTC, and notice period or earliest date you can join SecNinjaz.
More Info
Key Skills
dataset quality
machine learning pipelines
Hugging Face
LLM prompting
data leakage prevention
structured outputs
LLM fine-tuning
experiment configurations
retrieval
GPU environments
