Storage Engineer
larsen & toubro-vyoma- Posted 9 hours ago
- Be among the first 10 applicants
Job Description
Job Purpose
Provide high‑throughput, consistent storage tiers (Scratch/HPS + Object) for large‑scale training data ingest, checkpoints, and inference artifacts.
Roles & Responsibilities
- Implementation
Design/expand Lustre/BeeGFS HPS; NVMe‑oF and Object (S3) tiers; align with AI dataflow and GDS.
Establish namespace, OST/MDT layout, stripe/RAID policies; tiering for warm/cold datasets.
- Operations
Capacity/performance planning; rebalance and failover testing; automate snapshots and checkpoint retention.
Proactive detection of hot spots and metadata contention; schema for small‑file handling.
- Performance & Optimization
Tune RDMA paths, page cache, IO schedulers; validate end‑to‑end I/O profiles for LLM training/inference.
- Reliability & Incident
Lead P0/P1 critical incident response for enterprise-scale AI/ML storage infrastructure supporting NVIDIA GPU clusters, ensuring rapid service restoration and minimal impact to business-critical workloads.
Act as the storage SME during major incidents involving BeeGFS, Lustre, GPFS (IBM Spectrum Scale), NFS, NVMe-oF, Parallel File Systems, and Object Storage platforms.
Perform deep-dive troubleshooting and resolution of storage performance degradation, metadata bottlenecks, I/O latency spikes, filesystem corruption, capacity exhaustion, and hardware failures.
Conduct comprehensive Root Cause Analysis (RCA) for storage-related outages affecting GPU training, inferencing, AI pipelines, and high-performance computing (HPC) workloads.
Backup/DR for critical datasets; test restore time and RPO/RTO regularly; corruption and split‑brain handling.
- Security & Compliance
Multi‑tenant encryption at rest and in transit, POSIX/ACLs, S3 policies, retention/legal holds.
Experience & Educational Requirement
BE/B-Tech or equivalent with Computer Science or Electronics & Communication
Certification- SNIA; vendor (NetApp/Dell/VAST) preferred.
RELEVANT EXPERIENCE
- 7–12 years distributed storage; hands‑on with Lustre/BeeGFS/Ceph, NVMe‑oF, and S3 in GPU environments.
Tools / Tech
- Lustre/BeeGFS tooling, Ceph; GDS; Perf tools (fio, fs‑digests); Prometheus/Grafana; Ansible.

