Search Jobs

Search by job, company or skills

  • Posted 3 hours ago
  • Be among the first 10 applicants

Job Description

Job Description – Hardware Engineer

Experience: 5+ Years

Role Overview:

We are looking for an experienced and hands-on Hardware Engineer to support the operations, maintenance, troubleshooting and lifecycle management of server hardware across our infrastructure. The role will involve working with both AI/GPU-based servers and conventional CPU-based enterprise servers in mission-critical data center environments.

The ideal candidate should have strong experience in server hardware break-fix, fault diagnosis, component replacement, spare-parts management, RMA processes and coordination with OEM/ODM support teams.

Key Responsibilities:

1. Server Hardware Operations & Maintenance

  • Perform end-to-end troubleshooting, maintenance and repair of AI/GPU and conventional enterprise servers.
  • Ensure high availability and reliability of server hardware deployed across the infrastructure.
  • Monitor and identify hardware faults, degradation and recurring failure patterns.
  • Perform hardware diagnosis and identify faulty components for replacement.
  • Ensure servers are tested and operational after repair, replacement or maintenance activities.
  • Work within defined SLAs to ensure timely resolution of hardware-related incidents.

2. Server Hardware Break-Fix

  • Handle the complete hardware break-fix lifecycle, including:
  • Fault detection and diagnosis
  • Faulty component identification
  • Spare allocation
  • Onsite component replacement
  • Server testing and service restoration
  • Faulty part removal and tagging
  • RMA coordination and closure
  • Troubleshoot and replace server Field Replaceable Units (FRUs), including:
  • GPUs
  • CPUs
  • Motherboards
  • Memory
  • NICs
  • Power Supply Units (PSUs)
  • Fans
  • Storage and other server components

3. AI/GPU and Enterprise Server Support

  • Provide hardware support for NVIDIA GPU-based servers and compute nodes.
  • Work on high-density AI and HPC server environments.
  • Support GPU server platforms, including HGX/DGX architecture and other GPU-based server platforms.
  • Perform troubleshooting and replacement of GPU cards and associated server components.
  • Support conventional enterprise and CPU-based servers, including:
  • General-purpose compute servers
  • Application and database servers
  • Virtualization servers
  • High-performance compute servers

4. Spare Parts & Inventory Management

  • Maintain and manage server hardware spares required for break-fix activities.
  • Ensure proper receipt, inspection, storage and issue of server components.
  • Maintain accurate records of spare inventory and component movement.
  • Monitor availability of critical server FRUs and escalate requirements for replenishment.
  • Support inventory reconciliation of replaced, repaired and available spare components.
  • Ensure critical AI/GPU server components are available as per operational requirements.

5. Faulty Part & RMA Management

  • Identify, tag and maintain proper records of faulty or replaced components.
  • Follow the defined process for segregation and storage of faulty hardware.
  • Coordinate with OEMs/ODMs for raising and tracking RMA cases.
  • Prepare faulty components for shipment to designated OEM/ODM service centers.
  • Track repair and replacement status of faulty components.
  • Ensure repaired or replacement components are received and appropriately updated in the inventory.
  • Maintain accurate documentation to avoid unaccounted or misplaced components.

6. OEM/ODM & Vendor Coordination

  • Coordinate with OEMs, ODMs and hardware service partners for technical support and issue resolution.
  • Follow up on delayed parts, replacement requests and unresolved hardware issues.
  • Support warranty and service-related activities.
  • Escalate critical or recurring hardware failures to the relevant internal and external teams.
  • Coordinate with central engineering teams and onsite support teams for timely issue resolution.

7. Incident Management & Documentation

  • Respond to server hardware incidents within defined response and resolution timelines.
  • Maintain detailed records of hardware faults, repairs, replacements and RMA activities.
  • Document troubleshooting steps and resolutions for recurring issues.
  • Follow established operational processes, escalation procedures and SLAs.
  • Participate in shift or standby support for critical 24x7 environments, as required.

Required Skills & Experience:

  • 5+ years of experience in server hardware operations, data center hardware support or enterprise server infrastructure.
  • Strong hands-on experience in server hardware troubleshooting and break-fix operations.
  • Strong understanding of server hardware architecture and components.
  • Experience in diagnosing and replacing server components such as GPUs, CPUs, motherboards, memory, NICs, PSUs, fans and other FRUs.
  • Experience with spare-parts management and hardware inventory.
  • Experience in faulty part handling and RMA management.
  • Experience coordinating with OEMs, ODMs and hardware service partners.
  • Understanding of hardware incident management and SLA-driven support environments.
  • Experience working in 24x7 mission-critical data center or enterprise infrastructure environments.
  • Ability to troubleshoot issues independently and coordinate with multiple technical teams.

Preferred Skills:

Experience with one or more of the following will be an added advantage:

  • NVIDIA GPU-based servers and compute infrastructure
  • NVIDIA H100, H200, B200, B300, GB200 or GB300 platforms
  • NVIDIA HGX/DGX architecture
  • AI/HPC server environments
  • High-density compute infrastructure
  • Supermicro
  • ASUS
  • Gigabyte
  • Dell
  • HPE
  • GPUaaS, AI Cloud or large-scale AI data center environments

Educational Qualification:

Bachelor's degree or diploma in Computer Science, Electronics, Electrical Engineering, Information Technology, or a related technical discipline.

Relevant hardware, server or OEM certifications will be an added advantage.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

spare-parts management

RMA management

break-fix operations

server hardware architecture

server hardware troubleshooting

server hardware operations

hardware incident management

faulty part handling