← Browse all jobs

Y

Hardware Engineer

Yotta Data Services Private Limited · Navi mumbai, India

FULL TIME

Job Description

Job Description – Hardware Engineer


Experience: 5+ Years


Role Overview:

We are looking for an experienced and hands-on Hardware Engineer to support the operations, maintenance, troubleshooting and lifecycle management of server hardware across our infrastructure. The role will involve working with both AI/GPU-based servers and conventional CPU-based enterprise servers in mission-critical data center environments.

The ideal candidate should have strong experience in server hardware break-fix, fault diagnosis, component replacement, spare-parts management, RMA processes and coordination with OEM/ODM support teams .


Key Responsibilities:

1. Server Hardware Operations & Maintenance

  • Perform end-to-end troubleshooting, maintenance and repair of AI/GPU and conventional enterprise servers.
  • Ensure high availability and reliability of server hardware deployed across the infrastructure.
  • Monitor and identify hardware faults, degradation and recurring failure patterns.
  • Perform hardware diagnosis and identify faulty components for replacement.
  • Ensure servers are tested and operational after repair, replacement or maintenance activities.
  • Work within defined SLAs to ensure timely resolution of hardware-related incidents.


2. Server Hardware Break-Fix

  • Handle the complete hardware break-fix lifecycle, including:
  • Fault detection and diagnosis
  • Faulty component identification
  • Spare allocation
  • Onsite component replacement
  • Server testing and service restoration
  • Faulty part removal and tagging
  • RMA coordination and closure
  • Troubleshoot and replace server Field Replaceable Units (FRUs), including:
  • GPUs
  • CPUs
  • Motherboards
  • Memory
  • NICs
  • Power Supply Units (PSUs)
  • Fans
  • Storage and other server components


3. AI/GPU and Enterprise Server Support

  • Provide hardware support for NVIDIA GPU-based servers and compute nodes.
  • Work on high-density AI and HPC server environments.
  • Support GPU server platforms, including HGX/DGX architecture and other GPU-based server platforms.
  • Perform troubleshooting and replacement of GPU cards and associated server components.
  • Support conventional enterprise and CPU-based servers, including:
  • General-purpose compute servers
  • Application and database servers
  • Virtualization servers
  • High-performance compute servers


4. Spare Parts & Inventory Management

  • Maintain and manage server hardware spares required for break-fix activities.
  • Ensure proper receipt, inspection, storage and issue of server components.
  • Maintain accurate records of spare inventory and component movement.
  • Monitor availability of critical server FRUs and escalate requirements for replenishment.
  • Support inventory reconciliation of replaced, repaired and available spare components.
  • Ensure critical AI/GPU server components are available as per operational requirements.


5. Faulty Part & RMA Management

  • Identify, tag and maintain proper records of faulty or replaced components.
  • Follow the defined process for segregation and storage of faulty hardware.
  • Coordinate with OEMs/ODMs for raising and tracking RMA cases.
  • Prepare faulty components for shipment to designated OEM/ODM service centers.
  • Track repair and replacement status of faulty components.
  • Ensure repaired or replacement components are received and appropriately updated in the inventory.
  • Maintain accurate documentation to avoid unaccounted or misplaced components.


6. OEM/ODM & Vendor Coordination

  • Coordinate with OEMs, ODMs and hardware service partners for technical support and issue resolution.
  • Follow up on delayed parts, replacement requests and unresolved hardware issues.
  • Support warranty and service-related activities.
  • Escalate critical or recurring hardware failures to the relevant internal and external teams.
  • Coordinate with central engineering teams and onsite support teams for timely issue resolution.


7. Incident Management & Documentation

  • Respond to server hardware incidents within defined response and resolution timelines.
  • Maintain detailed records of hardware faults, repairs, replacements and RMA activities.
  • Document troubleshooting steps and resolutions for recurring issues.
  • Follow established operational processes, escalation procedures and SLAs.
  • Participate in shift or standby support for critical 24x7 environments, as required.


Required Skills & Experience:

  • 5+ years of experience in server hardware operations, data center hardware support or enterprise server infrastructure.
  • Strong hands-on experience in server hardware troubleshooting and break-fix operations .
  • Strong understanding of server hardware architecture and components.
  • Experience in diagnosing and replacing server components such as GPUs, CPUs, motherboards, memory, NICs, PSUs, fans and other FRUs.
  • Experience with spare-parts management and hardware inventory .
  • Experience in faulty part handling and RMA management .
  • Experience coordinating with OEMs, ODMs and hardware service partners .
  • Understanding of hardware incident management and SLA-driven support environments.
  • Experience working in 24x7 mission-critical data center or enterprise infrastructure environments .
  • Ability to troubleshoot issues independently and coordinate with multiple technical teams.


Preferred Skills:

Experience with one or more of the following will be an added advantage:

  • NVIDIA GPU-based servers and compute infrastructure
  • NVIDIA H100, H200, B200, B300, GB200 or GB300 platforms
  • NVIDIA HGX/DGX architecture
  • AI/HPC server environments
  • High-density compute infrastructure
  • Supermicro
  • ASUS
  • Gigabyte
  • Dell
  • HPE
  • GPUaaS, AI Cloud or large-scale AI data center environments


Educational Qualification:

Bachelor's degree or diploma in Computer Science, Electronics, Electrical Engineering, Information Technology, or a related technical discipline.

Relevant hardware, server or OEM certifications will be an added advantage.

Details

CompanyYotta Data Services Private Limited
LocationNavi mumbai, India
TypeFULL TIME
Nichetech

Similar Jobs

Y

Electronics Hardware Engineer

Yotuh Energy

O

Principal ADC Engineer

Omni Design Technologies, Inc.

O

Staff ADC Engineer

Omni Design Technologies, Inc.

C

Civil Engineer (AutoCAD / Civil 3D) - US Client Projects

Civion Design Pvt

P

Medical Sales Representative

Pfizer