Skip to content
← Back to job listings

Senior Systems Technician, Workstations & Servers

CORSAIR · Remote, California, United States

Other EngineeringRemoteExternal listingfull-timeabout 1 hour ago

About The Role

Corsair is seeking a Senior Systems Technician to validate and troubleshoot high-end GPU workstation and server platforms for the Corsair Pro product line. The role is responsible for diagnosing and resolving hardware faults at the component and system level, determining root cause and preventing the issue from recurrence.

This role serves as the technical escalation point for manufacturing on build, test, and burn-in failures, directing diagnosis remotely through production line technicians and managing escalation to component and GPU vendors where required. The role also performs hands-on bench work on prototypes, reference units, and returned systems, and is responsible for ensuring findings are reflected in the build, test, and QA documentation.

Key Responsibilities

Troubleshooting & Failure Analysis

  • Root-cause hardware faults on multi-GPU workstations and servers — no-POST, no-boot, intermittent instability, unexpected shutdowns, performance degradation.
  • Diagnose GPU-specific failures: cards dropping off the PCIe bus, link width/speed negotiation issues, ECC/memory errors and row remapping, XID/driver-level faults, NVLink and fabric errors, VBIOS/firmware mismatches.
  • Isolate power delivery problems — PSU sizing and transient behavior, 12V-2x6/12VHPWR connector seating and integrity, riser and cable faults, rack-level power budgeting.
  • Investigate thermal issues: airflow path validation, inlet/delta-T measurement, throttling behavior under sustained load, fan curve and heatsink verification.
  • Diagnose failures on units you cannot physically access — working from supplied logs, photos, and test output, and by directing on-site technicians through a structured isolation sequence.
  • Read and interpret BMC/IPMI SEL logs, and other diagnostic outputs (dmesg, etc)

Manufacturer Troubleshooting & Escalation

  • Serve as the escalation point for manufacturing build, test, and burn-in failures — triage inbound issues, determine severity, and direct remote technicians through structured diagnostic steps to isolate root cause.
  • Use remote out-of-band access (BMC/IPMI, remote KVM, serial console) to diagnose manufacturing floor units directly where available, rather than relying solely on relayed observations.
  • Recommend containment where analysis indicates a systemic issue and provide the technical basis for that recommendation.
  • Work with manufacturing to document and resolve systemic issues via corrective actions and recommended improvements to the SOPs
  • Escalate to component and GPU vendors on the behalf of manufacturing — consolidate supplied evidence into vendor-grade failure reports, manage the case to closure, and push the resolution back down to every affected site.
  • Conduct first-article inspection and build sign-off on new platforms and configurations before volume production is released.
  • Perform periodic on-site or remote process audits: verify the manufacturing site is building to the current documentation revision and using specified procedures and requirements.

Build, Test & Validation

  • Work with manufacturing to maintain established software stack deployment procedures and processes and ensure documentation and compliance.
  • Validate that new platforms are buildable, testable, and repeatable.
  • Work with manufacturing to implement burn-in and QA/Verification processes, ensure documentation and compliance.

Documentation & Quality

  • Document and maintain processes for manufacturing processes involving assembly, burn-in, qualification
  • Maintain escapes/failures log to track for systemic issues

Qualifications

Required

  • 5+ years hands-on building, testing, and troubleshooting server and/or workstation-class systems in a production or integration environment.
  • Demonstrated component-level and system-level troubleshooting ability — able to isolate a fault to a part without swapping the whole system.
  • Experience troubleshooting hardware remotely through third-party or offsite technicians, including defining diagnostic steps and evidence requirements for people you don't directly supervise.
  • Working knowledge of server hardware fundamentals: PCIe topology, CPU/memory population rules, power delivery, thermal design, rack integration and cabling.
  • Experience with BMC/IPMI out-of-band management and firmware/BIOS updates.
  • Ability to communicate technical instruction clearly across sites and shifts

Preferred

  • Direct experience with high-TDP multi-GPU systems (NVIDIA RTX PRO / data center GPUs, AMD Instinct, Intel Arc Pro).
  • Experience with GPU diagnostic and stress tooling (DCGM, nccl-tests, gpu-burn, vendor field diagnostics).
  • Familiarity with high-speed networking (InfiniBand, 100/200/400/800GbE) and cluster-level bring-up, cable management/labeling schemas and documentation.
  • Experience managing manufacturers as a technical escalation owner — directing failure analysis, driving corrective action, and auditing build compliance across multiple sites.
  • Exposure to formal quality processes (QMS documentation, 8D/CAPA, metrics and KPI reporting).

For roles that are based at our headquarters in Milpitas, CA: The starting base pay for this position is as shown below. The actual base pay is dependent upon a variety of job-related factors such as professional background, training, work experience, location, business needs and market demand. Therefore, in some circumstances, the actual salary could fall outside of this expected range. This pay range is subject to change and may be modified in the future.

Annual Salary Range $110,000—$120,000 USD

This is an external listing. JobSpring does not represent or verify the employer. Report this listing