Hardware Engineer, GPU Infrastructure
Tech Stack / Keywords
Firma i stanowisko
CoreWeave is The Essential Cloud for AI™, delivering technology, tools, and teams that enable innovators to build and scale AI. Founded in 2017, CoreWeave became publicly traded in 2025 and operates a platform combining infrastructure performance with technical expertise. The Warsaw team is CoreWeave's EMEA hub for Hardware Compute, focusing on deploying and operating the CoreWeave Cloud Platform inside customer and partner data centres, maintaining high-density GPU systems.
Wymagania
- Bachelor’s degree in Computer Science, Engineering, or related technical field (or equivalent experience).
- 4+ years professional experience in hardware, infrastructure, systems, or platform engineering in data centre, cloud, or HPC environments.
- 2+ years hands-on experience supporting and troubleshooting data centre class GPUs (H100 or newer), including InfiniBand and NVLink.
- In-depth knowledge of server hardware, GPUs, and PCIe devices.
- Proficiency in Python or Ansible, with experience using Redfish or IPMI (Redfish preferred) to interact with server BMCs.
- Experience integrating and automating GPU diagnostics, troubleshooting, and observability platforms like Prometheus and Grafana.
- Proven collaboration with hardware vendors on operational playbooks and RMAs.
- Experience in production infrastructure support, incident escalation, and on-call rotations.
- Excellent technical documentation, problem-solving skills, and fluency in English.
Preferred:
- Experience with new data centre region bring-ups or operating customer/partner facility infrastructure.
- Familiarity with rack-scale GPU platforms such as NVIDIA GB200 / GB300, including NVLink, NVSwitch, and InfiniBand fabrics.
- Experience with NVIDIA diagnostic tooling, GPU driver and firmware lifecycle management, and fleet-scale BMC management (Redfish/IPMI).
- Strong Linux system administration and kernel-level debugging skills, including PCIe enumeration.
- Familiarity with Kubernetes-based services and distributed production environments.
Obowiązki
- Own GPU and interconnect health across the fleet combining hands-on engineering with operational support.
- Automate GPU diagnostics and develop monitoring and alerting systems.
- Build playbooks for detecting and acting on GPU and PCIe failures.
- Lead deep root-cause analysis across hardware, firmware, drivers, and PCIe topologies.
- Partner with data centre operations teams, server OEMs, and vendors such as NVIDIA.
- Serve as regional escalation point for GPU issues across EMEA and APAC.
- Support new region bring-ups.
- Participate in on-call rotation.
- Feed operational insights back into automation and platform design.
Benefity
- Family-level Medical Insurance
- Family-level Dental Insurance
- Generous Pension Contribution
- Life Assurance at 4x Salary
- Critical Illness Cover
- Employee Assistance Programme
- Tuition Reimbursement
Benefits may vary by location.
Inne informacje
Employment offers are conditional upon passing a basic criminal record check in compliance with GDPR. This position requires access to export controlled information; applicants must be U.S. persons or eligible to access such information under U.S. Government export regulations. CoreWeave is an equal opportunity employer committed to inclusive hiring. Speculative CVs are not accepted and unsolicited CVs are considered company property with no fees owed. Privacy and personal data protections are governed according to GDPR and UK GDPR.
CoreWeave Europe
27 aktywnych ofert