Senior HPC Performance Engineer
221.3k - 383.5k PLN221 250 - 383 500 PLN/ mies.UoP
292.5k - 507k PLN292 500 - 507 000 PLN/ mies.UoP
SeniorFull-time·Umowa o pracę
#259650·Dodano 8 miesięcy temu·81
Źródło: NVIDIATech Stack / Keywords
NVSHMEMGoNodeNetworkingArchitectureScriptingPythonCloud
Firma i stanowisko
NVIDIA is a leader in Artificial Intelligence, High Performance Computing, and Visualization technologies. The company develops GPU communication libraries such as NCCL, NVSHMEM, and GPUDirect that are crucial for scaling Deep Learning and HPC applications at large scales, involving tens of thousands of GPUs connected by high-speed interconnects and networking technologies.
Wymagania
- M.S. or PhD in Computer Science or related field with relevant performance engineering and HPC experience.
- 3+ years of experience with parallel programming and at least one communication runtime (MPI, NCCL, UCX, NVSHMEM).
- Experience in performance benchmarking and triage on large scale HPC clusters.
- Good understanding of computer system architecture, hardware-software interactions, and operating system principles.
- Proficient in implementing micro-benchmarks in C/C++ and modifying code base as needed.
- Ability to debug performance issues across the entire hardware/software stack.
- Proficient in scripting languages, preferably Python.
- Familiarity with containers, cloud provisioning, and scheduling tools (Kubernetes, SLURM, Ansible, Docker).
- Adaptability and effective communication skills across teams and time zones.
Nice to have:
- Practical experience with Infiniband/Ethernet networks, including RDMA, topologies, and congestion control.
- Experience debugging network issues in large scale deployments.
- Familiarity with CUDA programming and/or GPUs.
- Experience with Deep Learning frameworks such as PyTorch, TensorFlow.
Obowiązki
- Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters.
- Study interactions of libraries with all hardware (GPU, CPU, Networking) and software stack components.
- Evaluate proof-of-concepts and conduct trade-off analysis when multiple solutions are available.
- Triage and root-cause performance issues reported by customers.
- Collect performance data and build tools and infrastructure for visualization and analysis.
- Collaborate with a dynamic team across multiple time zones.
Benefity
- Highly competitive salaries.
- Extensive benefits package.
- Work environment promoting diversity, inclusion, and flexibility.
Elastyczne godziny
NVIDIA
20 aktywnych ofert