NVIDIA
NVIDIA
New

Senior Engineer, NCX

292.5k - 507k PLN/ mies.UoP
375k - 650k PLN/ mies.UoP
SeniorFull-time·Umowa o pracę
#426391·Dodano 2 dni temu·0
Źródło: NVIDIA
Aplikuj teraz

Tech Stack / Keywords

CloudNetworkAINetworkingKubernetesNodeDevOpsLinux

Firma i stanowisko

NVIDIA is a leading technology company specializing in Artificial Intelligence, High-Performance Computing, and Visualization. The role is within the DSX team, focusing on NVIDIA Cloud Partner (NCP) infrastructure operations to support large-scale NVIDIA accelerated infrastructure in production.

Wymagania

  • BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or related field, or equivalent experience.
  • 8+ years experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, or systems engineering supporting large-scale production environments.
  • Strong experience operating Linux-based distributed systems and cloud infrastructure in production.
  • Deep understanding of Kubernetes, containers, cluster scheduling, and operational lifecycle of large multi-node environments.
  • Strong knowledge of production observability including metrics, logging, alerting, dashboards, and health checks.
  • Experience automating infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management.
  • Strong networking fundamentals and experience troubleshooting distributed systems across compute, network, and storage layers.
  • Programming and automation skills using Python, Go, shell scripting, or similar languages.

Nice to have:

  • Experience managing GPU or accelerated computing infrastructure for AI training and inference.
  • Experience with NVIDIA technologies such as DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, Network Operator.
  • Experience collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, or managed AI clouds.
  • Knowledge of infrastructure observability tools including Prometheus, Grafana, OpenTelemetry, Alertmanager.
  • Knowledge of failure modes related to large distributed AI workloads.

Obowiązki

  • Lead NCP Day 2 operational readiness efforts in collaboration with NVIDIA Cloud Partners.
  • Build and implement continuous infrastructure validation for GPU, CPU, storage, and network health.
  • Establish observability and operational telemetry across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads.
  • Develop automated detection and remediation workflows to manage unhealthy infrastructure.
  • Refine fleet lifecycle administration including NVIDIA driver and firmware, Kubernetes node maintenance, OS patching, and configuration management.
  • Operationalize NVIDIA reference architectures into production practices, automation, and measurable standards.
  • Define operational health and readiness through health signals, SLOs, metrics, and validation mechanisms.
  • Build reusable operational frameworks, including tooling, runbooks, and playbooks applicable across multiple NCP environments.

Benefity

  • Remote work from Germany, Spain, Czechia, or Poland.
  • Competitive base salary range for Poland: 292,500 PLN - 507,000 PLN (Level 4), 375,000 PLN - 650,000 PLN (Level 5).
NVIDIA

NVIDIA

22 aktywne oferty

Zobacz wszystkie oferty
Aplikuj teraz