Senior Engineer, NCX
292.5k - 507k PLN292 500 - 507 000 PLN/ mies.UoP
375k - 650k PLN375 000 - 650 000 PLN/ mies.UoP
SeniorFull-time·Umowa o pracę
#426391·Dodano 2 dni temu·0
Źródło: NVIDIATech Stack / Keywords
CloudNetworkAINetworkingKubernetesNodeDevOpsLinux
Firma i stanowisko
NVIDIA is a leading technology company specializing in Artificial Intelligence, High-Performance Computing, and Visualization. The role is within the DSX team, focusing on NVIDIA Cloud Partner (NCP) infrastructure operations to support large-scale NVIDIA accelerated infrastructure in production.
Wymagania
- BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or related field, or equivalent experience.
- 8+ years experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, or systems engineering supporting large-scale production environments.
- Strong experience operating Linux-based distributed systems and cloud infrastructure in production.
- Deep understanding of Kubernetes, containers, cluster scheduling, and operational lifecycle of large multi-node environments.
- Strong knowledge of production observability including metrics, logging, alerting, dashboards, and health checks.
- Experience automating infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management.
- Strong networking fundamentals and experience troubleshooting distributed systems across compute, network, and storage layers.
- Programming and automation skills using Python, Go, shell scripting, or similar languages.
Nice to have:
- Experience managing GPU or accelerated computing infrastructure for AI training and inference.
- Experience with NVIDIA technologies such as DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, Network Operator.
- Experience collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, or managed AI clouds.
- Knowledge of infrastructure observability tools including Prometheus, Grafana, OpenTelemetry, Alertmanager.
- Knowledge of failure modes related to large distributed AI workloads.
Obowiązki
- Lead NCP Day 2 operational readiness efforts in collaboration with NVIDIA Cloud Partners.
- Build and implement continuous infrastructure validation for GPU, CPU, storage, and network health.
- Establish observability and operational telemetry across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads.
- Develop automated detection and remediation workflows to manage unhealthy infrastructure.
- Refine fleet lifecycle administration including NVIDIA driver and firmware, Kubernetes node maintenance, OS patching, and configuration management.
- Operationalize NVIDIA reference architectures into production practices, automation, and measurable standards.
- Define operational health and readiness through health signals, SLOs, metrics, and validation mechanisms.
- Build reusable operational frameworks, including tooling, runbooks, and playbooks applicable across multiple NCP environments.
Benefity
- Remote work from Germany, Spain, Czechia, or Poland.
- Competitive base salary range for Poland: 292,500 PLN - 507,000 PLN (Level 4), 375,000 PLN - 650,000 PLN (Level 5).
NVIDIA
22 aktywne oferty