Infrastructure Engineer: GPU & Kubernetes

Brak informacji o wynagrodzeniu
MidFull-time
#396348·Dodano wczoraj·0
Źródło: nofluffjobs.com
Aplikuj teraz

Tech Stack / Keywords

KubernetesGPULinuxBashPythonHelmAnsibleTerraformCI/CDDockerGrafanaPrometeusPyTorch

Firma i stanowisko

EPOL IT operates within the EPOL HOLDING capital group. For over 15 years EPOL HOLDING group has completed over 100 projects that support business and give an advantage over competitors in the field of telecommunications, industry and health care. Our key domains are Telecommunications, Internet of things, Automation of business processes, Healthcare, Portal solutions and Artificial intelligence.

Wymagania

  • 3+ years of experience with Kubernetes in production environments
  • Hands-on experience with GPU infrastructure: NVIDIA driver/CUDA stack, GPU scheduling in Kubernetes (device plugins, MIG/time-slicing), and diagnosing GPU-related performance issues
  • Solid Linux systems fundamentals including networking, storage, and containers
  • Experience with infrastructure-as-code tools such as Terraform, Helm, Ansible, or similar
  • Proficiency with scripting languages like Python or Bash for automation and tooling
  • Familiarity with CI/CD practices and GitOps workflows

Nice to have:

  • Experience with distributed training frameworks (PyTorch DDP, NCCL) or inference serving (Triton, vLLM, KServe)
  • Exposure to telco cloud or network functions virtualization (NFV/CNF) environments
  • Experience with bare-metal GPU provisioning
  • Background in QA/testing or platform reliability
  • Familiarity with monitoring and deployment tools like Prometheus, Grafana, ArgoCD, or Rancher

Obowiązki

  • Design, build, and maintain Kubernetes clusters running GPU-accelerated workloads (training, inference, and validation test suites)
  • Manage GPU resource scheduling and sharing (MIG, time-slicing, device plugins) to optimize utilization
  • Own the infrastructure lifecycle: provisioning, upgrades, monitoring, cost/capacity planning for cloud and on-prem/bare-metal GPU nodes
  • Build and maintain CI/CD pipelines for infrastructure changes and containerized workloads
  • Troubleshoot performance issues across drivers, CUDA versions, network fabric, and pod scheduling
  • Work closely with QA and product teams to support Touchstone AI Factory's validation environments
  • Set up observability (metrics, logging, tracing) for GPU workloads and cluster health
  • Contribute to infrastructure-as-code practices

Benefity

  • Professional team
  • International projects
  • Agile methodology
  • Flexible working hours
  • Possibility to work hybrid from offices in Kraków or Białystok
  • Friendly working environment
  • Career development opportunities and skills growth
  • Private healthcare
  • Employment contracts offered: B2B, UoP (employment contract), Umowa zlecenie (mandate contract)
Elastyczne godziny
Opieka zdrowotna
EPOL IT Sp. z o.o.

EPOL IT Sp. z o.o.

Pracodawca

Aplikuj teraz