Gcore
Gcore
New

DevOps Engineer (AI Inference)

Brak informacji o wynagrodzeniu
SeniorFull-time·Umowa o pracę
#448552·Dodano 2 dni temu·3
Źródło: justjoin.it
Aplikuj teraz

Tech Stack / Keywords

AnsibleTerraformPrometheusGrafanaKubernetesCNICSIKubernetes clustersSlurm ClustersGPU

Firma i stanowisko

Gcore is a global provider of infrastructure and software solutions for AI, cloud, network, and security, with over 210 edge locations, 50+ cloud regions, and thousands of GPUs. The company collaborates with technology partners such as Intel, NVIDIA, Dell, and Equinix and employs over 550 professionals worldwide.

Wymagania

  • Strong understanding of Kubernetes architecture, including CNI, CSI, operators, ingress/gateway, and control plane components
  • Hands-on experience operating and troubleshooting production Kubernetes clusters
  • Strong Linux and networking troubleshooting skills, including DNS, routing, firewalling, TLS, MTU, connectivity, and performance issues
  • Ability to develop automation and operational tooling using Python, Go, or Bash
  • Experience with Terraform, Ansible, or similar IaC/configuration management tools
  • Experience with VictoriaMetrics, Grafana, or similar monitoring, alerting, and troubleshooting tools
  • Strong experience with Git-based workflows and CI/CD pipelines

Preferred Qualifications:

  • Familiarity with Cluster API or similar Kubernetes cluster lifecycle management technologies
  • Hands-on operation or administration of Slurm clusters
  • Knowledge of Argo CD, GitOps workflows, Helm, or Helmfile
  • Background working with managed platforms, PaaS, or cloud services
  • Exposure to bare metal, GPU, HPC, or other high-performance computing environments

Nice to have:

  • Familiarity with the NVIDIA GPU stack, RDMA/InfiniBand, or high-performance networking
  • Knowledge of OpenStack or similar cloud infrastructure platforms
  • Hands-on experience developing Kubernetes operators or controllers

Obowiązki

  • Design, develop, and maintain infrastructure for AI inference workloads, including GPU scheduling, model deployment pipelines, and data access patterns in on-premises environments
  • Build and manage monitoring and observability tools for AI inference platforms, including dashboards, alerts, and runbooks
  • Collaborate with ML engineers and platform teams to design system architecture for AI workloads, integrate inference runtimes, and test performance at scale

Benefity

  • Competitive compensation
  • Flexible working hours and hybrid or remote options
  • Work from anywhere in the world for up to 45 days per year
  • Private medical insurance
  • Extra paid vacation and sick leave days
  • Support for life’s important moments and celebrations
  • Language courses
  • Modern offices with snacks, drinks, and entertainment
  • Team sports and social activities
Elastyczne godziny
Opieka zdrowotna
Płatny urlop
Kursy językowe
Napoje w biurze
Darmowe przekąski
Karta sportowa
Spotkania integracyjne

Inne informacje

We provide equal opportunity to all applicants without regard to race, color, religion, sex, sexual orientation, age, gender identity, gender expression, national origin, disability, or any other legally protected characteristics.

Gcore

Gcore

11 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz