DevOps/SRE Engineer

30 - 37 EUR/ godz.B2B
MidFull-time·B2B
#386240·Dodano 24 dni temu·8
Źródło: nofluffjobs.com
Aplikuj teraz

Tech Stack / Keywords

CephStorageKubernetesDevOpsSite reliability engineeringAnsibleTerraformPythonBashHashiCorp VaultLinuxCI/CD PipelinesPrometheusELK StackGrafanaNetworkingAIAWS

Firma i stanowisko

The role involves designing, implementing, operating, and improving enterprise-scale Ceph storage platforms supporting business and research workloads across hybrid and cloud-native environments. The engineer will also work with technologies including Kubernetes, HashiCorp Vault, Weka, and Qumulo.

Wymagania

  • Minimum 2 years' experience with hands-on design, deployment, administration, and support of Ceph storage clusters
  • Experience operationalizing Ceph including lifecycle management, upgrades, expansion, health monitoring, capacity planning, troubleshooting, and performance tuning
  • Experience automating Ceph operations using Ansible, Terraform, Python, Bash, or similar automation frameworks
  • Strong understanding of distributed storage concepts including replication, erasure coding, CRUSH maps, OSDs, MONs, MGRs, CephFS, RBD, and RGW
  • Experience integrating Ceph with Kubernetes or cloud-native platforms
  • Experience with Linux platform administration
  • Knowledge of networking concepts supporting distributed storage platforms
  • Strong understanding of Site Reliability Engineering (SRE) principles
  • Experience implementing Infrastructure-as-Code using Terraform and/or Ansible
  • Experience building CI/CD pipelines for infrastructure automation
  • Experience with monitoring, alerting, logging, and observability platforms such as Prometheus, Grafana, and ELK Stack
  • Experience with incident management and root cause analysis
  • Strong scripting skills in Python, Bash, or similar

Nice to have:

  • Experience with other HPC storage platforms
  • Experience supporting hybrid cloud infrastructure, preferably AWS
  • Experience operating storage platforms supporting AI/ML, HPC, or large-scale Kubernetes environments

Obowiązki

  • Design, deploy, administer, and support Ceph storage clusters in production environments
  • Operationalize Ceph including lifecycle management, upgrades, expansion, health monitoring, capacity planning, troubleshooting, and performance tuning
  • Automate Ceph operations using Ansible, Terraform, Python, Bash, or similar frameworks
  • Build operational tooling, runbooks, and self-service capabilities to improve platform efficiency and reliability
  • Integrate Ceph with Kubernetes or cloud-native platforms
  • Improve platform reliability through automation, observability, proactive monitoring, incident response, capacity planning, and continuous improvement
  • Work with complementary storage technologies such as Weka and Qumulo, HashiCorp Vault, and modern infrastructure platforms
Ework Group

Ework Group

73 aktywne oferty

Zobacz wszystkie oferty
Aplikuj teraz