DevOps/SRE Engineer
30 - 37 EUR/ godz.B2B
MidFull-time·B2B
#386240·Dodano 24 dni temu·8
Źródło: nofluffjobs.comTech Stack / Keywords
CephStorageKubernetesDevOpsSite reliability engineeringAnsibleTerraformPythonBashHashiCorp VaultLinuxCI/CD PipelinesPrometheusELK StackGrafanaNetworkingAIAWS
Firma i stanowisko
The role involves designing, implementing, operating, and improving enterprise-scale Ceph storage platforms supporting business and research workloads across hybrid and cloud-native environments. The engineer will also work with technologies including Kubernetes, HashiCorp Vault, Weka, and Qumulo.
Wymagania
- Minimum 2 years' experience with hands-on design, deployment, administration, and support of Ceph storage clusters
- Experience operationalizing Ceph including lifecycle management, upgrades, expansion, health monitoring, capacity planning, troubleshooting, and performance tuning
- Experience automating Ceph operations using Ansible, Terraform, Python, Bash, or similar automation frameworks
- Strong understanding of distributed storage concepts including replication, erasure coding, CRUSH maps, OSDs, MONs, MGRs, CephFS, RBD, and RGW
- Experience integrating Ceph with Kubernetes or cloud-native platforms
- Experience with Linux platform administration
- Knowledge of networking concepts supporting distributed storage platforms
- Strong understanding of Site Reliability Engineering (SRE) principles
- Experience implementing Infrastructure-as-Code using Terraform and/or Ansible
- Experience building CI/CD pipelines for infrastructure automation
- Experience with monitoring, alerting, logging, and observability platforms such as Prometheus, Grafana, and ELK Stack
- Experience with incident management and root cause analysis
- Strong scripting skills in Python, Bash, or similar
Nice to have:
- Experience with other HPC storage platforms
- Experience supporting hybrid cloud infrastructure, preferably AWS
- Experience operating storage platforms supporting AI/ML, HPC, or large-scale Kubernetes environments
Obowiązki
- Design, deploy, administer, and support Ceph storage clusters in production environments
- Operationalize Ceph including lifecycle management, upgrades, expansion, health monitoring, capacity planning, troubleshooting, and performance tuning
- Automate Ceph operations using Ansible, Terraform, Python, Bash, or similar frameworks
- Build operational tooling, runbooks, and self-service capabilities to improve platform efficiency and reliability
- Integrate Ceph with Kubernetes or cloud-native platforms
- Improve platform reliability through automation, observability, proactive monitoring, incident response, capacity planning, and continuous improvement
- Work with complementary storage technologies such as Weka and Qumulo, HashiCorp Vault, and modern infrastructure platforms
Ework Group
73 aktywne oferty