Pipeline Engineer (worldwide remote, work anywhere)
Brak informacji o wynagrodzeniu
SeniorFull-time
#414511·Dodano 2 dni temu·2
Źródło: Cloudlinux⚠️Uwaga: ta oferta może już nie być aktualna. Sprawdź na stronie pracodawcy, czy rekrutacja jest nadal otwarta.
Tech Stack / Keywords
LinuxSecurityLLMAIGoAPIGitLabPrometheus
Firma i stanowisko
CloudLinux is a global remote-first company delivering high-volume, low-cost Linux infrastructure and security products. Imunify360 Security Suite is a product of CloudLinux Inc., designed for shared and VPS/Dedicated servers providing multi-layer security. CloudLinux serves more than 4,000 customers with over 500,000 product installations.
Wymagania
- Over 5 years of experience in backend, platform, or infrastructure engineering.
- Proven experience building and operating multi-stage data or automation pipelines (CI/CD, ETL/ELT, job orchestration, ML/data platforms).
- Deep knowledge in at least one of Python, Go, or Rust.
- Expertise in systems design focusing on reliability and correctness.
- Practical experience with workflow orchestration/job scheduling tools (e.g., Airflow, Temporal, Prefect, Dagster, Argo).
- Strong background in reliability engineering concepts such as idempotency, retries, checkpointing, graceful degradation, and backpressure.
- Hands-on observability experience with **Prometheus/**Grafana or equivalent, able to design metrics.
- Extensive CI/CD experience, especially with GitLab CI, including dynamic pipelines and self-hosted runners.
- Experience with object storage (S3/Ceph) and large-scale analytical stores such as ClickHouse.
- Proficient in designing state machines and managing concurrency in long-running processes.
- Excellent debugging skills across system, network, and data layers.
- Strong communication skills and ability to work effectively in distributed teams.
- Proficiency in spoken and written English.
Nice to have:
- Experience with progressive delivery techniques including canary deployments, feature flags, and automated rollback.
- Experience operating AI/LLM systems in production, including cost control and token accounting.
- Experience with fleet-scale telemetry and building quality gates based on noisy production data.
- Familiarity with WordPress, PHP, or WAF/ModSecurity concepts.
- Experience with configuration management tools like Ansible, Puppet, or Salt and Linux service operations.
Obowiązki
- Design, build, and operate end-to-end automated protection pipelines.
- Transform fragile batch jobs into resilient, idempotent, and observable systems.
- Define and enforce latency budgets and service level objectives (SLOs).
- Build observability layers including metrics, dashboards, alerting, and health gates.
- Design and implement guardrails such as automatic rollbacks, blast-radius limits, and safe fallback behaviors.
- Reduce manual operational interventions and increase system maintainability.
- Write and maintain unit and integration tests for complex concurrency and failure scenarios.
- Investigate and resolve issues across ClickHouse, GitLab CI, S3/object storage, Prometheus/Grafana, and third-party APIs.
- Collaborate with security analysts and server teams on architecture and production safety.
Benefity
- Professional development focus.
- Interesting, challenging projects.
- Fully remote work with flexible hours worldwide.
- 24 days paid vacation, 10 national holidays, unlimited sick leave.
- Compensation for private medical insurance.
- Reimbursements for co-working spaces and gym/sports.
- Education budget.
- Opportunity to receive rewards for innovative, patentable ideas.
Elastyczne godziny
Opieka zdrowotna
Karta sportowa
Dofinansowanie szkoleń
Płatny urlop
CloudLinux
8 aktywnych ofert