Lead Site Reliability Engineer - Imunify Reliability Platform (remote-only)
Brak informacji o wynagrodzeniu
SeniorFull-time
#413325·Dodano 2 dni temu·1
Źródło: CloudlinuxTech Stack / Keywords
LinuxSecurityCloudPythonGoRustPrometheusGrafana
Firma i stanowisko
CloudLinux is a company focused on Linux security, stability, and profitability, with over 500,000 product installations and 4,000 customers including Liquid Web, 1&1, and Dell. The role involves working remotely full-time with flexible hours within a company that emphasizes an open management style and remote work culture.
Wymagania
Required:
- Substantial production engineering or SRE experience including defining an SLO framework.
- Strong Python skills; comfortable reading/modifying Go or Rust.
- Deep practical experience on time-series and event telemetry at scale using tools like Prometheus/OpenMetrics, Grafana, Alertmanager, and columnar stores (ClickHouse or equivalent).
- Experience debugging distributed systems on bare metal and long-lived hosts outside Kubernetes.
- Familiarity with configuration management and CI tools such as Ansible, GitLab CI, and Jenkins.
- Ability to design measurement for systems not owned or scrappable, addressing push telemetry, sampling, clock skew, and privacy constraints.
- Strong written asynchronous communication skills.
Nice to have:
- Background in security products such as WAF, EDR, AV, and vulnerability management.
- Experience with monitoring under audit standards like SOC 2, ISO 27001, and NIST.
- Knowledge of OpenTelemetry, eBPF, Sentry.
- Expertise in cost- and cardinality-aware telemetry design.
- Familiarity with agentic development tooling and cursors.
- Experience with Kubernetes.
Obowiązki
SLI Definition:
- Define service level indicators (SLIs) with squad leads and senior engineers for about 70 components.
- Build a taxonomy of SLIs including Service SLIs, Fleet SLIs, Control-efficacy SLIs, Delivery SLIs, and Pipeline SLIs.
- Enforce measurement rules ensuring SLIs are externally measurable and attach SLOs, error budgets, and owners.
Collection System:
- Design and build a push-based, sampled, privacy-aware pipeline to collect telemetry from the fleet into a queryable store.
- Extend instrumentation on agent-side and service-side in Python, Go, and Rust.
- Consolidate dashboards and reporting into a defensible set of instruments.
Alerting and Alert Management:
- Implement symptom-based, SLO-anchored alerting with multi-window burn-rate semantics.
- Create a three-tier alert taxonomy: page, ticket, dashboard, with rules for paging humans.
- Ensure every alert has an owner, runbook, and documented failure mode.
- Maintain alert hygiene with quarterly reviews and track actionable rates.
Escalation:
- Maintain a component-to-owning squad ownership map, wired to routing systems for alert delivery.
- Define severity matrices, acknowledgement SLAs, and follow-the-sun on-call rotations across time zones.
- Support incident command practices with blameless postmortems within 24 hours.
- Coach squads to carry their own pagers and operate the escalation platform.
Benefity
- Professional development opportunities including challenging projects and mentorship programs.
- Fully remote work with flexible working hours from any location worldwide.
- Paid 24 vacation days per year, 10 national holidays, and unlimited sick leaves.
- Compensation for private medical insurance.
- Reimbursement for co-working spaces and gym/sports activities.
- Opportunity to receive rewards for innovative ideas eligible for patenting.
Elastyczne godziny
Płatny urlop
Opieka zdrowotna
Karta sportowa
CloudLinux
7 aktywnych ofert