Senior Site Reliability Engineer (Linux, Python, Prometeus)

Brak informacji o wynagrodzeniu
SeniorFull-time
#444887·Dodano wczoraj·0
Źródło: LinkGroup
Aplikuj teraz

Tech Stack / Keywords

LinuxPythonPrometheusGolangIaCKubernetesTerraformSaltStackAnsibleGrafana

Wymagania

  • Expert-level experience in Linux/Unix Systems Administration, Site Reliability Engineering (SRE), or DevOps in large-scale distributed production environments.
  • Proven expertise in Kubernetes and orchestration of large-scale containerized systems.
  • Advanced proficiency in at least one programming language, either Python or Golang, and familiarity with configuration and IaC tools like Terraform, SaltStack, or Ansible.
  • Hands-on experience establishing SLOs/SLIs and implementing observability stacks using Prometheus, Grafana, and tracing frameworks.
  • Practical experience in architecting software systems and cloud or bare-metal infrastructure at scale.
  • Collaborative mindset with strong accountability for system health and experience introducing SRE practices to teams.

Obowiązki

  • Design, build, and maintain automation tooling and scripts using Python, Golang, and Infrastructure as Code (IaC) to streamline software deployments and minimize manual intervention.
  • Define Service Level Objectives (SLOs) and enhance system telemetry using Prometheus, Grafana, and distributed tracing for improved reliability.
  • Drive capacity planning, workload scheduling, and auto-scaling strategies for AI compute infrastructure.
  • Participate in on-call rotations, providing technical guidance during critical incidents to accelerate recovery and root-cause analysis.
  • Coach developers on SRE principles, promote a culture of reliability, and foster continuous learning within engineering teams.
Link Group

Link Group

484 aktywne oferty

Zobacz wszystkie oferty
Aplikuj teraz