Senior Site Reliability Engineer (Linux, Python, Prometeus)
Brak informacji o wynagrodzeniu
SeniorFull-time
#444887·Dodano wczoraj·0
Źródło: LinkGroupTech Stack / Keywords
LinuxPythonPrometheusGolangIaCKubernetesTerraformSaltStackAnsibleGrafana
Wymagania
- Expert-level experience in Linux/Unix Systems Administration, Site Reliability Engineering (SRE), or DevOps in large-scale distributed production environments.
- Proven expertise in Kubernetes and orchestration of large-scale containerized systems.
- Advanced proficiency in at least one programming language, either Python or Golang, and familiarity with configuration and IaC tools like Terraform, SaltStack, or Ansible.
- Hands-on experience establishing SLOs/SLIs and implementing observability stacks using Prometheus, Grafana, and tracing frameworks.
- Practical experience in architecting software systems and cloud or bare-metal infrastructure at scale.
- Collaborative mindset with strong accountability for system health and experience introducing SRE practices to teams.
Obowiązki
- Design, build, and maintain automation tooling and scripts using Python, Golang, and Infrastructure as Code (IaC) to streamline software deployments and minimize manual intervention.
- Define Service Level Objectives (SLOs) and enhance system telemetry using Prometheus, Grafana, and distributed tracing for improved reliability.
- Drive capacity planning, workload scheduling, and auto-scaling strategies for AI compute infrastructure.
- Participate in on-call rotations, providing technical guidance during critical incidents to accelerate recovery and root-cause analysis.
- Coach developers on SRE principles, promote a culture of reliability, and foster continuous learning within engineering teams.
Link Group
484 aktywne oferty