Senior Site Reliability Engineer (Linux & Virtualization Platform)
Brak informacji o wynagrodzeniu
SeniorFull-time
#427464·Dodano 4 dni temu·0
Źródło: LinkGroupTech Stack / Keywords
LinuxCloudPythonShellScriptingAnsibleGoChef
Wymagania
- Expert-level knowledge of Linux internals and system administration.
- Practical experience with virtualization technologies, especially KVM/QEMU.
- Strong diagnostic skills to debug hardware, firmware, BIOS/UEFI, kernel drivers, and user-space problems.
- Proficient software developer with strong Python and shell scripting skills for building automation frameworks.
- Significant experience in Senior SRE or development roles focusing on large-scale, distributed, reliability-critical systems.
- Experience with infrastructure-as-code and configuration management tools such as SaltStack, Ansible, Chef, or Puppet.
- Excellent communication skills and ability to collaborate across engineering teams.
Obowiązki
- Develop an expert-level understanding of the virtualized host stack, including physical hardware, BIOS/UEFI, the Linux kernel, and KVM/QEMU virtualization layer.
- Write automation and operational tooling using Python and shell scripting to solve infrastructure problems and eliminate manual work.
- Design and implement advanced monitoring and observability to identify kernel, driver, or hardware issues proactively.
- Act as the escalation point for complex system-level issues across hardware, firmware, BIOS/UEFI, kernel drivers, and user-space applications.
- Participate in a 24/7 on-call rotation managing service-impacting incidents and guiding rapid restoration.
linkgroup
451 aktywnych ofert