Site Reliability Engineera
140 - 150 PLN/ godz.
SeniorFull-time
#412112·Dodano 12 dni temu·0
Źródło: LinkGroupTech Stack / Keywords
PythonKotlinDatadog
Firma i stanowisko
Join a team responsible for ensuring the reliability, availability, and operational excellence of large-scale distributed platforms in a fast-paced, global environment.
Wymagania
- Strong experience in Site Reliability Engineering (SRE) and incident management/incident command.
- Hands-on software engineering background with Python and/or Kotlin.
- Experience with distributed systems, systems design, and high-availability architectures.
- Knowledge of observability and monitoring tools such as Datadog and Chronosphere.
- Experience with incident management platforms including PagerDuty and Rootly.
- Strong troubleshooting, debugging, and production support skills.
- Understanding of SLA-driven operations, impact assessment, and severity management.
- Experience with APIs, system integrations, and operational automation.
- Excellent communication skills, including executive-level stakeholder communication.
- Ability to make decisions under pressure and work effectively in global follow-the-sun environments.
Obowiązki
- Lead and coordinate production incident response activities.
- Drive incident communication and status updates to internal and external stakeholders.
- Monitor platform health, identify reliability risks, and implement improvements.
- Investigate and resolve complex production issues across distributed systems.
- Develop automation solutions to reduce operational overhead and manual work.
- Collaborate with engineering, platform, and operations teams to improve service reliability.
- Manage observability, monitoring, and incident tooling.
- Support operational handoffs and continuous improvement of incident management processes.
linkgroup
430 aktywnych ofert