Site Reliability Engineera

140 - 150 PLN/ godz.
SeniorFull-time
#412112·Dodano 12 dni temu·0
Źródło: LinkGroup
Aplikuj teraz

Tech Stack / Keywords

PythonKotlinDatadog

Firma i stanowisko

Join a team responsible for ensuring the reliability, availability, and operational excellence of large-scale distributed platforms in a fast-paced, global environment.

Wymagania

  • Strong experience in Site Reliability Engineering (SRE) and incident management/incident command.
  • Hands-on software engineering background with Python and/or Kotlin.
  • Experience with distributed systems, systems design, and high-availability architectures.
  • Knowledge of observability and monitoring tools such as Datadog and Chronosphere.
  • Experience with incident management platforms including PagerDuty and Rootly.
  • Strong troubleshooting, debugging, and production support skills.
  • Understanding of SLA-driven operations, impact assessment, and severity management.
  • Experience with APIs, system integrations, and operational automation.
  • Excellent communication skills, including executive-level stakeholder communication.
  • Ability to make decisions under pressure and work effectively in global follow-the-sun environments.

Obowiązki

  • Lead and coordinate production incident response activities.
  • Drive incident communication and status updates to internal and external stakeholders.
  • Monitor platform health, identify reliability risks, and implement improvements.
  • Investigate and resolve complex production issues across distributed systems.
  • Develop automation solutions to reduce operational overhead and manual work.
  • Collaborate with engineering, platform, and operations teams to improve service reliability.
  • Manage observability, monitoring, and incident tooling.
  • Support operational handoffs and continuous improvement of incident management processes.
linkgroup

linkgroup

430 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz