Service Manager/Site Reliability Engineer
Tech Stack / Keywords
Firma i stanowisko
Join a global technology organization focused on ensuring the reliability, stability, and operational excellence of large-scale production systems.
Wymagania
- 5+ years of experience in Incident Operations, Site Reliability Engineering (SRE), Technical Operations, or a similar role.
- Experience working in on-call environments with SLA-driven responsibilities.
- Proven ability to operate effectively during high-pressure, real-time incident scenarios.
- Strong understanding of distributed systems and production environments.
- Experience with monitoring and alerting tools such as Datadog, Chronosphere, or similar.
- Experience with incident management platforms such as PagerDuty, Rootly, or comparable tools.
- Familiarity with APIs, system integrations, observability tools, and monitoring dashboards.
- Hands-on programming experience with Python and/or Kotlin.
- Understanding of the Software Development Lifecycle (SDLC) and production reliability principles.
- Excellent written and verbal communication skills.
- Strong operational judgment and ability to make decisions under uncertainty.
- Ability to manage multiple priorities simultaneously in a fast-paced environment.
- Highly organized with strong ownership and attention to detail.
- Strong collaboration skills with cross-functional engineering and business teams.
- Proactive mindset focused on continuous improvement, automation, and scalability.
Obowiązki
- Monitor, triage, and coordinate responses to production incidents and operational alerts.
- Act as a central coordination point between engineering teams and key stakeholders during incidents.
- Assess incident impact, determine severity, and coordinate communications according to SLA commitments.
- Manage incident lifecycles from detection through resolution and post-incident activities.
- Maintain external-facing incident communications and status updates.
- Support incident reporting, root cause analysis (RCA), and operational reviews.
- Contribute to process improvements, automation initiatives, and operational tooling enhancements.
- Collaborate with engineering teams to improve observability, monitoring, and incident response capabilities.
- Participate in reliability-focused development activities and support operational excellence initiatives.
Inne informacje
Informujemy, że administratorem danych jest Relyon IT Services z siedzibą w Warszawie, ul. ul. Wołoska 9 (dalej jako "administrator"). Masz prawo do żądania dostępu do swoich danych osobowych, ich sprostowania, usunięcia lub ograniczenia przetwarzania, prawo do wniesienia sprzeciwu wobec przetwarzania, a także prawo do przenoszenia danych oraz wniesienia skargi do organu nadzorczego. Dane osobowe przetwarzane będą w celu realizacji procesu rekrutacji. Podanie danych w zakresie wynikającym z ustawy z dnia 26 czerwca 1974 r. Kodeks pracy jest obowiązkowe. W pozostałym zakresie podanie danych jest dobrowolne. Odmowa podania danych obowiązkowych może skutkować brakiem możliwości przeprowadzenia procesu rekrutacji. Administrator przetwarza dane obowiązkowe na podstawie ciążącego na nim obowiązku prawnego, zaś w zakresie danych dodatkowych podstawą przetwarzania jest zgoda. Dane osobowe będą przetwarzane do czasu zakończenia postępowania rekrutacyjnego i przez okres możliwości dochodzenia ewentualnych roszczeń, a w przypadku wyrażenia zgody na udział w przyszłych postępowaniach rekrutacyjnych - do czasu wycofania tej zgody. Zgoda na przetwarzanie danych osobowych może zostać wycofana w dowolnym momencie. Odbiorcą danych jest serwis Rocket Jobs oraz inne podmioty, którym powierzyliśmy przetwarzanie danych w związku z rekrutacją.
RITS
256 aktywnych ofert