Service Manager/Site Reliability Engineer
160 - 180 PLN/ godz.B2B
SeniorFull-time·B2B
#446078·Dodano wczoraj·1
Źródło: RITS Professional ServicesTech Stack / Keywords
DatadogChronospherePagerDutyRootlyAPIsPythonKotlinSDLCdistributed systemsmonitoring
Wymagania
- 5+ years of experience in Incident Operations, Site Reliability Engineering (SRE), Technical Operations, or a similar role.
- Experience working in on-call environments with SLA-driven responsibilities.
- Proven ability to operate effectively during high-pressure, real-time incident scenarios.
- Strong understanding of distributed systems and production environments.
- Experience with monitoring and alerting tools such as Datadog, Chronosphere, or similar.
- Experience with incident management platforms such as PagerDuty, Rootly, or comparable tools.
- Familiarity with APIs, system integrations, observability tools, and monitoring dashboards.
- Hands-on programming experience with Python and/or Kotlin.
- Understanding of the Software Development Lifecycle (SDLC) and production reliability principles.
- Excellent written and verbal communication skills.
- Strong operational judgment and ability to make decisions under uncertainty.
- Ability to manage multiple priorities simultaneously in a fast-paced environment.
- Highly organized with strong ownership and attention to detail.
- Strong collaboration skills with cross-functional engineering and business teams.
- Proactive mindset focused on continuous improvement, automation, and scalability.
Obowiązki
- Monitor, triage, and coordinate responses to production incidents and operational alerts.
- Act as a central coordination point between engineering teams and key stakeholders during incidents.
- Assess incident impact, determine severity, and coordinate communications according to SLA commitments.
- Manage incident lifecycles from detection through resolution and post-incident activities.
- Maintain external-facing incident communications and status updates.
- Support incident reporting, root cause analysis (RCA), and operational reviews.
- Contribute to process improvements, automation initiatives, and operational tooling enhancements.
- Collaborate with engineering teams to improve observability, monitoring, and incident response capabilities.
- Participate in reliability-focused development activities and support operational excellence initiatives.
RITS
250 aktywnych ofert