Engineering Team Leader (Site Reliability Engineering)
27.2k - 34.6k PLN27 200 - 34 600 PLN/ mies.UoP
SeniorFull-time·Umowa o pracę
#410824·Dodano 4 dni temu·4
Źródło: nofluffjobs.comTech Stack / Keywords
PythonKubernetesAnsiblePrometheusGrafana
Firma i stanowisko
We are building XTB – a global investment company offering innovative technological solutions that allow clients to effectively manage their finances using the intuitive XTB app with over one million users worldwide. We are a certified Great Place to Work company.
Wymagania
- Several years of experience in Site Reliability Engineering, Infrastructure, or DevOps managing high-scale, distributed environments.
- Proven leadership experience in a formal management role leading and developing SRE or DevOps engineering teams.
- Extensive experience building and maintaining scalable, reliable, and observable infrastructure systems on Azure, Kubernetes, and on-premise.
- Demonstrated ability to deliver end-to-end reliability strategies and manage large-scale technical projects from design to production.
- Strong skills in Python for automation and tooling.
- Expertise in managing Kubernetes and configuration management using Ansible.
- Proficiency with observability tools like Prometheus, Grafana, OTEL, ELK, Tempo, and Thanos.
- Experience applying AI/ML for AIOps such as anomaly detection and log analysis.
Nice to have:
- Experience with commercial APM platforms (e.g., Datadog, Splunk, New Relic) and chaos engineering tools.
- Knowledge of cloud cost management and FinOps principles.
Obowiązki
Leadership & Management:
- Break down complex projects into actionable tasks and drive incremental value.
- Mentor and support the professional growth of team members.
- Facilitate workshops, build technical community, and resolve conflicts effectively.
- Drive operational excellence and reliability culture; lead incident management and champion post-mortems.
- Manage technical debt proactively and align team output with organizational goals.
Core Responsibilities:
- Shape and grow a high-performing Site Reliability Engineering team fostering technical excellence and continuous improvement.
- Define and drive the SRE platform strategy in collaboration with infrastructure and development teams.
- Oversee organization-wide 24/7 on-call and incident management processes, managing incident tooling, procedures, compliance, reporting, and continuous improvement.
- Define and track measurable objectives (KPIs) for team performance and operational health.
- Oversee the design, development, and evolution of the organization's observability ecosystem including standardized telemetry, structured logging, distributed tracing, and intelligent sampling.
Benefity
- Sport subscription
- Training budget
- Private health care
- An extra day off on your birthday
- An extra day off for parents
- Access to an e-learning platform for learning English
Karta sportowa
Dofinansowanie szkoleń
Opieka zdrowotna
Płatny urlop
Kursy językowe
XTB
16 aktywnych ofert