Engineering Team Leader (Site Reliability Engineering)
27.2k - 34.6k PLN27 200 - 34 600 PLN/ mies.UoP
SeniorFull-time·Umowa o pracę
#434737·Dodano miesiąc temu·7
Źródło: XTBTech Stack / Keywords
PythonKubernetesAnsibleAzurePrometheusGrafanaOTELELKTempoThanos
Firma i stanowisko
XTB is a global investment company offering innovative technological solutions through a single, intuitive app used by over one million users worldwide. It is a certified Great Place to Work company.
Wymagania
- Several years of experience in SRE, Infrastructure, or DevOps managing high-scale, distributed environments.
- Proven formal management experience leading and mentoring SRE or DevOps teams.
- Extensive experience with scalable, reliable, and observable infrastructure systems (including Azure, Kubernetes, on-premises).
- Ability to deliver end-to-end reliability strategies, architectural improvements, and manage large technical projects.
- Ability to drive cultural change and collaborate effectively with product engineering and remote teams.
- Break down complex projects into actionable tasks and drive incremental value.
- Mentor and support the professional growth of team members.
- Facilitate workshops, build a technical community, and resolve conflicts.
- Drive operational excellence, lead incident management and post-mortems.
- Proactively manage technical debt and align team outputs with organizational goals.
- Strong Python skills for building scalable automation, internal tools, and scripts.
- Expertise in managing Kubernetes, configuration management with Ansible, and designing resilient infrastructure on Azure and on-premises.
- Deep proficiency in building standardized telemetry systems with tools like Prometheus, Grafana, OTEL, ELK, Tempo, Thanos.
- Experience using AI/ML for AIOps, anomaly detection, log analysis, and optimizing reliability workflows.
Nice to have:
- Experience with commercial APM platforms (e.g., Datadog, Splunk, New Relic) and chaos engineering tooling.
- Proficiency with cloud cost management and FinOps principles.
Obowiązki
- Shape and grow a high-performing Site Reliability Engineering team fostering technical excellence and continuous improvement.
- Define and drive the SRE platform and reliability strategy aligned with business objectives.
- Oversee 24/7 on-call and incident management processes, including tooling, procedures, compliance, and reporting.
- Define and track KPIs to drive team performance improvements and operational health visibility.
- Lead the design and evolution of the observability ecosystem implementing standardized telemetry such as structured logging, distributed tracing, and intelligent sampling.
Benefity
- Real influence on company and product development.
- Experienced team sharing knowledge.
- Clear career development paths and regular feedback.
- Regular team-building meetings.
- Training budget for courses and conferences.
- Extra day off on your birthday and for parents.
- Equipment tailored to employee needs.
- Private medical care and group insurance.
- Access to e-learning platform for English and benefits platform.
- Access to wellbeing platform, workshops, and private therapy sessions.
- Remote work options: office in Warsaw or coworking space in your city.
Dofinansowanie szkoleń
Płatny urlop
Opieka zdrowotna
Ubezpieczenie
Elastyczne godziny
XTB
36 aktywnych ofert