Site Reliability Engineer
32k - 44k PLN32 000 - 44 000 PLN/ mies.UoP
SeniorFull-time·Umowa o pracę
#425108·Dodano 12 dni temu·9
Źródło: nofluffjobs.comTech Stack / Keywords
SREDevOpsKubernetesCloudAWSPythonBash
Firma i stanowisko
The company operates in the financial sector and focuses on artificial intelligence to support daily work, automate tasks, and improve efficiency.
Wymagania
- At least 5 years of experience in Site Reliability Engineering, DevOps, platform engineering, or a similar role.
- Strong experience working with complex, distributed, production-grade systems.
- Very good knowledge of observability tools, especially Prometheus, Grafana, Loki, Tempo, and OpenTelemetry.
- Hands-on experience with Kubernetes and Docker.
- Practical experience with both cloud and on-premises infrastructure.
- Preferred experience with AWS.
- Ability to automate tasks and workflows using Python, Bash, Go, or similar scripting languages.
- Good understanding of CI/CD practices, DevOps culture, and agile ways of working.
- Experience improving deployment pipelines, operational tooling, monitoring, and recovery procedures.
- Strong ownership mindset, attention to detail, and proactive approach to reliability improvements.
- Ability to communicate clearly with technical and non-technical stakeholders.
- Responsible approach to AI-assisted engineering, including validation, security awareness, critical thinking, and practical use of AI tools.
Obowiązki
- Help define and promote SRE practices, standards, and operating principles across engineering teams.
- Improve reliability, scalability, and performance of production systems and trading-related platforms.
- Build and enhance monitoring, logging, tracing, and observability solutions.
- Work with tools such as Prometheus, Grafana, Loki, Tempo, and OpenTelemetry.
- Review application reliability requirements within Kubernetes-based environments.
- Support better configuration of services regarding performance, cost, resilience, and operational stability.
- Create automation and internal tools to simplify deployments, health checks, recovery processes, and routine operational tasks.
- Cooperate with development teams to improve fault tolerance, service ownership, and production readiness.
- Support the implementation of SRE practices such as SLOs, incident reviews, and blameless post-mortems.
- Participate in an on-call rotation shared across the team.
Link Group
443 aktywne oferty