Site Reliability Engineer

32k - 44k PLN/ mies.UoP
SeniorFull-time·Umowa o pracę
#425108·Dodano 12 dni temu·9
Źródło: nofluffjobs.com
Aplikuj teraz

Tech Stack / Keywords

SREDevOpsKubernetesCloudAWSPythonBash

Firma i stanowisko

The company operates in the financial sector and focuses on artificial intelligence to support daily work, automate tasks, and improve efficiency.

Wymagania

  • At least 5 years of experience in Site Reliability Engineering, DevOps, platform engineering, or a similar role.
  • Strong experience working with complex, distributed, production-grade systems.
  • Very good knowledge of observability tools, especially Prometheus, Grafana, Loki, Tempo, and OpenTelemetry.
  • Hands-on experience with Kubernetes and Docker.
  • Practical experience with both cloud and on-premises infrastructure.
  • Preferred experience with AWS.
  • Ability to automate tasks and workflows using Python, Bash, Go, or similar scripting languages.
  • Good understanding of CI/CD practices, DevOps culture, and agile ways of working.
  • Experience improving deployment pipelines, operational tooling, monitoring, and recovery procedures.
  • Strong ownership mindset, attention to detail, and proactive approach to reliability improvements.
  • Ability to communicate clearly with technical and non-technical stakeholders.
  • Responsible approach to AI-assisted engineering, including validation, security awareness, critical thinking, and practical use of AI tools.

Obowiązki

  • Help define and promote SRE practices, standards, and operating principles across engineering teams.
  • Improve reliability, scalability, and performance of production systems and trading-related platforms.
  • Build and enhance monitoring, logging, tracing, and observability solutions.
  • Work with tools such as Prometheus, Grafana, Loki, Tempo, and OpenTelemetry.
  • Review application reliability requirements within Kubernetes-based environments.
  • Support better configuration of services regarding performance, cost, resilience, and operational stability.
  • Create automation and internal tools to simplify deployments, health checks, recovery processes, and routine operational tasks.
  • Cooperate with development teams to improve fault tolerance, service ownership, and production readiness.
  • Support the implementation of SRE practices such as SLOs, incident reviews, and blameless post-mortems.
  • Participate in an on-call rotation shared across the team.
Link Group

Link Group

443 aktywne oferty

Zobacz wszystkie oferty
Aplikuj teraz