linkgroup
linkgroup
New

Senior Site Reliability Engineer (AI Hardware & Infrastructure)

20k - 27k PLN/ mies.UoP
SeniorFull-time·Umowa o pracę
#427970·Dodano 2 dni temu·1
Źródło: nofluffjobs.com
Aplikuj teraz

Tech Stack / Keywords

PythonNetworkingBGPPrometheusGrafana

Wymagania

  • Deep background in Site Reliability or Production Engineering with a Computer Science foundation.
  • Exceptional programming skills in Python, building robust operational tools and automation frameworks.
  • Expert knowledge in Networking, advanced network topologies, high-bandwidth routing and switching, and BGP.
  • Hands-on experience with observability tools like Prometheus, Grafana, OpenTelemetry, and Loki.
  • Experience designing service rollout strategies, alerting thresholds, technical runbooks, and leading incident responses.
  • Strong problem-solving skills with ability to deliver production-grade automated solutions.
  • Excellent collaboration skills with external vendors and field technicians.

Obowiązki

  • Write sophisticated tooling and automation in Python to manage server lifecycle from provisioning to decommissioning.
  • Integrate operational systems like JIRA, Siebel, and PagerDuty through APIs to automate incident workflows.
  • Design and implement a custom observability stack including Grafana dashboards and Prometheus telemetry pipelines.
  • Participate in 24/7 on-call rotation and lead response during critical high-severity outages.
  • Drive blameless post-mortems to improve architecture based on incidents.
  • Use AI-driven tools and LLM-assisted development to enhance automation and system evaluation.
  • Collaborate with data center vendors and on-site technicians to ensure maximum uptime.

Benefity

  • Private healthcare
  • Sport subscription
  • Foreign languages classes
  • Life Insurance
  • Cafeteria system
Opieka zdrowotna
Karta sportowa
Kursy językowe
Ubezpieczenie
linkgroup

linkgroup

477 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz