Senior Site Reliability Engineer (AI Hardware & Infrastructure)
20k - 27k PLN20 000 - 27 000 PLN/ mies.UoP
SeniorFull-time·Umowa o pracę
#427970·Dodano 2 dni temu·1
Źródło: nofluffjobs.comTech Stack / Keywords
PythonNetworkingBGPPrometheusGrafana
Wymagania
- Deep background in Site Reliability or Production Engineering with a Computer Science foundation.
- Exceptional programming skills in Python, building robust operational tools and automation frameworks.
- Expert knowledge in Networking, advanced network topologies, high-bandwidth routing and switching, and BGP.
- Hands-on experience with observability tools like Prometheus, Grafana, OpenTelemetry, and Loki.
- Experience designing service rollout strategies, alerting thresholds, technical runbooks, and leading incident responses.
- Strong problem-solving skills with ability to deliver production-grade automated solutions.
- Excellent collaboration skills with external vendors and field technicians.
Obowiązki
- Write sophisticated tooling and automation in Python to manage server lifecycle from provisioning to decommissioning.
- Integrate operational systems like JIRA, Siebel, and PagerDuty through APIs to automate incident workflows.
- Design and implement a custom observability stack including Grafana dashboards and Prometheus telemetry pipelines.
- Participate in 24/7 on-call rotation and lead response during critical high-severity outages.
- Drive blameless post-mortems to improve architecture based on incidents.
- Use AI-driven tools and LLM-assisted development to enhance automation and system evaluation.
- Collaborate with data center vendors and on-site technicians to ensure maximum uptime.
Benefity
- Private healthcare
- Sport subscription
- Foreign languages classes
- Life Insurance
- Cafeteria system
Opieka zdrowotna
Karta sportowa
Kursy językowe
Ubezpieczenie
linkgroup
477 aktywnych ofert