Senior Site Reliability Engineer (AI/ML Platform)
20k - 27k PLN20 000 - 27 000 PLN/ mies.UoP
SeniorFull-time·Umowa o pracę
#427967·Dodano 2 dni temu·1
Źródło: nofluffjobs.comTech Stack / Keywords
PrometheusGrafanaPythonGoAIExcelSRE
Firma i stanowisko
Link Group is hiring a seasoned Site Reliability Engineer to maintain and enhance the backbone of their global AI/ML services, focusing on a massive distributed AI compute platform.
Wymagania
- Proven experience managing complex, large-scale distributed systems.
- Deep practical knowledge of Kubernetes and container orchestration under heavy load.
- Hands-on experience with observability tools: Prometheus, Grafana, and distributed tracing.
- Strong programming skills in Python and Go, including infrastructure-as-code with tools like Terraform.
- Familiarity or direct experience with AI/ML infrastructure challenges such as model serving, inference engines, and GPU workload management.
- Problem-solving mindset with full ownership of issues.
- Excellent collaboration and mentoring skills in adopting SRE principles.
Obowiązki
- Build and implement comprehensive observability, including telemetry, dashboards with Grafana, and alerting with Prometheus.
- Define and track Service Level Objectives (SLOs) and Indicators (SLIs).
- Write automation scripts and tooling in Python or Go to reduce manual work and develop self-healing systems.
- Lead incident response lifecycle, participate in blameless on-call rotation, and conduct post-mortems.
- Develop runbooks for operational tasks.
- Contribute to the CI/CD ecosystem by building integrations, automated safety checks, and rollback mechanisms.
- Act as a reliability advisor for product teams developing AI services, influencing architecture and operational readiness.
Benefity
- Private healthcare
- Sport subscription
- Foreign language classes
- Life insurance
- Cafeteria system
Opieka zdrowotna
Karta sportowa
Kursy językowe
Ubezpieczenie
Firmowa stołówka
linkgroup
477 aktywnych ofert