Senior Site Reliability Engineer (AI/ML Platform)
Brak informacji o wynagrodzeniu
SeniorFull-time
#427470·Dodano 5 dni temu·0
Źródło: LinkGroupTech Stack / Keywords
AIGrafanaPrometheusPythonGoCI/CDArchitectureKubernetes
Firma i stanowisko
LinkGroup is hiring for a role specializing in large-scale AI/ML infrastructure and services.
Wymagania
- Proven experience as a Site Reliability, Platform, or Infrastructure Engineer managing large-scale distributed systems.
- Deep practical experience with Kubernetes and container orchestration under heavy load.
- Hands-on expertise with observability tools including Prometheus, Grafana, and distributed tracing.
- Strong programming skills in Python, Go, and infrastructure-as-code tooling such as Terraform.
- Knowledge or experience with AI/ML infrastructure challenges, including model serving, inference engines, or GPU workloads.
- Proactive problem-solving attitude with ownership of issues through resolution and prevention.
- Collaborative mindset with experience mentoring engineers on SRE principles.
Obowiązki
- Design and implement comprehensive observability using telemetry, dashboards (Grafana), and alerting (Prometheus).
- Define and track SLOs/SLIs to maintain service reliability.
- Automate manual tasks by writing code in Python and Go to create self-healing systems and tooling.
- Lead incident response, participate in on-call rotations, and conduct post-mortems.
- Build and maintain a CI/CD ecosystem with automated safety checks and rollback capabilities.
- Partner with product teams to advise on reliability and operational best practices.
linkgroup
459 aktywnych ofert