linkgroup
linkgroup
New

Senior Site Reliability Engineer (AI/ML Platform)

Brak informacji o wynagrodzeniu
SeniorFull-time
#427470·Dodano 5 dni temu·0
Źródło: LinkGroup
Aplikuj teraz

Tech Stack / Keywords

AIGrafanaPrometheusPythonGoCI/CDArchitectureKubernetes

Firma i stanowisko

LinkGroup is hiring for a role specializing in large-scale AI/ML infrastructure and services.

Wymagania

  • Proven experience as a Site Reliability, Platform, or Infrastructure Engineer managing large-scale distributed systems.
  • Deep practical experience with Kubernetes and container orchestration under heavy load.
  • Hands-on expertise with observability tools including Prometheus, Grafana, and distributed tracing.
  • Strong programming skills in Python, Go, and infrastructure-as-code tooling such as Terraform.
  • Knowledge or experience with AI/ML infrastructure challenges, including model serving, inference engines, or GPU workloads.
  • Proactive problem-solving attitude with ownership of issues through resolution and prevention.
  • Collaborative mindset with experience mentoring engineers on SRE principles.

Obowiązki

  • Design and implement comprehensive observability using telemetry, dashboards (Grafana), and alerting (Prometheus).
  • Define and track SLOs/SLIs to maintain service reliability.
  • Automate manual tasks by writing code in Python and Go to create self-healing systems and tooling.
  • Lead incident response, participate in on-call rotations, and conduct post-mortems.
  • Build and maintain a CI/CD ecosystem with automated safety checks and rollback capabilities.
  • Partner with product teams to advise on reliability and operational best practices.
linkgroup

linkgroup

459 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz