linkgroup
linkgroup
New

Senior Site Reliability Engineer (AI Hardware & Infrastructure)

Brak informacji o wynagrodzeniu
SeniorFull-time
#427465·Dodano wczoraj·0
Źródło: LinkGroup
Aplikuj teraz

Tech Stack / Keywords

AIPythonJiraGrafanaPrometheusLLMNetworking

Wymagania

  • Deep background in Site Reliability or Production Engineering with experience managing large-scale infrastructure
  • Exceptional Python programming skills for building scalable operational tools
  • Strong networking expertise with advanced topologies, high-bandwidth routing and switching, BGP, dual-stack IPv4/IPv6
  • Expert-level hands-on experience with observability tools like Prometheus, Grafana, OpenTelemetry, and Loki
  • Experience in service rollout strategies, alerting thresholds, technical runbooks, and incident "war rooms"
  • Proven ability to independently solve complex technical problems with automated production-grade solutions
  • Strong collaboration skills working with external data center vendors and field technicians

Obowiązki

  • Write sophisticated tooling and automation in Python to manage lifecycle of server fleet from provisioning to decommissioning
  • Integrate core operational systems like JIRA, Siebel, and PagerDuty via APIs to automate workflows
  • Design and implement observability stack using Prometheus, Grafana, and AI-driven anomaly detection
  • Lead incident response including 24/7 on-call rotation, outage management, and post-mortems
  • Use AI utilities and LLM-assisted development to enhance automation and system evaluation
linkgroup

linkgroup

481 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz