Senior Site Reliability Engineer (AI Hardware & Infrastructure)
Brak informacji o wynagrodzeniu
SeniorFull-time
#427465·Dodano wczoraj·0
Źródło: LinkGroupTech Stack / Keywords
AIPythonJiraGrafanaPrometheusLLMNetworking
Wymagania
- Deep background in Site Reliability or Production Engineering with experience managing large-scale infrastructure
- Exceptional Python programming skills for building scalable operational tools
- Strong networking expertise with advanced topologies, high-bandwidth routing and switching, BGP, dual-stack IPv4/IPv6
- Expert-level hands-on experience with observability tools like Prometheus, Grafana, OpenTelemetry, and Loki
- Experience in service rollout strategies, alerting thresholds, technical runbooks, and incident "war rooms"
- Proven ability to independently solve complex technical problems with automated production-grade solutions
- Strong collaboration skills working with external data center vendors and field technicians
Obowiązki
- Write sophisticated tooling and automation in Python to manage lifecycle of server fleet from provisioning to decommissioning
- Integrate core operational systems like JIRA, Siebel, and PagerDuty via APIs to automate workflows
- Design and implement observability stack using Prometheus, Grafana, and AI-driven anomaly detection
- Lead incident response including 24/7 on-call rotation, outage management, and post-mortems
- Use AI utilities and LLM-assisted development to enhance automation and system evaluation
linkgroup
481 aktywnych ofert