Senior Site Reliability Engineer (Data Platform)
Brak informacji o wynagrodzeniu
SeniorFull-time
#428463·Dodano 4 dni temu·0
Źródło: LinkGroupTech Stack / Keywords
CloudDevOpsDatabasesCassandraMongoDBNoSQLPrometheusGrafana
Firma i stanowisko
On behalf of our client, a global leader in cloud and content delivery, this role is part of a critical team managing a massive data pipeline that collects telemetry from hundreds of thousands of servers and delivers data for analytics and reporting.
Wymagania
- Deep, hands-on experience engineering and operating large-scale, distributed data systems.
- Practical experience with modern NoSQL databases; preference for Cassandra or MongoDB expertise.
- At least 5 years of experience as an SRE or Systems/Infrastructure Engineer managing mission-critical distributed systems.
- Proficient programmer experienced with automation and tooling in Python, Go, or Java.
- Strong foundation in Linux/Unix administration and low-level system troubleshooting.
- Relentless drive to find root causes of complex problems and deliver production-grade solutions.
Obowiązki
- Engineer World-Class Data Systems: Focus on reliability and performance of massive data platforms, working extensively with distributed databases like Cassandra and MongoDB.
- Become the Ultimate Troubleshooter: Serve as the highest technical escalation for complex reliability and performance issues across the global stack.
- Build Insightful Observability: Design and build observability using tools like Prometheus and Grafana.
- Solve Problems with Code: Write automation and internal tooling using Python, Go, or Java to streamline operations and diagnostics.
- Drive Long-Term Reliability: Partner with Engineering, Product, and Network teams to improve architecture and prevent incidents.
linkgroup
459 aktywnych ofert