Senior Site Reliability Engineer

Brak informacji o wynagrodzeniu
SeniorFull-time
#435211·Dodano miesiąc temu·1
Źródło: Roche
Aplikuj teraz

Tech Stack / Keywords

AWSAzureKubernetesEKSAKSPythonTerraformCI/CDObservability

Firma i stanowisko

Roche is a global healthcare company dedicated to advancing science and ensuring access to healthcare worldwide. The team consists of more than 100,000 employees globally, delivering medicines and Diagnostics products that impact over 26 million people and conduct over 30 billion tests. The company fosters innovation and creativity to deliver life-changing healthcare solutions.

Wymagania

  • Bachelor’s degree in computer science, Engineering, or a related field, or equivalent professional experience
  • Experience in site reliability engineering, software engineering or related fields with production on-call experience
  • Solid experience with AWS and/or Azure, including setting up, monitoring, and maintaining cloud resources (including Kubernetes, EKS, AKS, GKE)
  • Proficiency with observability tools
  • Hands-on experience with incident management tools
  • Proficiency in scripting languages for automation purposes
  • Demonstrated proficiency in troubleshooting, especially in cloud and distributed system environments
  • Excellent communication, teamwork and documentation skills
  • Fluent spoken and written English communication
  • Encouraged diverse backgrounds and experiences

Obowiązki

Reliability Engineering & Architecture:

  • Define and implement SLIs, SLOs, and error budgets with product and engineering teams
  • Conduct reliability reviews for new and existing services
  • Design scalable, fault-tolerant architectures in AWS and Azure environments
  • Lead capacity planning, performance and cost optimization initiatives
  • Improve system resilience through automation and self-healing patterns
  • Drive organizational observability maturity (metrics, logs, traces, alert quality)

Incident Management & Continuous Improvement:

  • Perform complex root cause analysis and drive rapid mitigation
  • Participate in blameless postmortems and follow-through
  • Improve MTTR, reduce incident frequency, and elevate production standards
  • Collaborate seamlessly with engineering teams to enable timely and effective resolutions
  • Handle requests and incidents, create and maintain runbooks
  • Participation in a structured 24*7 on-call rotation

Automation & Platform Engineering:

  • Reduce operational toil through tooling and automation (Python or similar)
  • Improve CI/CD reliability and deployment safety mechanisms
  • Build and maintain infrastructure-as-code (Terraform or equivalent)
  • Enhance Kubernetes platform reliability (EKS, AKS, or similar)

Cross-Functional Leadership:

  • Partner with business, engineering, security, and cloud teams to embed reliability early in the software development life cycle
  • Mentor mid-level engineers and help shape SRE best practices
  • Champion a culture of ownership, accountability, and continuous improvement
Roche

Roche

76 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz