RecruTec
RecruTec
New

Senior Site Reliability Engineer SRE

70k - 90k EUR/ rok.B2B
SeniorFull-time·B2B
#395826·Dodano 2 dni temu·0
Źródło: nofluffjobs.com
Aplikuj teraz

Tech Stack / Keywords

SRESecurityCybersecurityIoTInfrastructure as CodeCloudSite reliability engineeringKubernetesAmazon EKSAKSIaCTerraformCloudFormationCloud platformAzureGoogle Cloud PlatformPrometheusGrafanaCI/CDScripting languageBashPythonGoGitHubGitLab CINoSQL

Firma i stanowisko

Our client is a product-driven technology company that focuses on engineering excellence and long-term product ownership. They operate with a strong ownership culture and seek experienced professionals dedicated to reliability, scalability, and operational excellence, especially in SaaS, cloud platforms, or IoT/embedded products.

Wymagania

  • Minimum 5 years of hands-on experience in Site Reliability Engineering, Platform Engineering, or related infrastructure roles.
  • Strong practical experience designing, operating, and scaling Kubernetes environments, including managed offerings like Amazon EKS, AKS, or GKE.
  • Solid experience with Infrastructure as Code (IaC) principles, preferably using Terraform or CloudFormation.
  • Proven experience working extensively with at least one major cloud platform such as AWS, Azure, or Google Cloud Platform.
  • Practical experience building and maintaining monitoring and logging ecosystems, including tools like Prometheus, Grafana, and Elasticsearch.
  • Ability to write and maintain automation scripts in languages such as Bash, Python, and optionally Go.
  • Experience designing and maintaining CI/CD workflows using GitHub Actions, GitLab CI, or similar tools.
  • Experience deploying and operating NoSQL database systems in production.

Nice to have:

  • Experience with Go programming.
  • Familiarity with GitHub and GitLab CI.
  • Knowledge of NoSQL databases.

Obowiązki

Security & Reliability Engineering:

  • Embed security principles throughout infrastructure and service lifecycle.
  • Contribute to security assessments and threat modeling.
  • Improve operational security standards.

Infrastructure Design & Automation:

  • Architect, implement, and maintain highly automated infrastructure and delivery pipelines using Infrastructure as Code.
  • Drive initiatives to improve system resilience including self-healing mechanisms.

Incident Leadership:

  • Manage production incidents and coordinate response efforts.
  • Facilitate blameless post-incident reviews and implement improvements.

On-Call Participation:

  • Participate in structured on-call rotation for production issue resolution.

Observability & Performance Optimization:

  • Design and enhance monitoring, logging, and tracing frameworks.
  • Refine alerting strategies and response procedures.

Cross-Team Collaboration:

  • Collaborate with engineering, product, and security stakeholders on reliability and operational best practices.

Cloud Efficiency & Cost Awareness:

  • Analyze cloud resource consumption and propose efficiency improvements.

Operational Support:

  • Provide technical support internally and assist external users to ensure service continuity.

Access & Permissions Governance:

  • Oversee service-level access management and secure user access control.

Benefity

  • Attractive financial conditions aligned with your skills.
  • 25 days of paid time off.
  • Paid maternity/paternity leave (3 months).
  • Coverage of business trips, training, and conference costs.
  • Payment during short-term illness (a few days).
Płatny urlop
RecruTec

RecruTec

Pracodawca

Aplikuj teraz