Senior Site Reliability Engineer SRE
70k - 90k EUR70 000 - 90 000 EUR/ rok.B2B
SeniorFull-time·B2B
#395826·Dodano 2 dni temu·0
Źródło: nofluffjobs.comTech Stack / Keywords
SRESecurityCybersecurityIoTInfrastructure as CodeCloudSite reliability engineeringKubernetesAmazon EKSAKSIaCTerraformCloudFormationCloud platformAzureGoogle Cloud PlatformPrometheusGrafanaCI/CDScripting languageBashPythonGoGitHubGitLab CINoSQL
Firma i stanowisko
Our client is a product-driven technology company that focuses on engineering excellence and long-term product ownership. They operate with a strong ownership culture and seek experienced professionals dedicated to reliability, scalability, and operational excellence, especially in SaaS, cloud platforms, or IoT/embedded products.
Wymagania
- Minimum 5 years of hands-on experience in Site Reliability Engineering, Platform Engineering, or related infrastructure roles.
- Strong practical experience designing, operating, and scaling Kubernetes environments, including managed offerings like Amazon EKS, AKS, or GKE.
- Solid experience with Infrastructure as Code (IaC) principles, preferably using Terraform or CloudFormation.
- Proven experience working extensively with at least one major cloud platform such as AWS, Azure, or Google Cloud Platform.
- Practical experience building and maintaining monitoring and logging ecosystems, including tools like Prometheus, Grafana, and Elasticsearch.
- Ability to write and maintain automation scripts in languages such as Bash, Python, and optionally Go.
- Experience designing and maintaining CI/CD workflows using GitHub Actions, GitLab CI, or similar tools.
- Experience deploying and operating NoSQL database systems in production.
Nice to have:
- Experience with Go programming.
- Familiarity with GitHub and GitLab CI.
- Knowledge of NoSQL databases.
Obowiązki
Security & Reliability Engineering:
- Embed security principles throughout infrastructure and service lifecycle.
- Contribute to security assessments and threat modeling.
- Improve operational security standards.
Infrastructure Design & Automation:
- Architect, implement, and maintain highly automated infrastructure and delivery pipelines using Infrastructure as Code.
- Drive initiatives to improve system resilience including self-healing mechanisms.
Incident Leadership:
- Manage production incidents and coordinate response efforts.
- Facilitate blameless post-incident reviews and implement improvements.
On-Call Participation:
- Participate in structured on-call rotation for production issue resolution.
Observability & Performance Optimization:
- Design and enhance monitoring, logging, and tracing frameworks.
- Refine alerting strategies and response procedures.
Cross-Team Collaboration:
- Collaborate with engineering, product, and security stakeholders on reliability and operational best practices.
Cloud Efficiency & Cost Awareness:
- Analyze cloud resource consumption and propose efficiency improvements.
Operational Support:
- Provide technical support internally and assist external users to ensure service continuity.
Access & Permissions Governance:
- Oversee service-level access management and secure user access control.
Benefity
- Attractive financial conditions aligned with your skills.
- 25 days of paid time off.
- Paid maternity/paternity leave (3 months).
- Coverage of business trips, training, and conference costs.
- Payment during short-term illness (a few days).
Płatny urlop
RecruTec
Pracodawca