Senior Site Reliability Engineer (remote within EMEA)
Tech Stack / Keywords
Firma i stanowisko
FYUL is a fast-growing global company formed in 2024 through the merger of Printful, Printify, and Snow Commerce, powering on-demand commerce at global scale from solo creators to entertainment giants with advanced technology, premium production, and global reach. The Platform Infrastructure team builds, operates, and evolves FYUL's container platform and cloud foundation, owning AWS cloud accounts, Kubernetes, cloud networking, observability stack, core databases, CI/CD pipelines, and infrastructure-as-code while fostering a DevOps culture with self-service tooling.
Wymagania
- Solid Linux systems administration background and scripting experience in Python.
- Strong AWS knowledge: EKS, IAM (roles, policies, IRSA), VPC networking, RDS, S3, SQS; familiarity with Well-Architected Framework and multi-account management.
- Hands-on experience operating and troubleshooting production-scale Kubernetes (EKS), including Helm chart development, CNI networking (Cilium), pod networking/IPAM concepts, container security (ECR, image scanning).
- Proficiency with Terraform (modules, state management), ideally Terragrunt for multi-environment management, and GitOps experience with ArgoCD.
- Experience with Postgres, MySQL, and/or MongoDB at production scale including Aurora.
- CI/CD experience with Jenkins (Jenkinsfile, shared libraries) and/or GitHub Actions, familiarity with blue-green and canary deployment strategies.
- Experience with the Grafana observability stack (Grafana, Prometheus, Loki, Tempo, Mimir) including metrics design, dashboarding, alerting, log aggregation, and distributed tracing, plus maintenance.
- Practical incident management experience: on-call rota, structured response, runbooks, and postmortems.
- Working knowledge of 12-Factor App principles and cost optimization/FinOps awareness.
- Methodical, data-driven troubleshooting approach.
- Strong written communication skills for runbooks, ADRs, and postmortems.
- Comfortable with ambiguous ownership and accountability.
- Experience mentoring less senior engineers and providing constructive feedback.
- Several years of hands-on production infrastructure/SRE experience with leadership in design, standards, and escalation.
Nice to have:
- Experience with GCP.
- Experience with Kafka / AWS MSK.
- Experience in regulated or compliance-sensitive environments (security best practices, access reviews).
- Experience contributing to platform/DevEx roadmaps consumed as self-service products.
Obowiązki
- Architect and manage highly available, secure, and scalable infrastructure across multiple AWS accounts and environments using infrastructure as code.
- Design and operate Amazon EKS clusters, including networking policies, persistent storage, and scaling strategies for containerized workloads.
- Own and evolve core platform services: cloud networking, Kubernetes, databases, and messaging systems.
- Drive large-scale automation projects and set standards for using Terraform, Terragrunt, and GitOps (ArgoCD) across teams.
- Lead adoption of automation to reduce manual operational work and keep environments consistent and repeatable.
- Solve complex, cross-service infrastructure problems related to observability and incident response.
- Improve reliability and observability using Grafana, Prometheus, Loki, Tempo, and Mimir.
- Participate in on-call rotation, lead incident response for production issues, and write runbooks, ADRs, and postmortems.
- Lead security efforts including IAM, encryption, secure logging, and mentor others on secure infrastructure practices.
- Audit infrastructure spending and drive cost optimization with FinOps practices.
- Mentor mid-level SREs, provide feedback, and support onboarding new team members.
- Communicate complex technical concepts to both technical and non-technical stakeholders.
- Partner with product engineering squads to understand needs and represent Platform Infrastructure in cross-team initiatives.
Benefity
- Work remotely or in a modern office in Riga.
- Flexible working hours with late start (up to 11 AM).
- Private health insurance.
- 2 extra paid days off for mental or physical well-being.
- 1 extra paid day off for birthday or other celebration.
- Internal and external learning opportunities.
- Access to mentorship, internal meetups, and hackathons (on-site and online).
- Free healthy lunch when working from Rīga office.
- Employee discount on company merch.
- Exciting team-building events and parties.
Inne informacje
We are an equal-opportunity workplace committed to diversity and inclusion, making hiring decisions based solely on qualifications, merit, and work experience. Candidates must send resumes in English.
Fyul
9 aktywnych ofert