Senior SRE / Production Reliability Engineer
Brak informacji o wynagrodzeniu
SeniorFull-time
#396464·Dodano 4 dni temu·1
Źródło: Symphony SolutionsTech Stack / Keywords
MicroservicesArchitectureScalaReactTypeScriptDevOpsLinuxNetworking
Firma i stanowisko
Symphony Solutions is a Cloud- and AI-driven technology company headquartered in the Netherlands, delivering world-class services and innovative products. The company operates globally with presence in over 20 countries and provides custom software solutions for Airline, Healthcare, iGaming, E-learning, e-Commerce, and Supply Chain sectors.
Wymagania
Must Have: Technical Skills:
- Confident troubleshooting of Linux terminal, logs, processes, networking basics, and resource usage.
- Strong hands-on understanding of Kubernetes / GKE workloads, pods, services, ingress/gateway, probes, RBAC, resources, autoscaling, and troubleshooting.
- Practical experience with GCP cloud infrastructure including GKE, IAM, networking, load balancing, artifact registry, and production diagnostics.
- Knowledge of Docker / Containers including images, registries, runtime debugging, and container lifecycle.
- Understanding of Helm releases, values, deployment state, and rollback.
- Familiarity with GitOps / FluxCD deployment flow, drift, image automation, and Git-based rollback.
- Understanding of CI/CD and ability to investigate failed pipelines and deployment issues.
- Observability tools: Prometheus, Alertmanager, Grafana; knowledge of metrics, logs, traces, dashboards, alerting, and production monitoring.
- Networking fundamentals: DNS, load balancers, ingress, gateways, TLS, routing, firewall/security rules.
- Ability to define and apply SLI, SLO, SLA reliability targets.
- Experience with incident response, rollback, service recovery, and RCA/postmortem.
- Basic operational knowledge of databases and messaging: PostgreSQL, Couchbase, Kafka, Elasticsearch/ELK or similar.
- Security/compliance basics including secrets, IAM/RBAC, audit trail, release/change evidence.
Must Have: Reliability / Operational Skills:
- Assess technical safety of releases.
- Define monitoring needs during and after deployment.
- Verify production health post-release.
- Prepare and validate rollback plans.
- Lead or support post-release watch.
- Improve runbooks and incident response procedures.
- Identify gaps in dashboards, alerts, deployment checks, and release evidence.
- Collaborate with DevOps without duplicating ownership.
Must Have: Soft Skills:
- Ownership and accountability through incident closure.
- Calm operation under pressure during incidents and rollbacks.
- Clear, concise communication including impact, action, ETA, and decisions.
- Structured troubleshooting across app, infrastructure, network, database, config, release diffs, and external dependencies.
- Risk awareness to identify and explain risky releases.
- Collaboration with Dev, QA, Product, Support, Release Manager, and DevOps.
- Discipline in documentation of runbooks, rollback steps, incident timelines, and evidence.
- Blameless mindset focusing on root cause and prevention.
- Proactivity in identifying missing alerts, dashboards, runbooks, and process gaps.
- Prioritization of real production impact over alert noise.
Nice To Have:
- Experience with Terraform / IaC.
- Advanced GCP infrastructure design.
- Experience with ArgoCD or other GitOps tools.
- Advanced database or performance tuning.
- Service mesh experience.
- On-call experience with PagerDuty, Opsgenie, or similar.
- Experience facilitating RCA/postmortem sessions.
- Betting/gaming domain experience.
- Good written English for operational documentation.
- Basic understanding of AI concepts: agents, skills, MCP.
Obowiązki
- Own production reliability practices together with DevOps and engineering teams.
- Monitor production health and improve observability.
- Support releases, hotfixes, rollback, and post-release watch.
- Validate deployment readiness and rollback readiness.
- Participate in incident response and RCA/postmortem.
- Maintain runbooks, dashboards, alerting rules, and operational documentation.
- Ensure release evidence is traceable and audit-ready.
- Coordinate with Dev, QA, Product, Support, Release Manager, and DevOps during production events.
Symphony Solutions
9 aktywnych ofert