Senior SRE / Production Reliability Engineer

Brak informacji o wynagrodzeniu
SeniorFull-time
#396464·Dodano 4 dni temu·1
Źródło: Symphony Solutions
Aplikuj teraz

Tech Stack / Keywords

MicroservicesArchitectureScalaReactTypeScriptDevOpsLinuxNetworking

Firma i stanowisko

Symphony Solutions is a Cloud- and AI-driven technology company headquartered in the Netherlands, delivering world-class services and innovative products. The company operates globally with presence in over 20 countries and provides custom software solutions for Airline, Healthcare, iGaming, E-learning, e-Commerce, and Supply Chain sectors.

Wymagania

Must Have: Technical Skills:

  • Confident troubleshooting of Linux terminal, logs, processes, networking basics, and resource usage.
  • Strong hands-on understanding of Kubernetes / GKE workloads, pods, services, ingress/gateway, probes, RBAC, resources, autoscaling, and troubleshooting.
  • Practical experience with GCP cloud infrastructure including GKE, IAM, networking, load balancing, artifact registry, and production diagnostics.
  • Knowledge of Docker / Containers including images, registries, runtime debugging, and container lifecycle.
  • Understanding of Helm releases, values, deployment state, and rollback.
  • Familiarity with GitOps / FluxCD deployment flow, drift, image automation, and Git-based rollback.
  • Understanding of CI/CD and ability to investigate failed pipelines and deployment issues.
  • Observability tools: Prometheus, Alertmanager, Grafana; knowledge of metrics, logs, traces, dashboards, alerting, and production monitoring.
  • Networking fundamentals: DNS, load balancers, ingress, gateways, TLS, routing, firewall/security rules.
  • Ability to define and apply SLI, SLO, SLA reliability targets.
  • Experience with incident response, rollback, service recovery, and RCA/postmortem.
  • Basic operational knowledge of databases and messaging: PostgreSQL, Couchbase, Kafka, Elasticsearch/ELK or similar.
  • Security/compliance basics including secrets, IAM/RBAC, audit trail, release/change evidence.

Must Have: Reliability / Operational Skills:

  • Assess technical safety of releases.
  • Define monitoring needs during and after deployment.
  • Verify production health post-release.
  • Prepare and validate rollback plans.
  • Lead or support post-release watch.
  • Improve runbooks and incident response procedures.
  • Identify gaps in dashboards, alerts, deployment checks, and release evidence.
  • Collaborate with DevOps without duplicating ownership.

Must Have: Soft Skills:

  • Ownership and accountability through incident closure.
  • Calm operation under pressure during incidents and rollbacks.
  • Clear, concise communication including impact, action, ETA, and decisions.
  • Structured troubleshooting across app, infrastructure, network, database, config, release diffs, and external dependencies.
  • Risk awareness to identify and explain risky releases.
  • Collaboration with Dev, QA, Product, Support, Release Manager, and DevOps.
  • Discipline in documentation of runbooks, rollback steps, incident timelines, and evidence.
  • Blameless mindset focusing on root cause and prevention.
  • Proactivity in identifying missing alerts, dashboards, runbooks, and process gaps.
  • Prioritization of real production impact over alert noise.

Nice To Have:

  • Experience with Terraform / IaC.
  • Advanced GCP infrastructure design.
  • Experience with ArgoCD or other GitOps tools.
  • Advanced database or performance tuning.
  • Service mesh experience.
  • On-call experience with PagerDuty, Opsgenie, or similar.
  • Experience facilitating RCA/postmortem sessions.
  • Betting/gaming domain experience.
  • Good written English for operational documentation.
  • Basic understanding of AI concepts: agents, skills, MCP.

Obowiązki

  • Own production reliability practices together with DevOps and engineering teams.
  • Monitor production health and improve observability.
  • Support releases, hotfixes, rollback, and post-release watch.
  • Validate deployment readiness and rollback readiness.
  • Participate in incident response and RCA/postmortem.
  • Maintain runbooks, dashboards, alerting rules, and operational documentation.
  • Ensure release evidence is traceable and audit-ready.
  • Coordinate with Dev, QA, Product, Support, Release Manager, and DevOps during production events.
Symphony Solutions

Symphony Solutions

9 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz