[MLA] Senior Site Reliability Engineer (SRE) – Kubernetes

Brak informacji o wynagrodzeniu
SeniorFull-time
#435484·Dodano 28 dni temu·0
Źródło: Software Mind
Aplikuj teraz

Tech Stack / Keywords

Site Reliability EngineeringKubernetesSplunkPrometheusGrafanaCI/CDHelmArgoCDFluxNode.js

Firma i stanowisko

Software Mind develops solutions for global companies, including tech giants and unicorns, focusing on transformative projects and emerging technologies. The role is within the AI Experience Framework team that builds the platform powering ServiceNow's AI-first user interfaces.

Wymagania

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Engineering roles.
  • 3+ years hands-on production Kubernetes experience, including deployment, scaling, rollout/rollback, resource tuning, and service-to-service troubleshooting.
  • Strong production incident response experience with on-call, runbooks, postmortems, and paging hygiene.
  • Experience with Splunk for log aggregation and production troubleshooting.
  • Experience building alerts and dashboards with Prometheus and Grafana.
  • Proficiency with CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux.
  • Strong Linux and networking fundamentals, including DNS, load balancing, TCP / HTTP, HTTP/2, and Kubernetes networking.
  • Production troubleshooting experience across Node.js and JVM/Java services, including heap snapshots, CPU profiling, event-loop and memory analysis, JVM GC log analysis, thread dumps, JVM tuning, and latency investigation.
  • Experience with service-to-service authentication including mTLS, certificate rotation and format conversion, and JWT-based authentication.
  • Very good spoken and written English.

Nice to have:

  • Experience with Web Components / Lit for UI-related debugging.
  • Server-side rendering or isomorphic runtime experience.
  • Canary rollout and multi-version production operations.
  • Distributed tracing and request-context correlation.
  • Experience with KEDA or event-driven autoscaling.
  • Experience with enterprise platform integration layers.

Obowiązki

  • Support the deployment, operation, and reliability of production services running on Kubernetes.
  • Monitor service health and investigate production incidents across distributed applications.
  • Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements.
  • Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams.
  • Support CI/CD, GitOps-based deployments, observability, and production monitoring.
  • Work within a client-directed backlog and established priorities.

Benefity

  • Flexible employment and remote work.
  • International projects with leading global clients.
  • International business trips.
  • Non-corporate atmosphere.
  • Language classes.
  • Internal and external training.
  • Private healthcare and insurance.
  • Multisport card.
  • Well-being initiatives.
Elastyczne godziny
Kursy językowe
Szkolenia wewnętrzne
Opieka zdrowotna
Ubezpieczenie
Karta sportowa
SOFTWARE MIND

SOFTWARE MIND

25 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz