Engineering Manager, SRE
Tech Stack / Keywords
Firma i stanowisko
Remote is solving modern organizations’ biggest challenge – navigating global employment compliantly with ease. The company enables businesses of all sizes to recruit, pay, and manage international teams through their best-in-class HR platform. The Site Reliability Engineering (SRE) team at Remote owns Kubernetes, AWS, PostgreSQL, CI infrastructure, observability stack, and reliability practices. The SRE team is part of Platform Engineering.
Wymagania
People leadership:
- Experience leading an SRE, infrastructure, or platform team with responsibility for growth, performance, and career progression.
- Coaching skills for both technical craft and soft skills.
- Ability to manage underperformance with clarity and empathy.
- Hiring experience and strong interview skills.
- Conflict resolution and fostering commitment to company goals.
Technical depth:
- Hands-on background in site reliability, DevOps, or cloud infrastructure engineering.
- Experience with Kubernetes in production.
- Experience with AWS at scale.
- Experience building, enabling, and scaling AI infrastructure.
- Knowledge of observability practices.
- Proficiency with infrastructure as code using Terraform.
- Experience with CI/CD systems like GitLab CI, GitHub Actions, or Jenkins.
- Skilled with Docker and shell scripting.
- Experience running a reliability practice including incident response, on-call, SLOs, and error budgets.
- Experience working in regulated environments.
Ways of working:
- Ability to prioritize operational and project work effectively.
- Strong written communication for asynchronous, distributed work.
- Skilled in building cross-team relationships.
Nice to have:
- Knowledge of backend languages such as Elixir, Java, Clojure, Node.js, or Python.
- Experience with modern observability tools like OpenTelemetry, distributed tracing, Honeycomb.
- Database operations experience, especially with PostgreSQL or Aurora.
- Experience with Linux systems outside cloud environments.
- Security expertise both defensive and offensive.
- Cloud cost management and FinOps experience.
- Experience growing and hiring for small teams.
Obowiązki
Your people:
- Manage full career lifecycle of reports: onboarding, feedback, performance assessment, progression, and hiring.
- Maintain team health, dynamics, and retrospective practices.
- Act as spokesperson for the team across engineering and senior leadership.
Delivery:
- Define and prioritize SRE goals and commitments.
- Oversee support rotation and on-call model.
The platform:
- Manage core infrastructure including Kubernetes, AWS, PostgreSQL, DNS, TLS, and CI infrastructure.
- Oversee reliability practices such as SLOs, error budgets, incident response, and observability stack.
- Partner with Security team on threats, patching, infrastructure controls, audits, and compliance.
- Manage vendor relationships, renewals, and commercial conversations with Director support.
Benefity
- Work from anywhere globally.
- Flexible paid time off.
- Flexible working hours with asynchronous work culture.
- 16 weeks paid parental leave.
- Budget for co-working spaces, learning, and wellness including gym memberships.
- Mental health support services.
- Stock options.
- Home office budget and IT equipment.
Inne informacje
Candidates overlapping with EMEA or APAC time zones are especially welcome; those helping cover the Americas gap are also encouraged. The position is fully remote. Applications and interviews are conducted in English. Equal opportunity and accommodations for under-represented groups are emphasized. The company embraces AI as a tool while prioritizing human creativity. Applications accepted on an ongoing basis.
Remote
7 aktywnych ofert