Database Reliability Engineer

Brak informacji o wynagrodzeniu
SeniorFull-time
#398336·Dodano 5 dni temu·0
Źródło: Alex Staff Agency
Aplikuj teraz

Tech Stack / Keywords

PostgreSQLMongoDBRedisLinuxAnsibleTerraformGitLabCI/CD

Wymagania

  • Deep hands-on experience with PostgreSQL in business-critical production environments, typically 5+ years.
  • Strong understanding of PostgreSQL internals and operations including MVCC, WAL, transactions, locks, indexes, query planning, replication, autovacuum, bloat, major upgrades, backups, PITR, and restore testing.
  • Proven experience with highly available databases and concepts like quorum, split-brain risk, failover, rollback, and recovery.
  • Strong Linux and infrastructure fundamentals: systemd, networking, storage, filesystems, CPU/memory/disk bottlenecks, TLS, DNS, firewalls, and root-cause troubleshooting.
  • Automation skills with Ansible and scripting; Terraform/OpenTofu, GitLab CI/CD, and merge-request based delivery are advantages.
  • Ability to support more than one database engine; willing to learn ClickHouse quickly and take responsibility.
  • Practical use of AI engineering assistants such as Claude and Codex for improving speed and quality of SQL, commands, scripts, and operational conclusions.
  • English proficiency: upper-intermediate or higher for effective communication.

Nice to have:

  • ClickHouse operations including replication, Keeper/ZooKeeper, MergeTree engines, distributed DDL, grants, row policies, backups, query troubleshooting, and cluster recovery.
  • MongoDB replica sets and Percona Backup for MongoDB.
  • Redis/Sentinel and broker/cache failure modes.
  • Database observability, SLOs, golden signals, alert tuning, and executable incident runbooks.
  • Experience building internal platforms, self-service portals, or DBaaS workflows for engineering teams.

Obowiązki

  • Own production PostgreSQL reliability: HA design, Patroni, PgBouncer, replication, failover, upgrades, vacuum/bloat control, query tuning, locks, indexes, capacity, backups, PITR, and restore validation.
  • Improve disaster recovery and operational evidence: tested restores, documented recovery paths, measurable RTO/RPO targets, runbooks, and safe maintenance plans.
  • Support the wider database estate: ClickHouse, MongoDB, and Redis by troubleshooting incidents, reviewing access and data-safety changes, and improving monitoring.
  • Automate DBA workflows with Ansible, Terraform/OpenTofu, GitLab CI/CD, scripts, and reproducible runbooks for provisioning, grants, backups, restores, health checks, and ownership metadata.
  • Help build DBaaS-style self-service capabilities for engineering teams to request databases, access, credentials, and operational checks with less manual DBA intervention.
  • Improve observability and incident response through Grafana, metrics, logs, SLOs, alert rules, and Opsgenie routing; maintain clear communication during production issues.

Benefity

  • A focus on professional development.
  • Interesting and challenging projects.
  • Fully remote work with flexible working hours.
  • Paid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leaves.
  • Compensation for private medical insurance.
  • Co-working and gym/sports reimbursement.
  • Budget for education.
  • Opportunity to receive a reward for the most innovative idea that the company can patent.
Elastyczne godziny
Płatny urlop
Opieka zdrowotna
Karta sportowa
Dofinansowanie szkoleń
Spotkania integracyjne
Alex Staff Agency

Alex Staff Agency

2 aktywne oferty

Zobacz wszystkie oferty
Aplikuj teraz