Hiberus
Hiberus
New

Spark / Data Lakehouse Developer

20.2k - 26.9k PLN/ mies.B2B
SeniorFull-time·B2B
#426273·Dodano 7 dni temu·0
Źródło: justjoin.it
Aplikuj teraz

Tech Stack / Keywords

MongoDBJenkinsPythonApache SparkPySparkPrometheusGrafanaApache Iceberg

Firma i stanowisko

Jesteśmy częścią Hiberus, firmy technologicznej wywodzącej się z Hiszpanii z silną pozycją na rynku polskim: 4700+ specjalistów, obecnej w 14+ krajach i realizującej projekty dla klientów na całym świecie. Pracujemy z technologiami AI, Data, Cloud, Software Development, BI i Cybersecurity. Dla Klienta z sektora finansowego poszukujemy doświadczonego Spark / Data Lakehouse Developera, który dołączy do zespołu Data Platform.

Wymagania

  • 4+ lata doświadczenia w pracy z Apache Spark w środowisku produkcyjnym.
  • Bardzo dobra znajomość PySpark, Spark SQL, DataFrames i Structured Streaming.
  • Doświadczenie z MongoDB i Spark Connector oraz znajomość modelowania i indeksowania danych.
  • Praktyczna znajomość Apache Iceberg i architektury Data Lakehouse (upsert/merge, optymalizacja zapisów).
  • Doświadczenie w CI/CD z Jenkins, w tym automatyzacji testów i deploymentów aplikacji Spark.
  • Bardzo dobra znajomość Python oraz Spark SQL.
  • Znajomość Prometheus i Grafana w obszarze monitoringu i observability.
  • Doświadczenie w pracy w Scrum/Kanban oraz znajomość Jira i Confluence.

Obowiązki

  • Projektowanie i implementacja procesów Spark SQL/PySpark do walidacji i porównywania danych (MongoDB/ODS).
  • Implementacja reguł jakości danych, detekcji duplikatów, anomalii i braków oraz raportowania/alertowania.
  • Integracja i optymalizacja Spark–MongoDB (MongoDB Spark Connector, partitioning, push-down, schema-aware reads).
  • Projektowanie warstwy danych Raw → Bronze → Silver → Gold z wykorzystaniem Delta Lake/Iceberg/Hudi oraz CI/CD dla schematów i migracji.
  • Monitoring i observability procesów Spark/batch/stream (Prometheus, Grafana, OpenTelemetry, Dynatrace) oraz definiowanie SLA/SLO.
  • Automatyzacja deploymentów Spark na Docker/Kubernetes (Spark Operator).
  • Diagnozowanie problemów z danymi, wydajnością i opóźnieniami procesów.
  • Tworzenie dokumentacji technicznej, runbooków i diagramów przepływu danych.
  • Współpraca w Scrum/Kanban z Data Engineerami, BA, testerami i DevOps; prowadzenie warsztatów i code review.

Inne informacje

Please be informed that the data controller is Hiberus Poland(hereinafter "controller"). You have the right to request access to your personal data, their rectification, erasure or restriction of processing, the right to object to processing, as well as the right to data portability and to lodge a complaint to the supervisory authority. Personal data will be processed for the purpose of the recruitment process. Provision of data to the extent resulting from the Act of 26 June 1974 Labour Code is mandatory. In the remaining scope, providing data is voluntary. Refusal to provide mandatory data may result in the impossibility to carry out the recruitment process. The Administrator processes mandatory data on the basis of a legal obligation incumbent upon him/her, while with regard to additional data, the basis for processing is consent. Personal data will be processed until the recruitment procedure is completed and for the period of the possibility of asserting potential claims, and in the case of consent to participate in future recruitment procedures - until the withdrawal of such consent. Consent to the processing of personal data can be withdrawn at any time. The recipient of the data is the Just Join IT service and other entities to whom we have entrusted the processing of data in connection with recruitment.

Hiberus

Hiberus

40 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz