Hiberus
Hiberus
New

Spark / Data Lakehouse Developer

20.2k - 26.9k PLN/ mies.B2B
SeniorFull-time·B2B
#425159·Dodano 3 dni temu·1
Źródło: Hiberus
Aplikuj teraz

Tech Stack / Keywords

SparkAICloudSoftware DevelopmentCybersecurity

Firma i stanowisko

We are part of Hiberus, a technology company founded in Spain with a strong presence in the Polish market: 4,700+ specialists, operating in 14+ countries and delivering projects for clients worldwide. We work with technologies including AI, Data, Cloud, Software Development, BI, and Cybersecurity. One of our key advantages is Hiberus University — an international learning environment offering real opportunities to develop and expand your skills. More than 1,000 people were trained last year in selected technologies. Your first project is just the beginning. With Hiberus, you can build a long-term career.

Wymagania

  • 4+ years of experience working with Apache Spark in production environments.
  • Very good knowledge of PySpark, Spark SQL, DataFrames, and Structured Streaming.
  • Experience with MongoDB and the Spark Connector, with a good understanding of data modeling and indexing.
  • Hands-on experience with Apache Iceberg and Data Lakehouse architecture (upsert/merge, write optimization).
  • Experience with CI/CD using Jenkins, including automation of testing and deployment of Spark applications.
  • Very good knowledge of Python and Spark SQL.
  • Knowledge of Prometheus and Grafana for monitoring and observability.
  • Experience working in Scrum/Kanban environments, with practical knowledge of Jira and Confluence.

Obowiązki

  • Design and implementation of Spark SQL/PySpark processes for data validation and comparison (MongoDB/ODS).
  • Implementation of data quality rules, duplicate, anomaly, and missing data detection, and reporting/alerting.
  • Integration and optimization of Spark–MongoDB workflows (MongoDB Spark Connector, partitioning, push-down, schema-aware reads).
  • Design of Raw → Bronze → Silver → Gold data layers using Delta Lake, Iceberg, Hudi, including CI/CD for schemas and migrations.
  • Monitoring and observability of Spark/batch/streaming processes (Prometheus, Grafana, OpenTelemetry, Dynatrace) and definition of SLAs/SLOs.
  • Automation of Spark application deployments using Docker/Kubernetes (Spark Operator).
  • Troubleshooting data quality, performance, and processing latency issues.
  • Creation and maintenance of technical documentation, runbooks, and data flow diagrams.
  • Collaboration in Scrum/Kanban teams with Data Engineers, Business Analysts, QA, and DevOps; conducting workshops and code reviews.

Benefity

  • Multisport card
  • Private medical care
Karta sportowa
Opieka zdrowotna
Hiberus

Hiberus

40 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz