Data Engineer Consultant

210 - 240 PLN/ godz.
SeniorFull-time
#448899·Dodano 9 dni temu·2
Źródło: Chabre IT Services
Aplikuj teraz

Tech Stack / Keywords

DatabricksPySparkDelta LakeUnity CatalogPythonSQLREST APIMLflowAzure DevOpsGitHub Actions

Firma i stanowisko

Chabre IT Services is a global professional IT services provider specializing in tailor-made solutions, smart outsourcing, try&hire, and success fee services. The client is a global technology services organisation delivering data and digital transformation projects focusing on building and enhancing a modern data platform based on Databricks, emphasizing data engineering, governance, automation, and analytics.

Wymagania

  • Recent hands-on experience delivering production solutions with Databricks.
  • Strong knowledge of PySpark, Delta Lake, and Unity Catalog.
  • Practical experience with the Databricks medallion architecture and production data pipelines.
  • Experience with Spark Declarative Pipelines / Delta Live Tables and Databricks Workflows.
  • Very good Python and advanced SQL skills.
  • Experience integrating REST APIs including authentication, pagination, rate limiting, and incremental or checkpointed data retrieval.
  • Strong understanding of data modelling including SCD Type 2 and current-state/CDC patterns.
  • Experience designing data quality controls, reconciliation processes, and handling invalid or quarantined data.
  • Knowledge of Git-based development and CI/CD for data platforms, preferably with Azure DevOps or GitHub Actions.
  • Experience implementing monitoring and observability for production data pipelines.
  • Ability to work within a fixed project scope, estimate work, meet milestones, and take ownership of deliverables.
  • Strong communication skills and ability to work independently in a consulting environment.
  • Excellent English and Polish skills (minimum B2 level).

Obowiązki

  • Design and implement data ingestion pipelines connecting external APIs and business systems with the Databricks data platform.
  • Develop reliable processes for retrieving and processing historical data, including rate limiting, checkpointing, retries, and safe reprocessing.
  • Build and maintain the Bronze layer of the data platform covering data standardisation, typing, historical tracking, and document storage.
  • Develop processes for linking digital documents with business records ensuring data consistency and idempotent processing.
  • Use Databricks capabilities to classify documents and extract structured information from PDFs and images with confidence scoring and manual review workflows.
  • Contribute to the Silver layer development using entity modelling, CDC and incremental processing patterns.
  • Implement deterministic entity resolution and source-priority logic across multiple data sources.
  • Apply data quality controls within the existing framework to ensure quality results are stored, monitored, and available for governance.
  • Support development of analytical and machine-learning data products including feature and metric stores and batch scoring workflows.
  • Work with MLflow and related processes for model registration and promotion.
  • Improve automation, monitoring, and reliability across data pipelines and contribute to CI/CD processes.
  • Collaborate with technical stakeholders to deliver agreed project milestones within consulting scope.

Benefity

  • Rate up to 240 PLN/h + VAT.
  • Remote work.
  • Subsidy for peripherals amounting to 500 PLN.
  • Working tool (MacBook Pro or Lenovo Legion 5).
  • Co-financing of courses related to the position.
  • Benefits including MultiSport and Medicover.
Karta sportowa
Opieka zdrowotna
Chabre IT Services

Chabre IT Services

35 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz