Databricks Data Architect
Tech Stack / Keywords
Firma i stanowisko
Our client is building a unified Master Data Management platform consolidating data into a single source of truth, with an AI layer integrated to deliver clean and enriched data. The project domain includes Occupational Health & Safety, Incident Management, Risk Management, and global regulatory frameworks, aiming to provide trustworthy and explainable data for compliance and risk assessment.
Wymagania
- Strong production experience with Databricks Lakehouse architecture, including Spark, PySpark, SQL, Delta Lake, and Unity Catalog
- Hands-on experience designing and building ETL/ELT pipelines for batch and incremental ingestion, cleansing, normalization, deduplication, and enrichment
- Practical experience with MDM concepts such as golden records, survivorship/merge rules, trust ranking, identity resolution, duplicate detection, SCD, and exception workflows
- Strong data modeling skills for analytical, operational, and semantic consumption patterns
- Experience designing a semantic layer with shared business definitions, governed metrics, reusable dimensions, and consistent entity definitions
- Experience with data quality and observability including pipeline SLAs, schema drift, CDC, data contracts, dead-letter handling, and source-to-master reconciliation
- Experience implementing data governance and security: Unity Catalog lineage, RBAC/ABAC, row/column-level security, PII handling, and regulatory traceability
- Ability to translate business requirements from product, compliance, and engineering stakeholders into scalable data architecture
Nice to have:
- Experience with Databricks Lakeflow Connect, Lakeflow Spark Declarative Pipelines, and Lakeflow Jobs
- Experience with Unity Catalog metric views or comparable semantic-layer technologies
- Experience with knowledge graphs and graph analytics such as GraphFrames or graph-based entity resolution
- Experience building AI/RAG solutions over enterprise data using AI Search / Vector Search, embeddings, metadata filtering, retrieval evaluation, and source-grounded generation with citations
- Experience with ML-based data enrichment, classification, anomaly detection, or entity matching
- Experience in regulated domains like occupational health and safety, incident management, risk, compliance, ESG, insurance, healthcare, or industrial operations
Obowiązki
- Own platform design and architecture
- Build reference ETL/ELT pipelines and data models
- Define AI/RAG patterns collaboratively with engineering team
Benefity
- Projects for clients including PayPal, Wargaming, Xerox, Philips, Adidas, and Toyota
- Competitive compensation dependent on qualifications and skills
- Career development system with clear skill qualifications
- Flexible working hours aligned to your schedule
- Options to work remotely
Inne informacje
Informujemy, że administratorem danych jest Itransition z siedzibą w Warsaw, Al. Jerozolimskie 123a. Masz prawo do żądania dostępu do swoich danych osobowych, ich sprostowania, usunięcia lub ograniczenia przetwarzania, prawo do wniesienia sprzeciwu wobec przetwarzania, a także prawo do przenoszenia danych oraz wniesienia skargi do organu nadzorczego. Dane osobowe przetwarzane będą w celu realizacji procesu rekrutacji. Podanie danych w zakresie wynikającym z Kodeksu pracy jest obowiązkowe. Odmowa podania danych może skutkować brakiem możliwości przeprowadzenia procesu rekrutacji. Dane będą przetwarzane do zakończenia rekrutacji i okresu dochodzenia roszczeń, a w przypadku zgody na udział w przyszłych rekrutacjach - do wycofania tej zgody.
Itransition
34 aktywne oferty