Senior Data Engineer
Brak informacji o wynagrodzeniu
SeniorFull-time
#444749·Dodano 4 dni temu·2
Źródło: emagineTech Stack / Keywords
Azure DatabricksPythonPySparkSpark SQLDelta LakeDatabricks WorkflowsDelta Live TablesUnity CatalogAPIJSON
Firma i stanowisko
This role is for a Senior Data Engineer to join the Labeling Data Platform team within the pharmaceutical industry. The platform manages labeling and eLabel content data supporting multiple systems and regulatory processes.
Wymagania
- Strong engineering experience with Azure Databricks, Python/PySpark, Spark SQL, and Delta Lake.
- Deep knowledge of Bronze → Silver → Gold medallion architecture.
- Experience ingesting structured JSON/API/file data and working with ADLS Gen2 and Azure storage.
- Familiarity with Databricks Workflows/Jobs; Delta Live Tables/Lakeflow is preferred.
- Expertise in schema enforcement, data-quality rules, validation, and reconciliation.
- Skills in data transformation, standardization, metadata management, and lineage.
- Experience with Unity Catalog and data governance.
- Competence in API-based ingestion and export patterns.
- Handling document metadata and references to PDF/eLabel assets.
- Performance optimization and partitioning of data workloads.
- Practice with CI/CD for notebooks/code and automated data tests.
- Monitoring, error handling, and reprocessing capabilities.
- Experience supporting regulated data pipelines ensuring traceability, auditability, reproducibility, and controlled releases.
- Background in working within regulated environments adhering to compliance and quality standards.
Obowiązki
- Develop and maintain data pipelines using Azure Databricks, Python, PySpark, Spark SQL, and Delta Lake.
- Design and implement solutions following the Bronze → Silver → Gold medallion architecture.
- Ingest and process data from APIs, JSON files, Azure Storage, and enterprise systems.
- Build and maintain data quality, validation, reconciliation, monitoring, and automated data testing frameworks.
- Optimize performance, partitioning, and scalability of data workloads.
- Implement CI/CD, automated testing, error handling, and reprocessing processes.
- Ensure traceability, auditability, reproducibility, and controlled releases for regulated data pipelines.
emagine
868 aktywnych ofert