Mid Data Engineer (Databricks)
120 - 150 PLN/ godz.B2B
MidFull-time·B2B
#426625·Dodano wczoraj·0
Źródło: nofluffjobs.comTech Stack / Keywords
PythonSQLSparkAIPostgreSQLMongoDBData modelling
Firma i stanowisko
Acaisoft is a software house cooperating with a US BioPharma client to deliver cloud software and services enhancing AI/ML capabilities in the BioPharma sector.
Wymagania
- Strong hands-on experience with Python, complex SQL, and production Spark on Databricks (Structured Streaming, Auto Loader, Delta Lake, Unity Catalog).
- Experience handling structured and unstructured data through ingestion, transformation, enrichment, indexing, and retention.
- Proven expertise ingesting from heterogeneous operational sources including relational (Oracle, PostgreSQL) and document stores (MongoDB) using Full load + CDC or equivalent replication patterns.
- Skills in data modelling across layered architecture with clear contracts between raw, curated, and serving layers supporting cross-entity queries and performant data access patterns.
Nice to have:
- Experience building unstructured document pipelines: parsing and extraction from PDF, Office, and scanned formats, chunking, embedding generation, and maintaining searchable indexes at scale.
Obowiązki
- Build, evolve, and operate data ingestion and processing capabilities for structured, semi-structured, and unstructured data, supporting transitions from prototypes through general release.
- Implement and maintain metadata and data quality practices enabling cross-record querying, traceability, and AI-ready data access across experiments, files, inventory, and workflows.
- Collaborate with architects, AI engineers, workflow engineers, and domain experts to ensure data usability, performance, and trustworthiness for GenAI and analytics use cases.
- Support engineering quality through code reviews and mentoring junior/mid-level engineers.
Inne informacje
Due to the client's location, the team works till 6:00 PM CEST.
Acaisoft
8 aktywnych ofert