Senior Data Engineer (AI/LLM)
150 - 170 PLN/ godz.B2B
SeniorFull-time·B2B
#402138·Dodano 2 dni temu·0
Źródło: nofluffjobs.comTech Stack / Keywords
PythonDockerSQLDatabricksPySparkCI/CDGCPAirflowBigQueryAPIFastAPIAzure DatabricksAzureKubernetes
Firma i stanowisko
DS STREAM is developing an advanced data ecosystem in a multicloud environment (Azure and GCP) that supports production deployments of generative AI, large language models (LLM), and Retrieval-Augmented Generation (RAG) architectures. The platform uses technologies including Azure Databricks, GCP Vertex AI (Vector Search), Unity Catalog, and Spark Structured Streaming, with a Medallion Architecture design for efficient management of structured and unstructured data.
Wymagania
- Advanced knowledge of Python and SQL for building scalable data pipelines.
- Minimum 3 years practical experience with Databricks, PySpark, Delta Lake, including familiarity with Unity Catalog and Databricks Asset Bundles.
- Experience building and maintaining DAGs in Apache Airflow.
- Practical knowledge of cloud data components in Azure (Data Factory, Databricks, Blob/ADLS, AI Search) and GCP (BigQuery, Cloud Composer, Vertex AI, Cloud Storage).
- Experience with data streaming and API data integration, specifically in real-time processing with Spark Structured Streaming.
- Understanding of LLM/RAG architectures, document chunking, embeddings, and vector databases.
- Software engineering skills including Git, Docker, CI/CD, and Infrastructure as Code (Terraform).
Nice to have:
- Certifications: Databricks Certified Data Engineer Senior/Associate, Azure Data Engineer Associate (DP-203), or Google Cloud Professional Data Engineer.
Obowiązki
- Develop scalable data processing software and pipelines for analytics and AI workloads.
- Design and develop data pipelines in Databricks using PySpark and SQL.
- Create and maintain streaming data pipelines using Spark Structured Streaming.
- Design and manage Unity Catalog objects according to Medallion Architecture (Bronze, Silver, Gold).
- Integrate and optimize data ingestion from diverse API interfaces, including external systems for AI use cases.
- Build ETL/ELT processes supporting extraction, cleaning, chunking, and embedding of textual data into vector databases in cloud platforms (Azure AI Search, GCP Vertex AI).
Benefity
- 100% remote work within Poland with optional use of office in Warsaw.
- Competitive compensation.
- Sport subscription.
- Training budget.
- Private healthcare.
- Flat organizational structure.
- Small teams.
- International projects.
- Free coffee and beverages.
- Bicycle parking.
- In-house trainings and hack days.
- Modern office environment.
- Startup atmosphere.
- No dress code.
Karta sportowa
Dofinansowanie szkoleń
Opieka zdrowotna
Napoje w biurze
Parking dla rowerów
Szkolenia wewnętrzne
DS STREAM
7 aktywnych ofert