Data Engineer
Tech Stack / Keywords
Firma i stanowisko
The customer is a technology company developing scalable software solutions for organizations across multiple industries, helping businesses improve operational efficiency, optimize workflows, and support digital transformation through modern technologies and data-driven approaches.
Wymagania
- Commercial experience in Data Engineering with ownership of end-to-end ELT/ETL pipelines and cloud data warehouse architecture for 4+ years.
- Expert-level SQL skills.
- Production experience with Amazon Redshift, including schema design, distribution/sort keys, WLM, query tuning, and reducing refresh latency on live warehouses.
- Strong MPP data warehouse experience; Snowflake and/or Azure Synapse SQL Pool experience valued but not a substitute for Redshift.
- Experience with incremental loads and CDC techniques such as watermarks, SCD2, change-data-capture, moving toward near-real-time processing.
- Experience re-architecting legacy ETL pipelines and migrating from script-based or on-prem legacy systems to scalable cloud platforms with minimal disruption.
- Experience using MySQL as a source system, including incremental extraction; familiarity with MariaDB.
- Experience with AWS S3 and IAM for data landing and access management.
- Production experience with orchestration tools like Airflow and dbt.
- Ability to work fully independently as sole owner of a platform, identifying and prioritizing technical debt without formal ticket tracking.
- Strong monitoring and failure recovery practices for maintaining healthy batch platforms.
- English language proficiency from Upper-Intermediate and above.
Nice to have:
- CDC tooling (e.g., Debezium, binlog-based) in production.
- Experience mentoring junior engineers and PR-review rigor (naming, readability, reuse, runtime budgets).
Obowiązki
- Monitoring and troubleshooting the batch pipelines day to day, checking every run, debugging failures.
- Tracing failures through the medallion layers (bronze/silver/gold), fixing understandable issues, validating in test environments before production.
- Setting up failure alerting with notifications showing batch success or failure and error details.
- Working towards more frequent data refresh cadence through incremental-load strategies.
- Handling requests to ingest new data sources independently, including setting up connections, landing data in S3, schema design in Redshift, and granting IAM access.
- Extracting data from sources like Postgres, MySQL, Google Sheets, Salesforce, and HR software via Glue, loading into Redshift, and transforming with dbt.
- Provisioning data for ML engineers and analysts by creating views or materialized views and managing access.
- Reviewing pull requests and mentoring junior engineers regarding naming, readability, reuse, and ETL runtime budgets.
- Onboarding to review pull requests and troubleshoot following the platform's architecture and philosophy.
Benefity
- Opportunity to work with leaders in FinTech, Healthcare, Retail, Telecom, and others.
- Ability to change projects or develop expertise in interesting business domains.
- Flexible work conditions: fully remote, office-based, or hybrid.
- Systems of mentoring and adaptation for new employees.
- Potential to earn up to an additional 1,000 EUR/month in annual bonus based on expertise.
- Access to a corporate training portal with an extensive and updated knowledge base.
- Corporate social events including parties and refreshments.
- Certification compensation (e.g., AWS, PMP).
- Referral program.
- Private health insurance and sports compensation, depending on employment type.
Inne informacje
Informujemy, że administratorem danych jest Andersen Soft UAB z siedzibą w Krakow, ul. Al. Pokoju 18, 31 - 564 dalej jako "administrator"). Masz prawo do żądania dostępu do swoich danych osobowych, ich sprostowania, usunięcia lub ograniczenia przetwarzania, prawo do wniesienia sprzeciwu wobec przetwarzania, a także prawo do przenoszenia danych oraz wniesienia skargi do organu nadzorczego. Dane osobowe przetwarzane będą w celu realizacji procesu rekrutacji. Podanie danych w zakresie wynikającym z ustawy z dnia 26 czerwca 1974 r. Kodeks pracy jest obowiązkowe. W pozostałym zakresie podanie danych jest dobrowolne. Odmowa podania danych obowiązkowych może skutkować brakiem możliwości przeprowadzenia procesu rekrutacji. Administrator przetwarza dane obowiązkowe na podstawie ciążącego na nim obowiązku prawnego, zaś w zakresie danych dodatkowych podstawą przetwarzania jest zgoda. Dane osobowe będą przetwarzane do czasu zakończenia postępowania rekrutacyjnego i przez okres możliwości dochodzenia ewentualnych roszczeń, a w przypadku wyrażenia zgody na udział w przyszłych postępowaniach rekrutacyjnych - do czasu wycofania tej zgody. Zgoda na przetwarzanie danych osobowych może zostać wycofana w dowolnym momencie. Odbiorcą danych jest serwis Just Join IT oraz inne podmioty, którym powierzyliśmy przetwarzanie danych w związku z rekrutacją.
Andersen
64 aktywne oferty