Senior Data Engineer

Brak informacji o wynagrodzeniu
SeniorFull-time
#434751·Dodano 3 dni temu·1
Źródło: Sigma Software
Aplikuj teraz

Tech Stack / Keywords

SQLPythonSparkPySparkAirflowCloud ComposerDagsterBigQueryGCPTerraform

Firma i stanowisko

Join Sigma Software to build large-scale data infrastructure powering a real-time AdTech platform processing hundreds of millions of auction requests daily. The role involves developing predictive modeling and optimization capabilities for a live advertising ecosystem combining streaming and batch pipelines in a cloud-native environment. The customer is a technology company operating supply-side infrastructure within the programmatic advertising ecosystem with a large-scale ad exchange.

Wymagania

  • 5+ years of experience in Data Engineering
  • At least 2 years of experience working with production ML or large-scale analytics pipelines
  • Expert-level SQL skills including window functions and incremental processing patterns
  • Strong Python skills for production-grade pipeline development
  • Hands-on experience with Spark or PySpark
  • Experience designing ETL / ELT pipelines with Airflow, Cloud Composer, Dagster, or similar tools
  • Experience working with cloud data warehouses at scale, preferably BigQuery
  • Strong understanding of data modeling and point-in-time correctness
  • Experience working with event-driven or clickstream datasets at very large scale
  • Experience supporting business-critical production pipelines
  • Upper-Intermediate English level or higher

Nice to have:

  • Experience with GCP services including Dataflow, Pub/Sub, GCS, and Beam
  • Experience building streaming or near-real-time ingestion systems
  • Understanding of feature stores, train/serve skew, and label leakage prevention
  • Experience in AdTech or auction-based environments
  • Experience handling delayed or incomplete labels in ML systems
  • Experience with dbt or similar transformation frameworks
  • Experience delivering solutions into Customer-owned infrastructure
  • Knowledge of GDPR/CCPA-related privacy engineering practices
  • Experience with experimentation infrastructure and statistical validation pipelines
  • Experience working in hybrid cloud/on-prem Linux environments
  • Experience with Terraform and Kubernetes
  • Experience optimizing warehouse cost and performance

Obowiązki

  • Write and defend diagnostic SQL queries against large-scale production datasets
  • Build and maintain ingestion pipelines for bid, win, and impression logs into BigQuery
  • Harmonize fields across independently designed datasets and maintain versioned field mappings
  • Develop point-in-time-correct feature tables and aggregation pipelines
  • Design and maintain conversion and labeling pipelines with delayed label handling
  • Own the data serving write path, schema contracts, publishing flows, and freshness SLOs
  • Build experimentation infrastructure including traffic splitting and reporting pipelines
  • Perform large-scale historical backfills and safe reprocessing after mapping changes
  • Implement data isolation and safe-aggregation controls for advertiser data protection
  • Develop automated data quality validation frameworks
  • Collaborate closely with Customer engineers and prepare operational documentation
  • Contribute to architecture discussions and platform scalability improvements
Sigma Software

Sigma Software

65 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz