Senior Data Engineer
Brak informacji o wynagrodzeniu
SeniorFull-time
#434751·Dodano 3 dni temu·1
Źródło: Sigma SoftwareTech Stack / Keywords
SQLPythonSparkPySparkAirflowCloud ComposerDagsterBigQueryGCPTerraform
Firma i stanowisko
Join Sigma Software to build large-scale data infrastructure powering a real-time AdTech platform processing hundreds of millions of auction requests daily. The role involves developing predictive modeling and optimization capabilities for a live advertising ecosystem combining streaming and batch pipelines in a cloud-native environment. The customer is a technology company operating supply-side infrastructure within the programmatic advertising ecosystem with a large-scale ad exchange.
Wymagania
- 5+ years of experience in Data Engineering
- At least 2 years of experience working with production ML or large-scale analytics pipelines
- Expert-level SQL skills including window functions and incremental processing patterns
- Strong Python skills for production-grade pipeline development
- Hands-on experience with Spark or PySpark
- Experience designing ETL / ELT pipelines with Airflow, Cloud Composer, Dagster, or similar tools
- Experience working with cloud data warehouses at scale, preferably BigQuery
- Strong understanding of data modeling and point-in-time correctness
- Experience working with event-driven or clickstream datasets at very large scale
- Experience supporting business-critical production pipelines
- Upper-Intermediate English level or higher
Nice to have:
- Experience with GCP services including Dataflow, Pub/Sub, GCS, and Beam
- Experience building streaming or near-real-time ingestion systems
- Understanding of feature stores, train/serve skew, and label leakage prevention
- Experience in AdTech or auction-based environments
- Experience handling delayed or incomplete labels in ML systems
- Experience with dbt or similar transformation frameworks
- Experience delivering solutions into Customer-owned infrastructure
- Knowledge of GDPR/CCPA-related privacy engineering practices
- Experience with experimentation infrastructure and statistical validation pipelines
- Experience working in hybrid cloud/on-prem Linux environments
- Experience with Terraform and Kubernetes
- Experience optimizing warehouse cost and performance
Obowiązki
- Write and defend diagnostic SQL queries against large-scale production datasets
- Build and maintain ingestion pipelines for bid, win, and impression logs into BigQuery
- Harmonize fields across independently designed datasets and maintain versioned field mappings
- Develop point-in-time-correct feature tables and aggregation pipelines
- Design and maintain conversion and labeling pipelines with delayed label handling
- Own the data serving write path, schema contracts, publishing flows, and freshness SLOs
- Build experimentation infrastructure including traffic splitting and reporting pipelines
- Perform large-scale historical backfills and safe reprocessing after mapping changes
- Implement data isolation and safe-aggregation controls for advertiser data protection
- Develop automated data quality validation frameworks
- Collaborate closely with Customer engineers and prepare operational documentation
- Contribute to architecture discussions and platform scalability improvements
Sigma Software
65 aktywnych ofert