Senior Data Engineer
Tech Stack / Keywords
Firma i stanowisko
Join Sigma Software to build large-scale data infrastructure powering a real-time AdTech platform processing hundreds of millions of auction requests daily. The company operates supply-side infrastructure within the programmatic advertising ecosystem managing a large-scale ad exchange with predictive decisioning technologies. The project focuses on building a predictive modeling and optimization platform on top of a live ad exchange environment with real-time scoring, filtering, and multi-objective optimization capabilities. The solution processes massive-scale event and auction datasets including feature engineering pipelines, streaming and batch ingestion, experimentation infrastructure, and ML-oriented data services.
Wymagania
- 5+ years of experience in Data Engineering
- At least 2 years of experience working with production ML or large-scale analytics pipelines
- Expert-level SQL skills including window functions and incremental processing patterns
- Strong Python skills for production-grade pipeline development
- Hands-on experience with Spark or PySpark
- Experience designing ETL / ELT pipelines with Airflow, Cloud Composer, Dagster, or similar tools
- Experience working with cloud data warehouses at scale, preferably BigQuery
- Strong understanding of data modeling and point-in-time correctness
- Experience with event-driven or clickstream datasets at very large scale
- Experience supporting business-critical production pipelines
- Upper-Intermediate English level or higher
Nice to have:
- Experience with GCP services including Dataflow, Pub/Sub, GCS, and Beam
- Experience building streaming or near-real-time ingestion systems
- Understanding of feature stores, train/serve skew, and label leakage prevention
- Experience in AdTech or auction-based environments
- Experience handling delayed or incomplete labels in ML systems
- Experience with dbt or similar transformation frameworks
- Experience delivering solutions into customer-owned infrastructure
- Knowledge of GDPR/CCPA-related privacy engineering practices
- Experience with experimentation infrastructure and statistical validation pipelines
- Experience working in hybrid cloud/on-prem Linux environments
- Terraform and Kubernetes experience
- Experience optimizing warehouse cost and performance
Obowiązki
- Write and defend diagnostic SQL queries against large-scale production datasets
- Build and maintain ingestion pipelines for bid, win, and impression logs into BigQuery
- Harmonize fields across independently designed datasets and maintain versioned field mappings
- Develop point-in-time-correct feature tables and aggregation pipelines
- Design and maintain conversion and labeling pipelines with delayed label handling
- Own the data serving write path, schema contracts, publishing flows, and freshness SLOs
- Build experimentation infrastructure including traffic splitting and reporting pipelines
- Perform large-scale historical backfills and safe reprocessing after mapping changes
- Implement data isolation and safe-aggregation controls for advertiser data protection
- Develop automated data quality validation frameworks
- Collaborate closely with customer engineers and prepare operational documentation
- Contribute to architecture discussions and platform scalability improvements
Sigma Software
59 aktywnych ofert