Senior Data Scientist
Brak informacji o wynagrodzeniu
SeniorFull-time
#446613·Dodano wczoraj·0
Źródło: Sigma SoftwareTech Stack / Keywords
Machine LearningPythonnumpypandasscikit-learnSQLXGBoostLightGBMCatBoostSpark
Firma i stanowisko
Join Sigma Software to build advanced machine learning solutions for a large-scale player in the programmatic advertising ecosystem. The project involves developing a predictive modeling platform integrated with a live ad exchange processing hundreds of millions of auction requests daily, focusing on bid optimization, calibration, counterfactual evaluation, and constrained decision-making systems. The customer is a technology company managing supply-side infrastructure and investing in predictive decisioning capabilities using advanced machine learning technologies.
Wymagania
- 5+ years of experience in Machine Learning or Data Science with production-grade models measured against business KPIs
- Strong Python skills including numpy, pandas, and scikit-learn
- Strong SQL skills and experience with large-scale datasets
- Experience with XGBoost, LightGBM, or CatBoost
- Strong understanding of regularization, calibration methods, and categorical feature handling
- Strong knowledge of probability, statistics, confidence intervals, and statistical power analysis
- Experience with feature engineering for structured and behavioral datasets
- Hands-on experience with Spark or PySpark
- Practical knowledge of experimentation frameworks and A/B testing methodologies
- Experience with advanced validation approaches including temporal splits, leakage detection, drift analysis, and slice-based metrics
- Understanding of explainability techniques such as SHAP and permutation importance
- Upper-Intermediate English level or higher
Nice to have:
- Experience in AdTech modeling including CTR/CVR prediction, bid-landscape modeling, audience segmentation, and RTB mechanics
- Experience working with sparse, delayed, or censored labels
- Knowledge of attribution modeling, survival analysis, and positive-unlabeled learning
- Practical experience with counterfactual and off-policy evaluation techniques
- Understanding of calibration methods including isotonic regression and Platt scaling
- Experience with hierarchical, empirical-Bayes, or partial-pooling models
- Knowledge of constrained or multi-objective optimization approaches
- Experience with uplift modeling and causal inference methods
- Experience with Vertex AI or similar managed ML training environments
- Publications, competitive modeling achievements, or open-source contributions related to Machine Learning or AdTech
Obowiązki
- Build and improve censored bid-landscape models to estimate clearing-price distributions from partially observed auction data
- Develop real-time win probability estimation models responsive to bid pricing dynamics
- Design and implement hierarchical lift estimation models with confidence-bound-based selection strategies
- Build conversion propensity models using sparse, delayed, and aggregate-only labels
- Develop look-alike audience modeling using positive-unlabeled learning and embedding-based nearest-neighbor techniques
- Implement advertiser-level calibration strategies and monitor ranking and calibration quality independently
- Design robust offline evaluation frameworks using inverse-propensity scoring, doubly-robust estimators, and importance reweighting
- Define exploration strategies and propensity logging approaches for reliable downstream correction and evaluation
- Develop constrained optimization mechanisms for campaign objectives, pricing constraints, and volume targeting
- Contribute to data diagnostics, capability assessments, and evidence-based model recommendations
- Collaborate with customer team during post-launch tuning and performance validation
- Prepare technical documentation and knowledge transfer materials for the customer’s data science team
- Participate in architecture discussions and contribute to scalable ML platform design decisions
Sigma Software
45 aktywnych ofert