Senior MLOps / ML Platform Engineer
Tech Stack / Keywords
Firma i stanowisko
We are looking for a Senior MLOps Engineer to join Sigma Software and help build a production-grade ML platform for a large-scale AdTech ecosystem. You will work on infrastructure powering predictive decision-making systems that process hundreds of millions of auction requests daily.
The customer is a technology company operating supply-side infrastructure within the programmatic advertising ecosystem, managing a high-load ad exchange platform and investing in advanced predictive decision-making capabilities.
Sigma Software is building a predictive modeling and optimization platform integrated with a live ad exchange environment, combining large-scale ML infrastructure, automated model lifecycle management, multi-tenant architecture, and advanced observability practices.
Key Technologies: Python, Kubernetes, Docker, GCP, Vertex AI, MLflow, Airflow, Kubeflow, Argo Workflows, Terraform
Wymagania
- 5+ years of experience in MLOps, ML platform engineering, or infrastructure engineering supporting production ML systems
- Strong Python skills and experience building platform-level tooling and automation
- Hands-on experience with Kubernetes and Docker
- Experience building CI/CD pipelines for ML workloads
- Hands-on production experience with MLflow, Kubeflow, Airflow, Argo Workflows, Vertex Pipelines, or similar orchestration and ML lifecycle platforms
- Experience with ML platforms and model lifecycle tools such as Vertex AI, MLflow, or Kubeflow
- Strong understanding of ML observability including drift detection, train/serve skew monitoring, and incident response
- Experience designing or supporting multi-tenant ML systems and isolated model environments
- Experience working with cloud platforms, preferably GCP
- Experience with infrastructure-as-code tools such as Terraform
- Experience with Linux environments
- Understanding of the ML lifecycle and productionization processes
- Upper-Intermediate English level or higher
Nice to have:
- Experience with feature stores and feature consistency management
- Experience with large-scale batch scoring systems operating under freshness SLAs
- Familiarity with experiment tracking platforms and evaluation gates
- Experience with on-premises Kubernetes or bare-metal Linux infrastructure
- Knowledge of DVC, lakeFS, or other data versioning tools
- Experience with Bigtable, Redis, Aerospike, or similar low-latency serving databases
- GPU scheduling and training cost optimization experience
- Familiarity with SOC 2, ISO 27001, or GDPR-related compliance requirements
Obowiązki
- Build and maintain ML training orchestration pipelines across hourly, daily, and weekly schedules
- Implement retries, backfills, and idempotent execution mechanisms
- Design and support model registry workflows including versioning, lineage, evaluation gates, and promotion processes
- Develop isolated per-advertiser model environments with namespace and configuration separation
- Build scalable refresh pipelines and publishing workflows for serving infrastructure
- Implement shadow mode and champion/challenger deployment strategies
- Develop monitoring and alerting for ML-specific metrics including feature drift, prediction drift, train/serve skew, and calibration decay
- Ensure reproducibility of ML workflows using containerized environments, pinned dependencies, and data snapshots
- Monitor training and scoring costs across tenants
- Collaborate with DevOps and SRE engineers on CI/CD and infrastructure automation
- Prepare operational documentation and platform handover materials
Sigma Software
59 aktywnych ofert