Senior MLOps Engineer
Brak informacji o wynagrodzeniu
SeniorFull-time
#441551·Dodano 3 dni temu·2
Źródło: Point WildTech Stack / Keywords
Google Cloud PlatformVertex AIGoogle Kubernetes EngineGKEGoogle Cloud StorageCloud RunGPUTPUTriton Inference ServervLLM MLflow Airflow GitHub Actions Docker Kubernetes Terraform Python SQL
Firma i stanowisko
Point Wild helps customers monitor, manage, and protect against the risks associated with their identities and personal information in a digital world. Backed by WndrCo, Warburg Pincus and General Catalyst, Point Wild is dedicated to creating the world’s most comprehensive portfolio of industry-leading cybersecurity solutions.
Wymagania
- At least 5 years of hands-on experience designing, deploying, and maintaining production ML workloads in cloud environments.
- Deep practical experience with Google Cloud Platform (GCP) ecosystem including Vertex AI, Cloud Storage, GKE, Cloud Run, and IAM/VPC configurations.
- Expertise with containerization tools Docker, Kubernetes/GKE, and serving tools such as Triton, vLLM, and MLflow.
- Proven experience with workflow orchestrators Airflow, Vertex AI Pipelines, and CI/CD tools GitHub Actions, ArgoCD.
- Solid experience managing cloud infrastructure using Terraform.
- Proficiency in Python and SQL for scripting, automation, API development, and data manipulation.
- Experience with monitoring and drift detection tools like Grafana, Prometheus, GCP Cloud Monitoring, or ML observability frameworks.
Nice to have:
- Experience running large-scale LLM or Deep Learning inference and training workloads.
- Certifications such as GCP Professional Machine Learning Engineer or GCP Professional Cloud Architect.
- Familiarity with feature stores such as Feast, Vertex AI Feature Store.
Obowiązki
- Architect and manage scalable GCP-based ML infrastructure using Vertex AI, Google Kubernetes Engine (GKE), Google Cloud Storage (GCS), Cloud Run, and GPU/TPU compute instances.
- Own the end-to-end deployment lifecycle for machine learning models, building high-throughput, low-latency inference services using containerization and serving frameworks such as Triton Inference Server, vLLM, and MLflow.
- Build automated, reproducible pipelines for model training, testing, evaluation, and deployment with tools like Airflow, Vertex AI Pipelines, and GitHub Actions.
- Implement monitoring systems for system health and ML-specific metrics including feature drift and prediction accuracy to enable automated retraining triggers.
- Provide scalable training environments, optimized runtime infrastructure, and standardized deployment templates to support AI and Research Engineers.
- Collaborate with Data Engineers to integrate model pipelines with feature stores, dataset versioning, and stream/batch data workflows.
- Lead the technical transition of AI prototypes to resilient, secure, auto-scaling microservices.
Benefity
- Opportunity to solve real customer problems with cybersecurity solutions.
- Work in a nimble organization with impactful individual contributions.
- Career acceleration with learning opportunities in new technologies, products, and fast-paced growth environment.
Point Wild
3 aktywne oferty