Senior MLOps Engineer

Brak informacji o wynagrodzeniu
SeniorFull-time
#441551·Dodano 3 dni temu·2
Źródło: Point Wild
Aplikuj teraz

Tech Stack / Keywords

Google Cloud PlatformVertex AIGoogle Kubernetes EngineGKEGoogle Cloud StorageCloud RunGPUTPUTriton Inference ServervLLM MLflow Airflow GitHub Actions Docker Kubernetes Terraform Python SQL

Firma i stanowisko

Point Wild helps customers monitor, manage, and protect against the risks associated with their identities and personal information in a digital world. Backed by WndrCo, Warburg Pincus and General Catalyst, Point Wild is dedicated to creating the world’s most comprehensive portfolio of industry-leading cybersecurity solutions.

Wymagania

  • At least 5 years of hands-on experience designing, deploying, and maintaining production ML workloads in cloud environments.
  • Deep practical experience with Google Cloud Platform (GCP) ecosystem including Vertex AI, Cloud Storage, GKE, Cloud Run, and IAM/VPC configurations.
  • Expertise with containerization tools Docker, Kubernetes/GKE, and serving tools such as Triton, vLLM, and MLflow.
  • Proven experience with workflow orchestrators Airflow, Vertex AI Pipelines, and CI/CD tools GitHub Actions, ArgoCD.
  • Solid experience managing cloud infrastructure using Terraform.
  • Proficiency in Python and SQL for scripting, automation, API development, and data manipulation.
  • Experience with monitoring and drift detection tools like Grafana, Prometheus, GCP Cloud Monitoring, or ML observability frameworks.

Nice to have:

  • Experience running large-scale LLM or Deep Learning inference and training workloads.
  • Certifications such as GCP Professional Machine Learning Engineer or GCP Professional Cloud Architect.
  • Familiarity with feature stores such as Feast, Vertex AI Feature Store.

Obowiązki

  • Architect and manage scalable GCP-based ML infrastructure using Vertex AI, Google Kubernetes Engine (GKE), Google Cloud Storage (GCS), Cloud Run, and GPU/TPU compute instances.
  • Own the end-to-end deployment lifecycle for machine learning models, building high-throughput, low-latency inference services using containerization and serving frameworks such as Triton Inference Server, vLLM, and MLflow.
  • Build automated, reproducible pipelines for model training, testing, evaluation, and deployment with tools like Airflow, Vertex AI Pipelines, and GitHub Actions.
  • Implement monitoring systems for system health and ML-specific metrics including feature drift and prediction accuracy to enable automated retraining triggers.
  • Provide scalable training environments, optimized runtime infrastructure, and standardized deployment templates to support AI and Research Engineers.
  • Collaborate with Data Engineers to integrate model pipelines with feature stores, dataset versioning, and stream/batch data workflows.
  • Lead the technical transition of AI prototypes to resilient, secure, auto-scaling microservices.

Benefity

  • Opportunity to solve real customer problems with cybersecurity solutions.
  • Work in a nimble organization with impactful individual contributions.
  • Career acceleration with learning opportunities in new technologies, products, and fast-paced growth environment.
Point Wild

Point Wild

3 aktywne oferty

Zobacz wszystkie oferty
Aplikuj teraz