Senior MLOps Engineer

Brak informacji o wynagrodzeniu
SeniorInne·B2B
#442225·Dodano 3 dni temu·1
Źródło: justjoin.it
Aplikuj teraz

Tech Stack / Keywords

GCPGKETroubleshootingPythonDockerMLOpsTritonvLLMVertexAI

Firma i stanowisko

Point Wild is a cybersecurity company backed by WndrCo, Warburg Pincus, and General Catalyst, focused on delivering industry-leading solutions to monitor, manage, and protect personal identities and information in the digital world.

Wymagania

  • At least 5 years of experience in designing, deploying, and maintaining production ML workloads in cloud environments.
  • Deep practical experience with Google Cloud Platform ecosystem including Vertex AI, Cloud Storage, GKE, Cloud Run, and IAM/VPC configurations.
  • Expertise with containerization using Docker, Kubernetes/GKE, and serving tools (Triton, vLLM, MLflow).
  • Experience with workflow orchestrators (Airflow, Vertex AI Pipelines) and CI/CD tools (GitHub Actions, ArgoCD).
  • Experience managing cloud resources using Terraform.
  • Proficiency in Python and SQL for scripting, automation, API development, and data manipulation.
  • Hands-on experience with logging, telemetry, and drift detection tools such as Grafana, Prometheus, and GCP Cloud Monitoring.

Nice to have:

  • Experience running large-scale LLM or Deep Learning inference/training workloads.
  • GCP Professional Machine Learning Engineer or GCP Professional Cloud Architect certifications.
  • Familiarity with feature stores like Feast or Vertex AI Feature Store.

Obowiązki

  • Architect and manage scalable GCP-based ML infrastructure using Vertex AI, Google Kubernetes Engine (GKE), Google Cloud Storage (GCS), Cloud Run, and GPU/TPU compute instances.
  • Own the end-to-end deployment lifecycle for ML models, building inference services using containerization and serving frameworks such as Triton Inference Server, vLLM, and MLflow.
  • Build automated pipelines for training, testing, evaluation, and deployment with tools like Airflow, Vertex AI Pipelines, and GitHub Actions.
  • Implement monitoring for system health and ML-specific metrics to enable automated retraining.
  • Provide scalable training environments and standardized deployment templates for AI engineers.
  • Collaborate with Data Engineers to integrate model pipelines with feature stores, dataset versioning, and data workflows.
  • Lead technical transition from AI prototypes to resilient, auto-scaling microservices.

Inne informacje

Please be informed that the data controller is Point Wild (hereinafter "controller"). You have the right to request access to your personal data, their rectification, erasure or restriction of processing, the right to object to processing, as well as the right to data portability and to lodge a complaint to the supervisory authority. Personal data will be processed for the purpose of the recruitment process. Provision of data to the extent resulting from the Act of 26 June 1974 Labour Code is mandatory. In the remaining scope, providing data is voluntary. Refusal to provide mandatory data may result in the impossibility to carry out the recruitment process. The Administrator processes mandatory data on the basis of a legal obligation incumbent upon him/her, while with regard to additional data, the basis for processing is consent. Personal data will be processed until the recruitment procedure is completed and for the period of the possibility of asserting potential claims, and in the case of consent to participate in future recruitment procedures - until the withdrawal of such consent. Consent to the processing of personal data can be withdrawn at any time. The recipient of the data is the Just Join IT service and other entities to whom we have entrusted the processing of data in connection with recruitment.

Point Wild

Point Wild

3 aktywne oferty

Zobacz wszystkie oferty
Aplikuj teraz