MLOps Engineer

Brak informacji o wynagrodzeniu
SeniorFull-time
#409095·Dodano 10 dni temu·0
Źródło: justjoin.it
Aplikuj teraz

Tech Stack / Keywords

MLflowDockerKubernetesTerraformAnsibleLLM

Firma i stanowisko

Andersen is hiring an MLOps Engineer for a project delivering scalable digital solutions and supporting reliable AI and technology platforms across diverse organizations.

The customer is a government-backed organization focused on socio-economic development through digital and business solutions, aiming to improve labor markets, enhance employment opportunities, and support public and private sector growth. Their services include digital platforms, automation, and training programs to develop human capital.

The project delivers digital and technology solutions for government, semi-government, and private-sector organizations, supporting diverse business needs with reliable, scalable, and efficient solutions.

Wymagania

  • Strong engineering experience across Software Engineering, DevOps/SRE, and MLOps for 5+ years.
  • Production ownership of machine learning systems: operated, supported, and accountable for ML in production.
  • Deep hands-on experience with Docker and Kubernetes, including resource management and GPU workloads.
  • Strong proficiency with GitLab CI, GitHub Actions, ArgoCD, or equivalent, including pipeline-enforced quality gates.
  • Extensive experience with MLflow, Kubeflow, Feast, BentoML, KServe, SageMaker, Vertex AI, Databricks, or comparable platforms.
  • Production engineering standard Python (typed, tested, packaged, and reviewed).
  • Infrastructure as Code (Terraform, Ansible, or equivalent) with real module and state management experience.
  • Production service engineering with stateless, configuration-driven services with structured logging, published API contracts, and centrally managed secrets.
  • Data services in production (PostgreSQL, caching, and message-driven asynchronous processing).
  • Monitoring and observability (metrics, logging, error tracking, and ML-aware monitoring such as Evidently, WhyLabs, or Arize).
  • Proven collaboration with platform or infrastructure engineering teams.
  • Hands-on LLMOps experience: model gateways, prompt versioning, RAG pipelines, vector databases, evaluation harnesses, and guardrail frameworks.
  • Generative AI consumption management: token metering, cost attribution, budget and quota controls, and cost optimization of LLM workloads.
  • Model evaluation and selection: building evaluation suites comparing models on quality, latency, cost, and safety.
  • Practical generative AI engineering: prompt and structured output design, context management, embedding selection, tool calling, and fine-tuning or model adaptation.
  • Experience operating agentic AI systems in production.
  • Experience establishing an ML platform capability from an early stage.
  • Cost engineering for AI workloads: GPU efficiency, inference optimization, and consumption attribution.
  • Familiarity with AI governance and regulated environments — model risk management, auditability, and fairness testing.
  • Experience with data quality frameworks and workflow orchestration.
  • English language: Upper-Intermediate and above.

Obowiązki

  • Designing and owning the end-to-end MLOps architecture for production Machine Learning and Generative AI systems.
  • Building and maintaining ML training, validation, deployment, serving, monitoring, and retraining pipelines.
  • Defining and implementing the model promotion lifecycle across development, staging, and production environments.
  • Building and maintaining model registry, experiment tracking, dataset/model versioning, and reproducible ML workflows.
  • Designing scalable training and inference infrastructure, including GPU-backed workloads.
  • Building and maintaining CI/CD pipelines for ML and AI workloads with automated testing, quality gates, approval controls, and rollback mechanisms.
  • Implementing progressive delivery approaches for ML models, including shadow testing, canary releases, and blue-green deployments.
  • Collaborating closely with Platform Engineering to run AI workloads on shared Kubernetes infrastructure.
  • Translating ML infrastructure requirements into technical requirements for platform, security, capacity, and architecture teams.
  • Establishing reusable MLOps project templates, shared pipeline components, and engineering standards.
  • Implementing automated ML testing, including data validation, model regression tests, training-serving consistency checks, and evaluation gates.
  • Owning production reliability of ML systems, including availability, latency, throughput, scalability, and operational stability.
  • Building observability across infrastructure, data quality, model performance, and business impact.
  • Configuring monitoring, alerting, incident response, runbooks, rollback procedures, and post-incident reviews for AI systems.
  • Implementing automated model retraining based on schedules, events, and data/model drift.
  • Managing the full model lifecycle, including deployment, monitoring, retraining, version promotion, and retirement.
  • Tracking and optimizing infrastructure, training, inference, and GPU-related costs.
  • Building and operating the production layer for Generative AI and LLM-based applications.
  • Implementing LLM model gateways, routing, prompt management, prompt versioning, caching, and rate limiting.
  • Implementing LLM guardrails, groundedness monitoring, sensitive-data protection, and prompt-injection mitigation.
  • Building observability and permission controls for agentic AI systems.
  • Implementing token-level and request-level consumption metering across Generative AI applications.
  • Building cost attribution mechanisms by business unit, use case, application, model, tenant, and environment.
  • Optimizing LLM workloads through model routing, prompt/context optimization, caching, batching, and selection of cost-efficient models.
  • Applying hands-on Generative AI engineering practices, including prompt engineering, structured outputs, context management, embeddings, tool/function calling, and model adaptation where required.
  • Evaluating emerging ML and Generative AI tools, models, and platforms and recommending technologies for adoption.
  • Making and documenting architectural and build-vs-configure decisions for new AI capabilities.
  • Producing architecture decision records, reference designs, technical documentation, and operational guidelines.
  • Participating in architecture, capacity planning, technical roadmap, and platform strategy discussions.
  • Performing code reviews, technical mentoring, pairing, and knowledge sharing to improve engineering standards across the team.

Benefity

  • Experience working with leaders in FinTech, Healthcare, Retail, Telecom, and other sectors.
  • Opportunity to change project or develop expertise in various business domains.
  • Work options: fully remote, office, or hybrid.
  • Professional, financial, and career growth with mentoring and adaptation systems.
  • Opportunity to earn up to an additional 1,000 EUR per month included in annual bonus.
  • Access to a corporate training portal with a continually updated knowledge base.
  • Bright corporate life: parties, pizza days, PlayStation, fruits, coffee, snacks, and movies.
  • Certification compensation (e.g., AWS, PMP).
  • Referral program.
  • Private health insurance and sports compensation depending on employment type.
Budżet konferencyjny
Dofinansowanie szkoleń
Spotkania integracyjne
Opieka zdrowotna
Karta sportowa
Ubezpieczenie
Płatne święta
Premie

Inne informacje

Informujemy, że administratorem danych jest Andersen Soft UAB z siedzibą w Krakow, ul. Al. Pokoju 18, 31 - 564 dalej jako "administrator"). Masz prawo do żądania dostępu do swoich danych osobowych, ich sprostowania, usunięcia lub ograniczenia przetwarzania, prawo do wniesienia sprzeciwu wobec przetwarzania, a także prawo do przenoszenia danych oraz wniesienia skargi do organu nadzorczego. Dane osobowe przetwarzane będą w celu realizacji procesu rekrutacji. Podanie danych w zakresie wynikającym z ustawy z dnia 26 czerwca 1974 r. Kodeks pracy jest obowiązkowe. W pozostałym zakresie podanie danych jest dobrowolne. Odmowa podania danych obowiązkowych może skutkować brakiem możliwości przeprowadzenia procesu rekrutacji. Administrator przetwarza dane obowiązkowe na podstawie ciążącego na nim obowiązku prawnego, zaś w zakresie danych dodatkowych podstawą przetwarzania jest zgoda. Dane osobowe będą przetwarzane do czasu zakończenia postępowania rekrutacyjnego i przez okres możliwości dochodzenia ewentualnych roszczeń, a w przypadku wyrażenia zgody na udział w przyszłych postępowaniach rekrutacyjnych - do czasu wycofania tej zgody. Zgoda na przetwarzanie danych osobowych może zostać wycofana w dowolnym momencie. Odbiorcą danych jest serwis Just Join IT oraz inne podmioty, którym powierzyliśmy przetwarzanie danych w związku z rekrutacją.

Andersen

Andersen

60 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz