AI Platform & LLMOps Expert (Principal / Lead)

170 - 230 PLN/ godz.B2B
SeniorFull-time·B2B
#448594·Dodano wczoraj·1
Źródło: nofluffjobs.com
Aplikuj teraz

Tech Stack / Keywords

LLMOpsMLOps,CloudKubernetesTerraformPythonCI/CDObservabilityDevOps/Solutions Architect,TOGAF

Firma i stanowisko

This role is for an international FinTech & Telecom group building a world-class AI capability. The position involves building, operating, and maintaining AI platforms that support AI agents, models, and applications for next-generation financial services across multiple markets.

Wymagania

  • Minimum 8–10 years in software, platform, cloud, or Site Reliability Engineering (SRE).
  • 2–3 years leading or guiding engineering teams.
  • Hands-on experience building and running enterprise-grade cloud platforms, Infrastructure as Code (IaC), containers, and automated pipelines.
  • Practical experience with LLMOps and MLOps, focusing on model evaluation, monitoring, tracing, and cost management.
  • Strong proficiency with cloud platforms (AWS, Azure, GCP), Kubernetes, Infrastructure as Code tools (Terraform, Ansible), and secrets management.
  • Experience with on-call shifts, incident response, and root-cause analysis in high-availability environments.
  • Exposure to regulated sectors requiring formal change management, audit, and security.

Nice to have:

  • Recognized cloud, platform, or SRE certifications (e.g., AWS, Azure DevOps, Solutions Architect, CKA, TOGAF).
  • Postgraduate qualification (Master's degree) in a technical field.
  • Experience in FinTech, banking, or other high-volume transactional platforms.

Obowiązki

AI Platform & Release Mechanics:

  • Build and maintain multi-environment AI platform infrastructure, CI/CD pipelines, container orchestration, and release mechanics.

Model Operations & LLMOps:

  • Manage production model lifecycle, including provisioning, dynamic routing, versioning, fallback strategies, latency optimization, and cost control.

Evaluation & Quality Gates:

  • Build and own evaluation harnesses, regression test suites, and pre-release quality gates to ensure model and agent safety and performance.

Observability & Incident Leadership:

  • Design the full observability stack including monitoring, tracing, alerting.
  • Hold production accountability including on-call shifts, incident response leadership, and root-cause analysis.

Application Interfaces & APIs:

  • Deliver full-stack applications, developer interfaces, and robust APIs for AI service consumption.
Emerge Soft Sp. z o.o.

Emerge Soft Sp. z o.o.

6 aktywnych ofert

Zobacz wszystkie oferty
Aplikuj teraz