Data Scientist - Product Catalog Data

Brak informacji o wynagrodzeniu
MidFull-time
#389808·Dodano 21 dni temu·0
Źródło: nofluffjobs.com
Aplikuj teraz

Tech Stack / Keywords

SparkPySparkMLOpsPythonStatisticsMachine learningAIAirflow

Firma i stanowisko

Our team drives the solutions that build a trustworthy marketplace. As a collaborative group of Data Scientists, Data Engineers, and Analysts, we develop the statistical models, products and business rules necessary to analyze and continuously enhance our Product Catalog. We oversee the quality and correctness of all product components: titles, descriptions, images, and parameters. Our scope also covers ensuring consistency with sellers offers, automated products duplicates detection, building selection models and more. We work in close collaboration with developers and business teams, participating in the entire project lifecycle - from ideation through to implementation. Our ultimate goal is to create a single, reliable source of truth that powers a professional shopping experience for buyers and a fair, frictionless environment for our partners.

Wymagania

  • Hands-on experience with Spark or PySpark for large dataset processing.
  • Practical experience in MLOps and deploying machine learning models into production.
  • Strong proficiency in Python with clean, object-oriented code practices and unit testing.
  • Deep understanding of statistics and machine learning theory and methods.
  • Enthusiasm and practical understanding of Generative AI and Agentic AI, including prompting and validation.
  • Excellent communication skills with the ability to translate business challenges into ML problems and explain technical results.
  • Degree in a quantitative field such as Mathematics, Physics, Computer Science, or Economics, or equivalent practical experience in Data Science.

Nice to have:

  • Experience with Apache Airflow for pipeline orchestration.

Obowiązki

  • Co-creating projects at each stage from concept to productization to solve business problems.
  • Using diverse modeling techniques including gradient boosting, Bayesian methods, causal inference, optimization, deep learning, and Generative AI (LLMs, Agentic AI).
  • Collaborating closely with Data Engineers on data processing on GCP.
  • Integrating solutions with Analysts, Product Managers, and Developers.
  • Leveraging diverse data types including spatial data, natural language (NLP), images, and time series.
  • Implementing both offline and online machine learning models.
  • Participating in brainstorming, knowledge sharing, and continuous professional development.

Benefity

  • Flexible working hours with a hybrid model (4 days office, 1 day remote) plus 30 days of occasional remote work.
  • Well-located modern offices equipped with kitchens, bicycle parking, terraces, ergonomic furniture, and interactive conference rooms.
  • Choice of a 16" or 14" MacBook Pro or Dell Windows laptop with accessories.
  • Access to a cafeteria plan with medical, sports, lunch packages, insurance, and purchase vouchers.
  • Employer-paid English classes relevant to the job.
  • Training budget, inter-team tourism, hackathons, and an internal learning platform with multiple trainings.
  • Additional day off for volunteering.
  • Social events such as Spin Kilometers, Family Day, Fat Thursday, and Advent of Code.
Elastyczne godziny
Karta sportowa
Dofinansowanie szkoleń
Opieka zdrowotna
Firmowa stołówka
Kursy językowe
Spotkania integracyjne
Telefon
Parking dla aut
Parking dla rowerów
Prysznic
Napoje w biurze
Allegro

Allegro

94 aktywne oferty

Zobacz wszystkie oferty
Aplikuj teraz