Senior Data Scientist – AI/LLM Evaluation
140 - 170 PLN/ godz.B2B
SeniorFull-time·B2B
#423997·Dodano 5 dni temu·0
Źródło: nofluffjobs.comTech Stack / Keywords
Data scienceStatistical analysisExperimantation skillsAILLMAgentic workflow3DGeometryRenderingLLM-as-a-JudgeLLM-based evaluation
Firma i stanowisko
The role is with Square One Resources on an international project focused on Artificial Intelligence, Large Language Models (LLMs), and AI evaluation, specifically developing and improving methodologies for evaluating AI models and agentic workflows.
Wymagania
- Proven experience as a Senior Data Scientist or in a similar role.
- Strong data science, statistical analysis, and experimentation skills.
- Hands-on experience with AI/LLM evaluation.
- Strong understanding of LLMs and agentic workflows.
- Experience designing, validating, and optimizing evaluation metrics for AI/ML models.
- Strong analytical skills and experience investigating model errors, including false positives and false negatives.
- Experience working with ground-truth datasets and defining quality criteria.
- Ability to independently analyze experimental results and translate findings into actionable recommendations.
- Strong problem-solving and analytical mindset.
Nice to have:
- Basic knowledge of 3D, geometry, or rendering.
- Experience with LLM-as-a-Judge / LLM-based evaluation systems.
- Experience with human-in-the-loop evaluation methodologies.
- Previous experience evaluating AI agents or agentic systems.
- Familiarity with advanced approaches to evaluating LLM-generated content and outputs.
Obowiązki
- Design and validate evaluation metrics and ground-truth datasets.
- Analyze evaluation quality, including accuracy, false positives, and false negatives.
- Optimize metric normalization and scoring methodologies.
- Develop and improve LLM-based judges for automated evaluation.
- Design and implement human-feedback-based evaluation approaches.
- Explore and develop new evaluation metrics, including complexity, prompt adherence, output quality, and correctness.
- Analyze experiment results and identify opportunities to improve model performance and evaluation methodologies.
- Collaborate with Data Science, AI/ML, and Engineering teams to develop evaluation solutions for AI models and agentic workflows.
Benefity
- Private healthcare
- Sport subscription
Opieka zdrowotna
Karta sportowa
SQUARE ONE RESOURCES
168 aktywnych ofert