Senior AI Data Engineer
Tech Stack / Keywords
Firma i stanowisko
Billennium is a global technology company providing IT services and digital solutions with over 1,700 professionals in offices across Poland, Canada, Malaysia, Germany, Switzerland, and India. The company focuses on harnessing technology and innovation with core values summarized by TIGER: Trust, Innovation, Growth, Energy, and Responsibility.
Wymagania
- 5+ years of professional experience in Data Engineering, Applied Data Science, Analytics Engineering, or related fields.
- Proven experience owning and maintaining production-grade data pipelines.
- Strong skills in Python for data processing, parsing, cleaning, normalization, and pipeline development.
- Hands-on experience with unstructured enterprise data such as documents, PDFs, Office files, wikis, or knowledge bases.
- Proven experience preparing and managing data for retrieval, NLP, or RAG-based use cases including embedding preparation and metadata enrichment.
- Strong understanding of data quality engineering including validation, monitoring, lineage, freshness, and refresh cycles.
- Understanding of RAG and retrieval concepts such as chunking, metadata, embeddings, and optimization.
- Ability to collaborate effectively with AI Engineers, Architects, and business stakeholders on data requirements.
- Strong analytical and problem-solving skills with hands-on approach to data quality issues.
Nice to have:
- Experience with Postgres + pgvector or other vector databases/store.
- Knowledge of hybrid search, filtered retrieval, and reranking techniques.
- Experience with AI observability and monitoring tools such as Langfuse.
- Familiarity with RAG evaluation frameworks and metrics like RAGAS or DeepEval.
- Experience with PII detection, masking, and enterprise privacy workflows like Microsoft Presidio.
- Experience with enterprise data governance, lineage, and auditability frameworks.
- Familiarity with LLM gateway patterns and AI application architectures.
- Experience collaborating on production AI agents, copilots, or RAG systems.
Obowiązki
- Lead data discovery and triage for AI use cases, identifying authoritative and relevant data sources.
- Analyze enterprise repositories to identify duplicates, outdated, or low-quality content.
- Clean, normalize, deduplicate, and standardize large volumes of structured and unstructured data.
- Define data inclusion and exclusion criteria based on relevance, quality, freshness, and sensitivity.
- Build and maintain ingestion pipelines for SharePoint and enterprise document repositories.
- Process documents such as PDFs, Word/Excel files, and wikis; implement text extraction, document parsing, and metadata capture.
- Define and maintain document normalization standards including taxonomy and canonical identifiers.
- Prepare enterprise data specifically for retrieval and RAG-based AI applications.
- Design and optimize chunking strategies, metadata enrichment, and document structures.
- Prepare and maintain AI-ready knowledge sets for embedding and serving through Postgres + pgvector.
- Support retrieval optimization including filtered retrieval, hybrid search, and reranking.
- Collaborate on retrieval interfaces and integration of knowledge assets into AI applications.
- Define and implement data quality gates covering freshness, completeness, relevance, duplication, and metadata.
- Establish monitoring and refresh processes for reliable AI knowledge.
- Work with AI Engineers to evaluate retrieval and RAG performance using frameworks like RAGAS and DeepEval.
- Establish human feedback loops and targeted data audits.
- Maintain traceability and lineage from original source to production.
- Apply data governance, privacy, and security requirements; implement PII detection and masking using tools like Microsoft Presidio.
- Create reusable approaches for data cleanup and RAG readiness including ingestion templates and chunking playbooks.
- Build repeatable data foundations to accelerate future AI use cases.
Benefity
- Comprehensive benefits including Udemy for Business, private medical care, Multisport card, veterinary package, language lessons, and shopping vouchers.
- Career growth opportunities and learning perks with partnerships including Microsoft, AWS, Snowflake, and Salesforce.
- Global collaboration with a diverse international team.
- Innovative and forward-thinking work environment.
- Team-building events including an annual event in Mazury.
- Welcome pack to start the journey with the company.
Inne informacje
Informujemy, że administratorem danych jest Billennium S.A. z siedzibą w Warszawie, ul. Koszykowa 61 (dalej jako "administrator"). Masz prawo do żądania dostępu do swoich danych osobowych, ich sprostowania, usunięcia lub ograniczenia przetwarzania, prawo do wniesienia sprzeciwu wobec przetwarzania, a także prawo do przenoszenia danych oraz wniesienia skargi do organu nadzorczego. Dane osobowe przetwarzane będą w celu realizacji procesu rekrutacji. Podanie danych w zakresie wynikającym z ustawy z dnia 26 czerwca 1974 r. Kodeks pracy jest obowiązkowe. W pozostałym zakresie podanie danych jest dobrowolne. Odmowa podania danych obowiązkowych może skutkować brakiem możliwości przeprowadzenia procesu rekrutacji. Administrator przetwarza dane obowiązkowe na podstawie ciążącego na nim obowiązku prawnego, zaś w zakresie danych dodatkowych podstawą przetwarzania jest zgoda. Dane osobowe będą przetwarzane do czasu zakończenia postępowania rekrutacyjnego i przez okres możliwości dochodzenia ewentualnych roszczeń, a w przypadku wyrażenia zgody na udział w przyszłych postępowaniach rekrutacyjnych - do czasu wycofania tej zgody. Zgoda na przetwarzanie danych osobowych może zostać wycofana w dowolnym momencie. Odbiorcą danych jest serwis Just Join IT oraz inne podmioty, którym powierzyliśmy przetwarzanie danych w związku z rekrutacją.
Billennium S.A.
74 aktywne oferty