Senior Data Engineer - REMOTE
Brak informacji o wynagrodzeniu
SeniorFull-time
#432893·Dodano 2 dni temu·1
Źródło: nofluffjobs.comTech Stack / Keywords
PythonETL
Firma i stanowisko
Switch Intelligence Inc is offering a position to build and improve the data backbone of the Switch Intelligence Platform, a multi-tenant, source-agnostic data streaming and intelligence system.
Wymagania
- 5+ years of professional Data Engineering experience.
- Strong Python development skills, including OOP, clean architecture, testing, and writing maintainable production code.
- Strong experience designing and implementing ETL/ELT data pipelines.
- Event streaming/messaging experience with Kafka, Google Pub/Sub, RabbitMQ, or similar.
- Hands-on experience integrating with REST APIs, third-party services, databases, and external data sources.
- Experience building multi-tenant data platforms or SaaS integrations with tenant isolation and per-tenant customization.
- Experience with PostgreSQL and strong SQL skills, including JSONB for semi-structured fields.
- Experience ingesting data from Salesforce, HubSpot, Dynamics, or comparable commercial source systems.
- Experience with modern warehouses or lakehouses such as Snowflake, Databricks, or BigQuery.
- Experience with Change Data Capture like Debezium and reliable write-back patterns.
- Experience with data processing frameworks and/or distributed data processing.
- Experience with Docker and Kubernetes.
- Strong understanding of Git and software development workflows.
- Experience with CI/CD pipelines such as GitHub Actions, Jenkins, Bitbucket Pipelines; including automated testing, deployment, and monitoring.
- Experience building and maintaining production-grade data applications.
- Strong understanding of data modeling, data quality, validation, schema evolution, and pipeline reliability.
- Ability to troubleshoot and resolve complex technical problems independently.
Nice to have:
- Experience with Apache Airflow, Dagster, Prefect, or similar orchestration frameworks.
- Experience with Redis or similar caching/NoSQL technologies.
- Experience with Neo4j or production graph databases.
- Experience with metering, usage aggregation, or systems with strict correctness guarantees.
- Experience with schema evolution/migration tooling or code generation from declarative schemas (e.g., YAML/DSL-driven codegen).
- Experience with Spark, PySpark, or other distributed data processing technologies.
- Experience with Google Cloud Platform (GCP) or other cloud providers (AWS, Azure).
- Experience with Google Cloud Storage, S3, or similar object storage systems.
- Experience with data provenance, lineage, or field-level assertion/audit models.
- Experience with Ansible, SonarQube, Artifactory, or similar DevOps tooling.
- Experience with data observability and monitoring platforms.
Obowiązki
- Design, build, and maintain scalable, reliable, production-grade data pipelines for batch and real-time streaming processing.
- Develop integrations with multiple data sources, APIs, databases, SaaS platforms, and third-party services.
- Build reusable, tenant-aware ingestion and transformation frameworks that support per-tenant customization.
- Implement Change Data Capture (CDC) and reliable write-back flows with idempotency and exactly-once or at-least-once semantics.
- Develop robust ETL/ELT pipelines with validation, cleansing, normalization, transformation, and enrichment workflows.
- Build API-driven data services and integrations using Python.
- Contribute to schema evolution tooling including declarative schema definitions, validation, diffing, and code generation.
- Containerize and deploy data services using Docker and Kubernetes (GKE).
- Design systems handling large volumes of structured and unstructured data in pipelines.
- Implement reliable error handling, retries, logging, monitoring, and data-quality checks.
- Work with relational databases and optimize data storage and retrieval.
- Build asynchronous and event-driven data processing workflows.
- Build automated testing, CI/CD, deployment, and monitoring for data applications.
- Collaborate closely with AI/ML, backend, and engineering teams.
- Research and evaluate new data technologies, APIs, frameworks, and integration patterns.
- Troubleshoot complex production data and integration issues.
- Collaborate on graph translation in the data pipeline to create a governed knowledge graph model.
- Build and operate metering and usage-aggregation workloads with windowed aggregation, late-arriving data handling, and billing-grade correctness.
Switch Intelligence Inc
Pracodawca