Senior Data Engineer
Switch Intelligence Inc
- Praca zdalna
We are looking for a strong Data Engineer who is also an excellent Python developer to think through already existing architecture, improve, modify or redesign, and build the data backbone of the Switch Intelligence Platform — a multi-tenant, source-agnostic data streaming and intelligence system . You will build scalable pipelines that connect and integrate with multiple data sources, APIs, databases, and external services, which will feed a Neo4j knowledge graph , a postgres , lakehouse , and an AI-driven signal loop . The ideal candidate has 5+ years of professional experience , strong software engineering fundamentals, and hands-on experience building production-grade, multi-tenant data systems .
Key Responsibilities
- Design, build, and maintain scalable, reliable, production-grade data pipelines for batch and real-time (streaming) processing.
- Develop integrations with multiple data sources, APIs, databases, SaaS platforms, and third-party services.
- Build reusable, tenant-aware ingestion and transformation frameworks that support per-tenant object and field customization without sacrificing data integrity.
- Implement Change Data Capture (CDC) and reliable write-back flows to and from source systems, with idempotency and exactly-once/at-least-once semantics as appropriate.
- Develop robust ETL/ELT pipelines with validation, cleansing, normalization, transformation, and enrichment workflows.
- Build API-driven data services and integrations using Python .
- Contribute to schema evolution tooling: declarative (YAML-based) schema definitions, validation, diffing, and code generation.
- Containerize and deploy data services using Docker and Kubernetes (GKE) .
- Design systems for handling large volumes of structured and unstructured data in the pipelines .
- Implement reliable error handling , retries , logging , monitoring , and data-quality checks .
- Work with relational databases and optimize data storage and retrieval.
- Build asynchronous and event-driven data processing workflows.
- Build automated testing, CI/CD, deployment, and monitoring for data applications.
- Collaborate closely with AI/ML, Backend, and Engineering teams .
- Research and evaluate new data technologies, APIs, frameworks, and integration patterns.
- Troubleshoot complex production data and integration issues.
- Collaborate on graph translation in the data pipeline: ingest data from source systems and translate it into a governed knowledge graph model.
- Build and operate metering and usage-aggregation workloads with windowed aggregation, late-arriving data handling, and billing-grade correctness.
Ideal Candidate
The ideal candidate combines strong Data Engineering expertise with excellent software engineering and Python skills . You should be comfortable taking an ambiguous data integration requirement through to the pipeline architecture, writing the Python code, connecting the external systems, building the pipeline, testing it, and deploying it.
You should be able to reason about real architectural trade-offs — for example, when to use typed columns versus JSONB versus extension objects for per-tenant flexibility, and why a metering engine or money chain cannot be built on a fully generic EAV model. We value honest technical pushback grounded in correctness, integrity, and operability.
You should be a hands-on engineer and strong coder , rather than someone focused primarily on data analysis or low-code ETL tools.
Nice to Have
- Experience with Apache Airflow, Dagster, Prefect , or similar orchestration frameworks.
- Experience with Redis or similar caching/NoSQL technologies.
- Experience with Neo4j or another graph database in production.
- Experience with metering, usage aggregation, or other systems with strict correctness guarantees.
- Experience with schema evolution/migration tooling or code generation from declarative schema definitions (e.g., YAML/DSL-driven codegen).
- Experience with Spark, PySpark , or other distributed data processing technologies.
- Experience with Google Cloud Platform (GCP) — our target platform — or AWS/Azure.
- Experience with Google Cloud Storage, S3 , or similar object storage systems.
- Experience with data provenance, lineage, or field-level assertion/audit models.
- Experience with Ansible, SonarQube, Artifactory , or similar DevOps tooling.
- Experience with data observability and monitoring platforms.