
Sr. Data Engineer
YemPover Inc
Job description
-
Experience: 8+ years in data engineering, with a strong focus on large-scale, distributed, cloud-native systems.
-
Snowflake depth (primary): Proven, hands-on production experience with Snowflake โ warehouse tuning and cost optimization, RBAC and data governance, Streams/Tasks, Snowpipe, Time Travel, and Snowpark. You can reason about query profiles and micro-partitioning, not just write SQL.
-
Core languages: Expert Python and strong, advanced SQL; PySpark/Snowpark a plus.
-
Transformation & orchestration: Production experience with dbt and a workflow orchestrator (Airflow strongly preferred), designing for portability across engines.
-
AWS ecosystem: Solid hands-on experience with S3, Glue, Step Functions, IAM, and event services (SNS/SQS).
-
Streaming: Track record building streaming/CDC pipelines with Kinesis or Kafka and tools like Debezium, Fivetran, or Airbyte.
-
Data quality: Demonstrated ownership of automated testing, data profiling, reconciliation, and pipeline validation โ you own the QA of your own pipelines.
-
Ways of working: Strong documentation habits (playbooks, technical specs), a bias toward reproducibility and version control, and a clear ownership mindset.
-
Top desirable โ Automotive domain: Direct experience with automotive repair-order data, dealership fixed-operations, or DMS feeds is a top differentiator for this role.
-
Top desirable โ Databricks: Hands-on Databricks experience (Spark, Delta Lake, Unity Catalog, Workflows) used alongside Snowflake. We value portability across engines and may extend the stack over time, so multi-platform depth is highly valued.
-
Certifications: SnowPro Core or Advanced (Data Engineer), Databricks Certified Data Engineer Professional, or AWS Certified Data Engineer.
-
Governance & lineage: Exposure to cataloging/lineage tooling (OpenMetadata) and modern CI/CD for data (tested, peer-reviewed deployments).
Key Responsibilities
- Snowflake Warehouse Engineering
-
Design and maintain governed, well-modeled schemas in Snowflake using dimensional and Medallion (bronze/silver/gold) patterns, applying dbt for modular, version-controlled transformations.
-
Optimize Snowflake for performance and cost โ right-size virtual warehouses, manage clustering keys and micro-partition pruning, tune query profiles, and control credit consumption through resource monitors and warehouse scaling policies.
-
Leverage native Snowflake features such as Streams and Tasks (or external orchestration), Snowpipe / Snowpipe Streaming for continuous ingestion, Dynamic Tables, Time Travel, Zero-Copy Cloning, and Secure Data Sharing.
-
Implement governance with role-based access control (RBAC), row-access and masking policies, and object tagging for PII and sensitive automotive/customer data.
- Pipeline Development & AWS Data Lake Engineering
-
Build and orchestrate complex ETL/ELT pipelines through a central control plane (e.g. Airflow with astronomer-cosmos) that runs consistently across AWS, Snowflake, and Databricks, favoring portable, engine-agnostic designs over transformations locked into any single platform's native constructs.
-
Manage the data lake on AWS S3 as the primary landing zone, optimizing open storage formats (Iceberg, Parquet, Delta) and integrating them with Snowflake external tables and Iceberg tables.
-
Build decoupled, event-driven architectures using SNS and SQS for high-throughput messaging between data services.
- Real-Time Data Streaming & Ingestion
-
Develop low-latency ingestion using AWS Kinesis or Kafka, feeding operational analytics through Snowpipe Streaming and Kafka connectors.
-
Implement Change Data Capture (CDC) with Debezium, Fivetran, or Airbyte to keep Snowflake in near-real-time sync with source systems windowed on update timestamps.
- Data Quality, Reconciliation & QA Ownership
-
Own end-to-end data quality by embedding automated tests directly into ETL/ELT flows โ using portable frameworks such as dbt tests and Great Expectations that apply uniformly whether transformations run on AWS, Snowflake, or Databricks โ rather than treating QA as a separate downstream step.
-
Build cross-engine reconciliation between source systems and Snowflake โ control totals, partition/row hashing, and drill-down on discrepancies โ to guarantee completeness and correctness at scale.
-
Enforce data contracts and schema-evolution guidelines, and stand up proactive observability and alerting to catch data drift and pipeline anomalies before they reach users.
- Engineering for ML/AI
-
Engineer ML-ready datasets and feature pipelines to support the Data Science team, managing feature freshness and lineage.
-
Operationalize AI workflows using Snowflake Cortex, Snowpark for Python, and AWS Bedrock where appropriate.
- Technical Leadership & Collaboration
-
Mentor engineers on coding standards, SQL and Snowflake optimization, dbt modeling, and Python development, and lead code reviews.
-
Partner with architects, product, and analytics teams to translate designs into reliable, documented, production code, and help drive SDLC and CI/CD maturity across the data platform.