YemPover Inc logo

Sr. Data Engineer

YemPover Inc

On-siteseniorPosted 22h ago

Job description

  • Experience: 8+ years in data engineering, with a strong focus on large-scale, distributed, cloud-native systems.

  • Snowflake depth (primary): Proven, hands-on production experience with Snowflake โ€” warehouse tuning and cost optimization, RBAC and data governance, Streams/Tasks, Snowpipe, Time Travel, and Snowpark. You can reason about query profiles and micro-partitioning, not just write SQL.

  • Core languages: Expert Python and strong, advanced SQL; PySpark/Snowpark a plus.

  • Transformation & orchestration: Production experience with dbt and a workflow orchestrator (Airflow strongly preferred), designing for portability across engines.

  • AWS ecosystem: Solid hands-on experience with S3, Glue, Step Functions, IAM, and event services (SNS/SQS).

  • Streaming: Track record building streaming/CDC pipelines with Kinesis or Kafka and tools like Debezium, Fivetran, or Airbyte.

  • Data quality: Demonstrated ownership of automated testing, data profiling, reconciliation, and pipeline validation โ€” you own the QA of your own pipelines.

  • Ways of working: Strong documentation habits (playbooks, technical specs), a bias toward reproducibility and version control, and a clear ownership mindset.

  • Top desirable โ€” Automotive domain: Direct experience with automotive repair-order data, dealership fixed-operations, or DMS feeds is a top differentiator for this role.

  • Top desirable โ€” Databricks: Hands-on Databricks experience (Spark, Delta Lake, Unity Catalog, Workflows) used alongside Snowflake. We value portability across engines and may extend the stack over time, so multi-platform depth is highly valued.

  • Certifications: SnowPro Core or Advanced (Data Engineer), Databricks Certified Data Engineer Professional, or AWS Certified Data Engineer.

  • Governance & lineage: Exposure to cataloging/lineage tooling (OpenMetadata) and modern CI/CD for data (tested, peer-reviewed deployments).

Key Responsibilities

  1. Snowflake Warehouse Engineering
  • Design and maintain governed, well-modeled schemas in Snowflake using dimensional and Medallion (bronze/silver/gold) patterns, applying dbt for modular, version-controlled transformations.

  • Optimize Snowflake for performance and cost โ€” right-size virtual warehouses, manage clustering keys and micro-partition pruning, tune query profiles, and control credit consumption through resource monitors and warehouse scaling policies.

  • Leverage native Snowflake features such as Streams and Tasks (or external orchestration), Snowpipe / Snowpipe Streaming for continuous ingestion, Dynamic Tables, Time Travel, Zero-Copy Cloning, and Secure Data Sharing.

  • Implement governance with role-based access control (RBAC), row-access and masking policies, and object tagging for PII and sensitive automotive/customer data.

  1. Pipeline Development & AWS Data Lake Engineering
  • Build and orchestrate complex ETL/ELT pipelines through a central control plane (e.g. Airflow with astronomer-cosmos) that runs consistently across AWS, Snowflake, and Databricks, favoring portable, engine-agnostic designs over transformations locked into any single platform's native constructs.

  • Manage the data lake on AWS S3 as the primary landing zone, optimizing open storage formats (Iceberg, Parquet, Delta) and integrating them with Snowflake external tables and Iceberg tables.

  • Build decoupled, event-driven architectures using SNS and SQS for high-throughput messaging between data services.

  1. Real-Time Data Streaming & Ingestion
  • Develop low-latency ingestion using AWS Kinesis or Kafka, feeding operational analytics through Snowpipe Streaming and Kafka connectors.

  • Implement Change Data Capture (CDC) with Debezium, Fivetran, or Airbyte to keep Snowflake in near-real-time sync with source systems windowed on update timestamps.

  1. Data Quality, Reconciliation & QA Ownership
  • Own end-to-end data quality by embedding automated tests directly into ETL/ELT flows โ€” using portable frameworks such as dbt tests and Great Expectations that apply uniformly whether transformations run on AWS, Snowflake, or Databricks โ€” rather than treating QA as a separate downstream step.

  • Build cross-engine reconciliation between source systems and Snowflake โ€” control totals, partition/row hashing, and drill-down on discrepancies โ€” to guarantee completeness and correctness at scale.

  • Enforce data contracts and schema-evolution guidelines, and stand up proactive observability and alerting to catch data drift and pipeline anomalies before they reach users.

  1. Engineering for ML/AI
  • Engineer ML-ready datasets and feature pipelines to support the Data Science team, managing feature freshness and lineage.

  • Operationalize AI workflows using Snowflake Cortex, Snowpark for Python, and AWS Bedrock where appropriate.

  1. Technical Leadership & Collaboration
  • Mentor engineers on coding standards, SQL and Snowflake optimization, dbt modeling, and Python development, and lead code reviews.

  • Partner with architects, product, and analytics teams to translate designs into reliable, documented, production code, and help drive SDLC and CI/CD maturity across the data platform.