Evlo AI logo

Data Engineer

Evlo AI

RemotemidPosted 2h ago

Job description

About The Role

The Data Engineer will design and operate the data infrastructure that powers product analytics, operational reporting, and machine learning initiatives. The role will build reliable batch and streaming pipelines across event, transactional, and third-party data sources, with a focus on data quality, observability, and scalable cloud architecture.

Working closely with analytics engineers, data analysts, software engineers, and product stakeholders, the role will turn complex data requirements into well-modeled, discoverable datasets. The team values engineers who can own systems in production while understanding how downstream users interpret and apply the data.

Key Responsibilities

  • Design and maintain batch and streaming data pipelines using Python, SQL, Apache Spark, and orchestration tools such as Airflow or Dagster

  • Build scalable data platforms on AWS, including services such as S3, Glue, EMR, Lambda, and Redshift, with appropriate security and cost controls

  • Model raw and curated data in Snowflake, BigQuery, or Redshift to support reporting, experimentation, and operational use cases

  • Implement data quality checks, lineage, monitoring, and alerting to detect freshness, completeness, schema, and accuracy issues before they affect users

  • Develop reusable ingestion frameworks for application databases, APIs, event streams, and third-party systems using tools such as Kafka, Fivetran, or dbt

  • Partner with analysts and analytics engineers to define data contracts, dimensional models, metrics, and documentation for trusted self-service analytics

  • Review designs and code, improve pipeline performance and reliability, and contribute to engineering standards for testing, deployment, and incident response

What We Are Looking For

  • 3โ€“8 years of experience in data engineering, analytics engineering, or a closely related software engineering role

  • Strong proficiency in Python and SQL, including experience optimizing complex queries and working with large datasets

  • Hands-on experience building production ETL or ELT pipelines with Apache Spark, Airflow, Dagster, dbt, or comparable technologies

  • Experience with cloud data platforms and services, preferably AWS, plus a strong understanding of distributed systems, storage, and compute patterns

  • Practical knowledge of data modeling, dimensional modeling, schema evolution, data governance, and automated data quality testing

  • Bachelorโ€™s degree in computer science, engineering, mathematics, statistics, or a related technical field, or equivalent professional experience

  • Bonus: Experience with Kafka or Kinesis, Terraform, Kubernetes, CI/CD, real-time analytics, and observability platforms such as Datadog or Monte Carlo