GoGuardian logo

Data Engineer

GoGuardian

On-sitemid$130k–$150kPosted 12h ago

Job description

  • We’re looking for a Data Engineer II to help design, build, and continuously improve the GoGuardian Analytics and AI/ML ecosystem. This position sits on the Data Engineering team, a group responsible for building and maintaining the core data platform that powers analytics, product insights, and machine learning across the company. You’ll collaborate closely with Data Science, Business Intelligence, and other teams to enable the next generation of data-driven products and AI capabilities

  • Design, build, and optimize ETL pipelines that power analytics, data science, and ML workflows using tools such as Databricks, PySpark, and Airflow

  • Develop and maintain labeling and retraining pipelines for machine learning models, ensuring quality, reproducibility, and observability

  • Implement and support MLOps practices, including model versioning, CI/CD for ML, and model monitoring in production environments

  • Collaborate with data scientists to productionize and scale model training, inference, and evaluation pipelines

  • Contribute to the design and evolution of the data lakehouse, including schema design, partitioning strategies, and performance optimization

  • Document and communicate data architecture, lineage, and dependencies to ensure transparency and maintainability across teams

  • Champion data quality and governance, ensuring that datasets are accurate, well-structured, and compliant with organizational standards

  • Leverage infrastructure-as-code and containerization to build reproducible, maintainable environments

  • Participate in code reviews and continuous improvement of engineering best practices within the team

Benefits

  • Health insurance: 90% for you (100% HMO) and 50% for your dependents

  • 401(k) retirement savings plan with company match

  • Comprehensive time-off package, including flexible time off, sick time, 13 company holidays, quarterly wellness days, and a paid year-end holiday break

  • 24 hours and 7 days a week, free confidential services include personal coaching, counseling, self-care apps, and much more!

  • Employee stock option

  • Company-supported growth

  • Paid parental leave for up to 14 weeks

  • Company paid short and long-term disability

  • A fertility and adoption program to help you support your family needs

  • Relax, recharge, and refresh with weekly yoga classes and guided meditation- The ideal candidate combines strong software engineering and data architecture skills with curiosity about machine learning systems and a drive to automate, optimize, and scale data workflows

  • Bachelor’s degree in Computer Science, Engineering, or related field

  • 2–4 years of experience building and operating large-scale data systems, ideally supporting analytics and ML workloads

  • Excellent problem-solving, collaboration, and communication skills; comfortable working in a dynamic, fast-paced environment

  • Hands-on experience with workflow orchestration tools such as Airflow, Dagster, or Prefect

  • Experience with DBT

  • Familiarity with MLOps concepts (e.g., feature stores, model registries, CI/CD for ML)

  • Strong understanding of data modeling, ETL design, and distributed data systems

  • Experience with modern data warehousing and lakehouse platforms, preferably Databricks

  • Experience using Infrastructure as Code, preferably Terraform

  • Proficiency in Python and SQL, with experience in PySpark, pandas, or similar data processing frameworks

  • Experience with AWS data and compute services (S3, Lambda, ECS, CloudWatch, etc.) or equivalent cloud platforms