Snorkel AI logo

Software Engineer (Platform)

Snorkel AI

On-siteNew York City, NYsenior$220k–$300kPosted 12h ago

Job description

  • We’re looking for Platform Engineers who combine strong infrastructure chops with real backend engineering depth

  • You’ll build and operate the systems, services, and agent infrastructure that let product teams move fast and reliably — from data access layers and event-driven pipelines to the agent first architecture that will transform our ability to scale. Your work will directly allow us to scale the amount of high quality data we are able to deliver to our customers

  • We are looking to grow our team of Platform Engineers, and are hiring at multiple levels

  • Design and build agent infrastructure that allows us to safely and reliably speed up workflows from engineering and operation teams

  • Design and implement event-driven data flows using event brokers, CDC connectors, schema registries, event routing, and dead letter queues — ensuring events flow reliably and failures are visible and recoverable

  • Build the systems that track how data moves through the platform (lineage), enforce who can access what (governance and RBAC), and log what happened (auditing), including PII handling, retention policy enforcement, and audit infrastructure for enterprise and regulatory compliance

  • Set strategy and architecture for build systems, testing frameworks, and CI/CD pipelines, and drive the transition toward robust, automated continuous deployment

  • Instrument services with OpenTelemetry, define and monitor SLOs (query latency, pipeline success rates, service reliability), and build alerting that catches issues before they become incidents — you will be on-call for the systems you build

  • Contribute to infrastructure cost visibility and optimization — query cost estimation, workload right-sizing, and routing data to the most cost-effective storage tier for its access pattern

  • Collaborate with engineers, product managers, and designers to bring consistency and high standards to codebases, infrastructure, and processes

  • You’ll have meaningful ownership over the infrastructure and services that every product team and customer deployment depends on. This isn’t a maintenance role — you’ll be making foundational architecture decisions that shape how the company scales, with a team that cares deeply about reliability, craft, and moving fast without breaking things

Benefits

  • Health: Snorkelers and their dependents are covered by comprehensive medical, dental, and vision plans.

  • Wellness: All Snorkelers get a yearly wellness stipend to use on anything related to health and well-being.

  • PTO and rest days: We provide unlimited personal time off so you can find your work-life balance. Plus, two company-wide rest days per quarter in addition to company holidays.

  • One global team: From our headquarters in Redwood City, CA, to regional offices in San Francisco and New York, we work together as one global team.

  • Events: Get to know your fellow Snorkelers outside of work! We host regular events and learning opportunities for our teams across our offices and virtually to connect, have fun, and grow.

  • Family leave: Snorkel offers generous parental leave for both birthing and non-birthing parents based on eligibility.- Strong background in distributed systems and cloud platforms (AWS preferred) — hands-on experience with services like S3, RDS, EKS, EventBridge, and IAM, and comfort working in a Terraform-managed environment

  • Familiarity with data orchestration tools (Prefect, Airflow, or Dagster) and transformation frameworks (dbt)

  • Track record of leading complex engineering initiatives, influencing stakeholders, and delivering measurable impact

  • 5+ years building platform infrastructure, backend services, or data systems in production — you have built and operated pipelines, data access layers, distributed services, or ETL/ELT systems at scale

  • Ability to work in a fast-paced environment with strong technical communication skills

  • Fluency with modern developer tooling and a willingness to adopt new tools quickly — the team evaluates and integrates new tooling regularly to improve velocity and reliability

  • Strong proficiency in Python, and experience designing REST APIs for internal services and developers

  • Understanding of data governance concepts — RBAC, PII handling, audit logging, data lineage

  • Experience building shared libraries or SDKs consumed by multiple teams — versioning, backwards compatibility, migration support

  • Experience with event-driven architectures — CDC, event buses, schema registries, at-least-once delivery semantics

  • Experience with OpenTelemetry, ClickHouse, or similar observability infrastructure

  • Experience with Ray or similar frameworks for distributed compute workloads

  • Prior work in regulated environments (SOC 2, FedRAMP, HIPAA) where compliance requirements shaped system design

  • Experience in hyper-growth startup environments or scaling engineering orgs

  • Prior experience as a Tech Lead, Team Lead, or hands-on Engineering Manager