Exa logo

Software Engineer (Distributed Data Systems)

Exa

On-siteSan Francisco, CAmid$150k–$300kPosted 8h ago

Job description

  • As a Data Engineer, you’ll architect and build the data infrastructure that powers everything we do—from crawling billions of pages to training our embedding models to serving real-time search

  • You’ll have enormous autonomy in designing systems that scale to hundreds of petabytes

  • Design a lakehouse architecture that handles 100+ PB of web crawl data

  • Build streaming pipelines that process billions of documents per day for real-time indexing

  • Architect the data layer for our embedding training infrastructure on Ray

  • Scale our ClickHouse deployment to handle analytical queries across petabytes of search logs- If you’ve ever wanted to build data pipelines at a scale that most companies only dream about, this is your chance

  • An obsessive focus on reliability and building systems that don’t page you at 3am

  • Familiarity with Ray, Spark, or ClickHouse at production scale

  • Deep understanding of lakehouse architectures (Delta Lake, Iceberg, Hudi) and when to use them

  • Hands-on experience with streaming data systems (Kafka, Flink, or similar)

  • Experience building and operating large-scale distributed data processing pipelines

  • Experience with Lance or other vector-native storage formats

  • Background in GPU-accelerated data processing (RAPIDS, cuDF)