
Software Engineer (Distributed Data Systems)
Exa
Job description
-
As a Data Engineer, you’ll architect and build the data infrastructure that powers everything we do—from crawling billions of pages to training our embedding models to serving real-time search
-
You’ll have enormous autonomy in designing systems that scale to hundreds of petabytes
-
Design a lakehouse architecture that handles 100+ PB of web crawl data
-
Build streaming pipelines that process billions of documents per day for real-time indexing
-
Architect the data layer for our embedding training infrastructure on Ray
-
Scale our ClickHouse deployment to handle analytical queries across petabytes of search logs- If you’ve ever wanted to build data pipelines at a scale that most companies only dream about, this is your chance
-
An obsessive focus on reliability and building systems that don’t page you at 3am
-
Familiarity with Ray, Spark, or ClickHouse at production scale
-
Deep understanding of lakehouse architectures (Delta Lake, Iceberg, Hudi) and when to use them
-
Hands-on experience with streaming data systems (Kafka, Flink, or similar)
-
Experience building and operating large-scale distributed data processing pipelines
-
Experience with Lance or other vector-native storage formats
-
Background in GPU-accelerated data processing (RAPIDS, cuDF)