Elastic logo

Principal Software Engineer (Performance Tuning - Elasticsearch)

Elastic

Remotelead$160k–$304kPosted 9h ago

Job description

  • Elasticsearch powers search, observability, and AI retrieval (RAG) for the world’s largest organizations. We are seeking a Principal Software Engineer to join the Elasticsearch Performance team. In this role, you will lead architectural and code-level performance engineering initiatives

  • Your goal is to drive the continuous optimization and predictability of Elasticsearch performance, partnering with area-specific teams to unlock the full potential of our software

  • Owning core performance engineering initiatives from architecture to production, focusing on the delivery of high-impact optimizations. Leading the technical design, plan, and execution for major architectural and code-level performance improvements

  • Developing foundational performance models and methodologies for complex, distributed systems

  • Driving optimization strategies to ensure Elasticsearch remains performant, predictable, and scalable in diverse environments

  • Profiling and analyzing system behavior to identify bottlenecks in logging, metrics, vector search, and ES|QL

  • Ensuring robust performance benchmarks and regression detection for both stateful and stateless (Serverless) architectures

  • Collaborating across the company to embed performance-first thinking into new features from the outset

  • Drive automation efforts by designing and building AI-assisted optimization harnesses that streamline profiling, hypothesis testing, and benchmarking

  • Mentoring and coaching other engineers, fostering a culture of technical excellence and performance-aware development

Benefits

  • Toast to your health: Fully paid health coverage for you and your family, in many locations.

  • Craft your calendar: Flexible location and schedule for most roles.

  • Create space for you: Distributed by design workforce, plus generous number of vacation days each year.

  • Embrace parenthood: Minimum of 16 weeks of parental leave, plus generous family formation benefits.

  • Give back your time: 40 hours each year to use toward volunteering with organizations and causes you’re passionate about.

  • Amplify your impact: Double your charitable giving — we match donations up to $1500 USD (or local currency equivalent).- You possess the ability to collaborate across functions and teams, acting as a force multiplier for performance engineering across the organization

  • You have deep knowledge of Java internals and JVM memory management. You understand how concurrency models work. You can write code that is high-performance, thread-safe, and lock-free. This experience includes working with large open-source and enterprise codebases

  • You have proven experience in profiling and optimizing distributed systems. This includes deep experience with benchmarking tools (e.g., flamegraphs, JMH, Rally), identifying performance regressions, and implementing algorithmic or hardware-aware optimizations

  • You have a solid comprehension of distributed systems architecture, including partition tolerance, cluster state propagation, and scaling challenges in large-scale data stores

  • You possess the ability to collaborate across functions and teams and seamlessly transition between different projects, codebases, or teams based on business priorities

  • You can work autonomously, drive decisions and result in a distributed team by leveraging asynchronous, direct, and transparent communication

  • You have a proven track record of using AI or advanced tooling to accelerate optimization, debug complex performance issues, and automate benchmarking workflows

  • Experience integrating high-performance native libraries (e.g., C++, Rust, SIMD-accelerated code) into Java applications

  • Deep knowledge of modern storage engine performance, index modes, or vector search optimizations

  • Experience working on the internals of a large-scale data store or search engine

  • Experience defining and managing Performance SLAs and success criteria for distributed systems

  • Experience working on the internals of a data store or search engine