
Senior Generative AI Engineer
Cleary Gottlieb Steen & Hamilton LLP
Job description
Cleary Gottlieb is a pioneer in globalizing the legal profession. We have 14 offices in major financial centers around the world, but we operate as a single, integrated global partnership and not as a U.S. firm with a network of overseas locations. The firm employs approximately 1,100 lawyers from more than 50 countries.
Since 1946 our lawyers and staff have worked across practices, industries, jurisdictions, and continents to provide clients with simple, actionable approaches to their most complex legal and business challenges, whether domestic or international. We support every client relationship with intellectual agility, commercial acumen, and a human touch.
We're an internal team at Cleary Gottlieb building bespoke AI solutions for legal work, drawing on software engineering, data science, and deep domain expertise. We work at the boundary of enterprise-grade software, and we are regularly building services and applications nobody here has built before. That makes for a dynamic learning environment โ and one where a junior engineer will gain exposure to real-world AI systems from day one.
We are a tight-knit, remote-first team that values growth, honesty, and curiosity. We work collaboratively across multiple disciplines โ engineering, data science, legal domain experts, and product โ to deliver meaningful impact. Junior engineers are paired with experienced mentors and participate in structured code reviews, pair-programming sessions, and weekly knowledge-sharing stand-ups.
You will join as a Senior Generative AI Engineer on the AI Acceleration Team, reporting to the Data Science Manager, and working day to day with our engineers, data scientists, and legal domain experts. Much of the infrastructure and tooling described below is already under way; what we need is someone to own and advance the AI systems that sit on top of it.
Nobody arrives knowing legal workflows or our stack, and you will have the Data Science Manager and the rest of the team alongside you while you pick them up. Judgement, curiosity, and an appetite for more responsibility as we grow count for more here than the length of a CV.
Responsibilities
The Role:
Weโre looking for a Senior Generative AI Engineer who owns the full lifecycle of LLM-powered products, from rapid prototyping through production deployment. You'll design, ship, and operate AI systems that extract intelligence from complex legal documents, orchestrate multi-step agent workflows, and deliver measurable accuracy improvements under real-world latency and cost constraints.
You should bring a track record of shipping LLM-powered features to production users, not just notebooks or demos. We'll ask you to walk us through a system you built end-to-end: the architecture decisions, the failure modes you hit in production, and the evaluation methodology you used to prove it worked.
You should have deep, demonstrable experience across most of these areas:
-
Advanced RAG & Retrieval: Chunking strategies, hybrid search (dense + sparse), metadata filtering, re-ranking pipelines, and vector-store lifecycle management (e.g. Pinecone, Weaviate, pgvector)
-
Prompt Engineering & LLM Integration: Designing reliable prompt pipelines with structured outputs, chain-of-thought reasoning, and function/tool calling across multiple model providers
-
Agentic Orchestration: Multi-agent coordination (e.g. LangGraph, AutoGen, CrewAI), state and memory management, human-in-the-loop patterns, and tool-use orchestration
-
Systematic Evaluation: Defining golden test sets and automated eval pipelines (RAGAS, G-Eval, LLM-as-a-judge) to measure accuracy, faithfulness, and hallucination rates rather than relying on manual spot checks
-
Production Serving & Cost Management: Deploying LLM-backed services within latency SLAs and token-cost budgets, including caching strategies, rate-limit handling, fallback/retry logic, and observability (tracing, logging, alerting)
-
Governance & Guardrails: Building prompt firewalls, output filters, and PII-handling pipelines to ensure compliance with data-privacy frameworks and prepare for EU AI Act obligations
-
Document Intelligence: Extraction, classification, and structuring of complex multi-format documents (PDFs, scanned images, tables) using LLM and multi-modal pipelines
You should understand transformer architectures well enough to reason about practical trade-offs: why retrieval-augmented generation outperforms fine-tuning for certain tasks, when to use smaller distilled models vs. frontier APIs, and how context-window limits affect pipeline design.
This is a software engineering adjacent role. You must write clean, tested, production-ready Python. You should be comfortable with CI/CD pipelines, code review, version control, and shipping behind feature flags. Research fluency (reading papers, reproducing techniques) is a plus, not a substitute.
A legal background is not required; however, a genuine interest in legal work is essential. Candidates who are intrigued by contracts and legal processes will find this position well suited to their interests. Experience with document automation or familiarity with legal workflows will be considered a significant advantage.
This is a hands-on role focused on building practical solutions that lawyers will use daily, not academic research.
What Youโll Actually Be Doing
-
Own production AI systems end-to-end: Design, build, deploy, and monitor LLM-powered document analysis pipelines (extraction, classification, risk flagging) that serve lawyers daily, meeting defined latency SLAs and accuracy benchmarks
-
Data Engineering: Transform legal data into structured, high-quality datasets that power our AI systems.
-
Design and ship agent workflows: Architect multi-step, multi-agent systems for complex legal tasks (e.g. due diligence, contract review, regulatory analysis) with robust state management, tool calling, and human-in-the-loop checkpoints
-
Define evaluation and governance frameworks: Build golden test sets, automated eval pipelines, and regression suites. Implement guardrails (prompt firewalls, output filters, PII redaction) to ensure safety and regulatory readiness
-
Optimise cost and performance: Manage token budgets across model providers, implement caching and batching strategies, and make data-driven build-vs-buy decisions on model selection (API-served, open-source vs. proprietary)
-
Collaborate with lawyers and product: Translate ambiguous legal workflows into structured AI problems, define acceptance criteria with domain experts, and iterate based on user feedback and eval results
Qualifications
What We Need You To Have:
-
At least 3 years of professional experience building and deploying AI systems, with at least 1-2 years focused on LLM/GenAI applications in production (not just prototypes or research)
-
Deep experience with document-heavy NLP: extraction from complex layouts (PDFs, tables, scanned documents), entity recognition, and structured output generation
-
Proven ability to design and optimise RAG pipelines and prompt architectures for accuracy, cost, and latency in production
-
Strong Python engineering skills. Familiarity with at least one LLM orchestration framework (LangGraph, LlamaIndex, or equivalent) and at least one vector database (Pinecone, Weaviate, pgvector, or equivalent)
-
Experience with cloud platforms (AWS or Azure preferred) for deploying and monitoring LLM-backed services, including CI/CD, containerisation, and observability tooling
-
Experience designing systematic evaluation methodologies for generative AI (automated evals, golden test sets, faithfulness/hallucination metrics)
-
Ability to own a workstream end-to-end: scope it, build it, evaluate it, ship it, and clearly communicate tradeoffs and results to non-technical stakeholders
Extra Credit
-
Experience in a startup or high-velocity AI team where you shipped frequently and wore multiple hats
-
Master's or PhD in Computer Science, Computational Linguistics, Mathematics, or a related quantitative field
-
Vector databases, retrieval systems, or knowledge graphs experience
-
Familiarity with model-serving infrastructure (vLLM, TGI, Triton) and GPU-aware deployment
-
Domain experience in legal tech, compliance tech, or other regulated industries where accuracy and auditability are non-negotiable
-
Published research or significant open-source contributions in NLP, information retrieval, or generative AI
-
Experience with knowledge graphs, ontologies, or semantic reasoning over structured legal data
-
Experience with Spark or Databricks and related technologies
The estimated base salary range for this position is $200,000 to $240,000 at the time of posting. The actual salary offered will depend on a variety of job-related factors, including skills, education, training, credentials, experience, scope and complexity of role responsibilities, geographic location, and performance. This role is exempt meaning it is not overtime pay eligible.
Cleary provides a comprehensive benefits package, including health care benefits. More information can be found here: Benefits
We are an equal opportunity employer and prohibit discrimination based on any category protected by law. Cleary provides reasonable accommodations to enable otherwise qualified employees to perform the essential functions of their position, provided the accommodation does not pose an undue hardship to the Firm.
Click on the following links to view the California Privacy Policy and Notice at Collection for California Residents .