
Junior Generative AI Engineer
Cleary Gottlieb Steen & Hamilton LLP
Job description
Cleary Gottlieb is a pioneer in globalizing the legal profession. We have 14 offices in major financial centers around the world, but we operate as a single, integrated global partnership and not as a U.S. firm with a network of overseas locations. The firm employs approximately 1,100 lawyers from more than 50 countries.
Since 1946 our lawyers and staff have worked across practices, industries, jurisdictions, and continents to provide clients with simple, actionable approaches to their most complex legal and business challenges, whether domestic or international. We support every client relationship with intellectual agility, commercial acumen, and a human touch.
We're an internal team at Cleary Gottlieb building bespoke AI solutions for legal work, drawing on software engineering, data science, and deep domain expertise. We work at the boundary of enterprise-grade software, and we are regularly building services and applications nobody here has built before. That makes for a dynamic learning environment โ and one where a junior engineer will gain exposure to real-world AI systems from day one.
We are a tight-knit, remote-first team that values growth, honesty, and curiosity. We work collaboratively across multiple disciplines โ engineering, data science, legal domain experts, and product โ to deliver meaningful impact. Junior engineers are paired with experienced mentors and participate in structured code reviews, pair-programming sessions, and weekly knowledge-sharing stand-ups.
You'll join the AI Acceleration Team as a Junior Generative AI Engineer, reporting to the Data Science Manager and working day to day with our engineers, data scientists, and legal domain experts. Nobody arrives knowing legal workflows or our exact stack โ curiosity and an appetite for increasing responsibility count for more here than the length of a CV.
Responsibilities
The Role:
Under the guidance of senior engineers, you'll help prototype, build, test, and deploy LLM-powered products that extract intelligence from complex legal documents and deliver measurable accuracy improvements. Bring a solid foundation in software engineering and genuine enthusiasm for AI โ we'll ask about projects you've worked on (personal, academic, or professional) where you used AI or ML tools: what you built, what you learned, and what you'd do differently.
What Youโll Be Doing
-
Contribute to production AI systems: Under the guidance of senior engineers, help build, test, and improve LLM-powered document analysis features (extraction, classification, risk flagging) that serve lawyers daily
-
Data Preparation & Engineering: Help clean, structure, and label legal data to power our AI systems โ building the data pipelines that make models work well in practice
-
Build features within agent workflows : Implement bounded components of multi-step AI workflows for legal tasks (e.g. due diligence, contract review) โ writing tool integrations, prompt chains, and retrieval logic under the supervision of senior engineers
-
Support evaluation, quality, and performance: Write test cases, contribute to automated eval pipelines, and help maintain regression suites. Gain exposure to model selection, token budgets, caching strategies, and the guardrails (output filters, PII redaction) that keep production AI systems reliable
-
Collaborate with lawyers, engineers, and product: Participate in cross-functional sessions to understand legal workflows, contribute to acceptance criteria, and iterate based on user feedback and evaluation results
How We'll Support You
We believe in building engineers, not just hiring them. As a junior member of the team, you'll receive:
-
Dedicated mentor: A senior engineer assigned as your day-to-day mentor for at least your first six months โ available for architecture discussions, code review, and career guidance
-
Structured onboarding: A 90-day onboarding plan covering our stack, legal domain fundamentals, and AI evaluation practices so you're never left guessing what to learn next
-
Pair programming: Regular pair-programming sessions with senior engineers on production features โ the fastest way to absorb patterns, tooling, and judgement
-
Code review as learning: Every pull request receives a thorough, constructive review. You'll also review others' code early on, reading good code is one of the best ways to write it
-
AI safety and bias training: Guided training on identifying non-deterministic failure modes, prompt injection risks, bias in LLM outputs, and the responsible-AI practices that govern our systems
-
Growth path: A clear progression framework from Junior to Mid-Level to Senior, with regular check-ins, stretch goals, and the opportunity to take on increasing ownership as your skills develop
Qualifications
What We Need You To Have:
-
Some software development experience (internships, co-ops, or substantial academic/personal projects count), including hands-on exposure to LLM/GenAI tools or APIs and some experience working with text data (parsing PDFs or HTML, basic text processing, or using an LLM for extraction or classification)
-
Solid Python skills and comfort with basic SQL. Conceptual understanding of how RAG pipelines work, with ideally some hands-on experience. Familiarity with at least one LLM API (OpenAI, Anthropic, or similar). Exposure to an orchestration framework (LangChain, LlamaIndex) or a vector database is a plus
-
Basic familiarity with cloud platforms (AWS or Azure) and an understanding of how to evaluate generative AI outputs (accuracy, faithfulness, hallucination). You don't need deep expertise in either area, but you should understand the fundamentals and be willing to learn
-
Strong communication skills: ability to explain what you've built, ask good questions, and clearly describe problems and trade-offs to both technical and non-technical teammates
Nice To Have
-
Experience in a hackathon, open-source project, or fast-paced team environment where you shipped working software quickly
-
A degree (Bachelor's or Master's) in Computer Science, Mathematics, Computational Linguistics, or a related quantitative field
-
Hands-on experience with a vector database or retrieval system (Pinecone, Weaviate, pgvector, FAISS)
-
Curiosity about model-serving infrastructure (vLLM, TGI) or how open-source models are deployed โ you don't need production experience, but interest matters
-
Any exposure to legal tech, compliance tech, or other regulated industries where accuracy and auditability are critical
-
A personal blog, portfolio project, or open-source contribution demonstrating your interest in NLP, information retrieval, or generative AI
-
Interest in or coursework on knowledge graphs, ontologies, or semantic reasoning
-
Any experience with Spark, Databricks, or data pipeline tooling
The estimated base salary range for this position is $120,000 to $160,000 at the time of posting. The actual salary offered will depend on a variety of job-related factors, including skills, education, training, credentials, experience, scope and complexity of role responsibilities, geographic location, and performance. This role is exempt meaning it is not overtime pay eligible.
Cleary provides a comprehensive benefits package, including health care benefits. More information can be found here: Benefits
We are an equal opportunity employer and prohibit discrimination based on any category protected by law. Cleary provides reasonable accommodations to enable otherwise qualified employees to perform the essential functions of their position, provided the accommodation does not pose an undue hardship to the Firm.
Click on the following links to view the California Privacy Policy and Notice at Collection for California Residents .