🔍 Browse Jobs 🏢 Companies 🗂️ Categories 📝 Career Blog 🔒 Privacy
figma
📍 San Francisco, CA • New York, NY • United States

Director, Research - AI Evals

About the Role
<div class="content-intro"><p>Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us!</p></div><p class="font-claude-response-body break-words whitespace-normal">The Figma Research team is hiring an Director, Research - AI Evals to own how we measure the quality of Figma's AI-powered experiences. As Figma ships more AI capabilities across our products, the question "is this really good?" has never mattered more — and answering it rigorously is what this role exists to do. You'll define what "good" means for our AI features, build the frameworks and quality bars to measure it, and turn that into trusted signal that product teams rely on to decide what to ship.</p> <p class="font-claude-response-body break-words whitespace-normal">The ideal candidate brings deep, hands-on experience evaluating AI/LLM-powered products — blending human evaluation with automated, model-based approaches — along with the product instinct and communication skills to make evaluation genuinely useful. Partnering with Product, Design, Engineering, and Data Science, you'll sit upstream of nearly every AI shipping decision at Figma and directly shape the quality of features used by millions of people.</p> <p class="font-claude-response-body break-words whitespace-normal">This is a full time role that can be held from one of our US hubs or remotely in the United States.</p> <p class="font-claude-response-body break-words whitespace-normal"><strong>What you'll do at Figma:</strong></p> <ul class="[li_&]:mb-0 [li_&]:mt-1 [li_&]:gap-1 [&:not(:last-child)_ul]:pb-1 [&:not(:last-child)_ol]:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3"> <li class="font-claude-response-body whitespace-normal break-words pl-2">Own AI evaluation methods and operations for Figma's AI-powered experiences — define quality dimensions, design how we measure them, and turn results into decision-ready signal</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Build and maintain evaluation frameworks, rubrics, golden datasets, and quality bars, combining human evaluation with automated/model-based approaches (e.g., LLM-as-judge) where appropriate</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Partner with engineering to stand up repeatable, reproducible evaluation pipelines and regression testing, so evaluation is a routine part of how AI features are built and shipped</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Produce clear readouts and dashboards that let stakeholders confidently make go/no-go and prioritization decisions</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Socialize a shared definition of quality so evaluation standards are adopted across teams rather than re-invented — and advocate for evaluation as a strategic partner in the product process</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Manage a small team to execute our AI evals in partnership with contractors, internal staff, and/or LLMs</li> </ul> <p class="font-claude-response-body break-words whitespace-normal"><strong>We'd love to hear from you if you have:</strong></p> <ul class="[li_&]:mb-0 [li_&]:mt-1 [li_&]:gap-1 [&:not(:last-child)_ul]:pb-1 [&:not(:last-child)_ol]:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3"> <li class="font-claude-response-body whitespace-normal break-words pl-2">10+ years of experience in product, research, applied research, or a closely related field, including 2+ years of management experience</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Direct, hands-on experience owning the evaluation of AI/LLM-powered products</li> <li class="font-claude-response-body whitespace-normal break-words pl-2">Expertise designing and running AI evaluation — human evaluation programs, rubric and benchmark/golden-dataset construction, inter-rater reliability — and sound judgment about when and how to apply automated/model-based approaches (e.g., LLM-as-judge), including their l
Apply Now →
📮 Post a Remote Job
Reach thousands of remote job seekers. We review every listing before it goes live.