figma ai eval process.md

The Art of Evals: How Figma Put People at the Center of Its AI Product

AI tools are changing the entire process of product-building, where people of all skill levels can turn an idea into something that actually works — something they can see, feel, interact with and iterate on. As Apple engineering leader Michael Lopp says, democratization is a good thing because it makes this capability available to anyone, but it also makes for an extremely crowded market.

What’s that mean for builders who want to create a standout product? Maybe taste matters more than ever. Or maybe it’s speed.

Figma’s newest tool, Figma Make, places human craft and creativity at the center of that product-building process. The company’s new prompt-to-functional-app experience that just launched at Config (along with three other products: Sites, Buzz and Draw) further blurs the line between design and production, reducing the technical skills required to actually bring a product to life.

David Kossnick, Figma’s Head of Product, AI, didn’t just make humans the focus of the product’s experience; they were also the focus of the product’s development, and specifically, its evaluation process.

Much like using an AI product, building an AI product requires a different approach. Unlike traditional software where there’s a clearer path to see what’s possible, the capabilities of an AI product exist in a foggy middle ground that’s only validated through actual testing.

Here on The Review, we’ve looked at a few of the different ways Figma makes decisions rooted in how real humans use its products. FigJam was born out of people in the community using Figma as a whiteboarding and brainstorming tool during the pandemic, according to the company’s CPO, Yuhki Yamashita. Figma Slides was a “bottoms-up project that came to life via a series of internal viral moments,” says its founding PM, Mihika Kapoor.

In this interview, we explore the evaluation process Kossnick and team used to launch Figma Make, one that kept humans at the center of every step — from defining success metrics, to the process for gathering qualitative feedback and then exploring how you can assess this data. If you’re interested in building, testing and validating AI products, this one is for you.

Developing the infrastructure shared by different AI products

Figma Sites was a massive infrastructure undertaking across the whole company, involving different rendering technologies, Kossnick says. But that’s the groundwork that made it possible to build Figma Make so quickly.

Sites allows you to publish a Figma design as a public website. To do this, Figma had to bridge the gap between design tools and web publishing — developing entirely new systems that could translate design elements into functional web code. Kossnick gives the example of converting a blue rectangle with specific dimensions in a design to HTML and CSS that browsers could render correctly. This was done by deterministic code-gen, not AI code-gen.

But as they were developing Sites, AI code-gen models were rapidly advancing. Then, a designer had an idea: What if, when you’re designing a static website, you could make these components functional with AI?

“That turned into a hackathon project and was extremely compelling,” says Kossnick. “It drew a ton of internal attention and excitement, people playing with it, showing examples. But it barely worked end-to-end. Failure rate was high. When it did work, it was incredible.

This idea led the team to develop the second crucial pre-component of Figma Make — Code Layers in Figma Sites, which added three important capabilities. The first was making code a primitive on Figma’s canvas, allowing users to write code directly inside of Figma. The second was converting designs into code, and specifically React, not just HTML and CSS. And third, it created a chat interface where users could prompt to add coded interactions and behaviors. These innovations tackled some of the most difficult aspects of design-to-code conversion.

Code Layers in Figma

With this technology being tested internally, another hackathon led to the concept of Figma Make. A designer, tinkering on the side, created a standalone prototype of Code Layers as its own surface, where the interface lent itself to users asking AI to build a whole site or app, not just a component on a page. “It worked surprisingly well, a surprising percentage of the time,” Kossnick says.

Figma’s decision tree to determine AI product viability

Kossnick says developing an AI product is difficult because of its malleability. “It’s easy to look at a product and imagine any part of the surface where AI could fit,” he says. “So deciding what not to do is really important.”

These are the four different paths for product development he uses to assess if allocating more time and resources into any AI project is worth it:

Path 1: The technology isn’t ready yet
“In AI product development, a prototype is becoming the gold standard as a validation mechanism before really starting on projects,” Kossnick says. This is especially true as the cost of prototyping has decreased tremendously.

Path 2: It’s almost possible (with a lot of custom development)
This is really a consideration in how much work you’re willing to put into a project and its ability to scale.

Path 3: It’s possible, but you need to adjust the product
The path here is a bit clearer if you’re able to ruthlessly prioritize and narrow scope.

Path 4: It works
This is the happy path, where you’re striking the exact right intersection of technology and product capabilities.

Kossnick says builders need to ask themselves where they’re at in this decision tree. And once you’ve identified a path, speed is important — when prototyping happens fast, so can validation.

Tips for constructing your AI product team

Each AI product has its own set of goals and constraints, which determine the structure of the team building it. Kossnick shares some of what he learned from staffing Figma Make:

  1. Role blending lets you keep the team small (even if you’re at a big company): AI tools bleed the stark lines in skillsets between different functions. Designers can code. PMs can prototype. Engineers can design. “A designer wrote the first system prompt for Figma Make,” he says.
  2. Almost everyone should be touching code: AI tooling makes this far more possible than it was years or even months ago.
  3. Treat AI products as centralized teams: Code Layers and Make operated as one large integrated team sharing technologies and infrastructure with two different UX treatments.
  4. Get your target persona involved in the eval process: It was a conscious decision to have designers and PMs in the eval process because they’d be the ones using the product.

AI tooling is changing how all teams operate, not just engineering, product and design. But on the technical teams actually building AI products, embracing fluidity across roles allows processes to adapt to this new reality, resulting in greater pace and efficiency.

Figma’s three-step, human-centric eval process

Qualitative feedback is an important part of assessing traditional software, whether it’s user research programs or behavior tracking. Scaling these approaches can work for deterministic software where behavior is mostly predictable once coded. But AI products require a more continuous and widely-scoped approach to qualitative feedback — in large part due to the probabilistic nature of AI outputs and creative subjectivity in determining “good” performance.

Prototyping is so valuable here, especially when it’s increasingly more common, cheaper and expressive. It’s a tool to obtain feedback and sharpen a product quickly. “Continuous prototyping and refining enabled rapid validation,” says Kossnick.

Figma’s process for defining, obtaining and evaluating the product was rooted in what real people expected of it, how they used it and their quality assessment of its outputs. To use that human feedback — whether in the next prototype or the features that’d make it into the final version — Kossnick and his team had to figure out how to scale human taste while also making it actionable.

1. Define the success metrics that actually matter to your persona

Picking goal metrics is absolutely critical. As part of that, it’s important to have good coverage of classes and scenarios you care about (a mockup vs. a prompt, one shot vs. long shot conversion, desktop vs. mobile). “Every slice adds more scenarios to cover, so be rigorous about how many are core and in what order,” Kossnick says.

To determine what good looked and felt like, Kossnick considered three key elements of usability: the personas, the scenarios in which they’d use the product and which deliverables they expected.

For Figma Make, Kossnick used two key evaluation metrics: design score and functionality score, where each would be graded on a scale of 1 - 4:

2. Gather qualitative, human feedback at scale

Toil away in the qualitative metrics all you want, but there’s no substitute for getting your product into the hands of people and seeing what they do with it, even (and especially) at the earliest possible stages.

Figma had four increasingly broadening concentric circles of user feedback as part of their eval process: the internal AI team building the product, the PM and design teams as target personas, the entire company to see what was possible, and an alpha group of customers for final viability assessments.

The first group of 30 people on the AI team received an “unoptimized prototype,” says Kossnick.

At this point, there wasn’t a thumbs up or thumbs down feature in the product. But given the relatively small set of testers, the team could be scrappy.

To find it, Kossnick and the team broadened the feedback group to the entire PM and design teams — their intended target persona.

This time, instead of using Slack threads, they made a giant FigJam board where they asked users for the same feedback: prompt, result, and scores.

The next phase of feedback expanded company-wide, where the team hosted “The Great Figma Bakeoff,” which was a much larger version of the FigJam board they used with the PM and design teams.

“It’s easy to over-engineer your eval stack, your data set, some part of the quality loop — but it all depends on what users want to use your product for,” says Kossnick.

3. Figure out how to assess the data you’ve gathered

Kossnick follows one simple rule to make sure evals are working: are we moving in the right direction? “What do you really want to test and what do you want the outcome to be?” he asks.

Even though he admits the initial data collection wasn’t the most sophisticated, it was extremely valuable — so he and the team had to figure out the best form factor in order to make it useful at scale.

Still focusing on product quality for a core persona (designers) and use case (designers bringing prototypes to life), Kossnick used four evaluation types:

Deterministic: Pretty straightforward, like a pass/fail class: did it do the thing or not?
Taste and judgement: This requires humans, and humans at scale, but the considerations are cost and speed.

AI as judge: “Can you take human judgement and teach AI how to do it?” Kossnick says.

Usage analytics: A/B testing in production is also a form of evaluation.

Each stage of Figma Make’s ideation, development and testing can be traced back to product experience — people use the product, and thus should play a large role in the evaluation of its quality. Kossnick was intentional about bringing Make’s target users into evals.