PM Framework

HEART Framework

Google's five categories of user-experience metric, the Goals-Signals-Metrics process that makes them work, and why you should never use all five.

The paper that introduced HEART tells you not to use all five categories. Almost every team that adopts it builds a five-row dashboard anyway, and then stops looking at it within a quarter.

That gap between the framework as published and the framework as practised is the whole subject of this page.

What is the HEART framework?

HEART is a set of five categories for user-experience metrics, published by Kerry Rodden, Hilary Hutchinson and Xin Fu of Google in Measuring the User Experience on a Large Scale: User-Centered Metrics for Web Applications at CHI 2010. The categories are Happiness, Engagement, Adoption, Retention and Task success.

It exists because product teams were reporting page views and server logs, which say how much traffic a product got and nothing about whether anyone was better off. HEART was an attempt to put the user's experience into a number a large organisation would actually track.

  • Happiness. Attitude. How users say they feel about the product. Survey territory.
  • Engagement. Depth of voluntary interaction in a period. Actions per user, sessions per week.
  • Adoption. New users taking up a product or a feature in a window.
  • Retention. Users from a starting cohort who are still present later.
  • Task success. Efficiency and completion on a defined task. Completion rate, time on task, error rate.

What is Goals, Signals, Metrics?

This is the half of HEART that gets dropped, and it is the half that does the work.

The paper pairs the five categories with a three-step process. You start from the goal for a category, which is a sentence about what success means for users, written before any measurement exists. You then find a signal, an observable behaviour or attitude that would move if the goal were met. Only then do you choose a metric, a specific number you can instrument and chart.

Teams that skip to the third step end up with whatever their analytics tool already tracks, which is how a HEART dashboard fills up with counts of clicks. The chain matters because it keeps the number attached to a claim about users that somebody can disagree with.

One recommendation from the paper is worth repeating on its own: use ratios. A raw count of engaged users rises whenever the user base rises, which means it measures your growth team. Normalise by the number of active users in the period or the number stops meaning anything.

Which HEART categories should you actually use?

Pick the ones where the goal is real for your product. The table gives the question each category answers, an example signal and metric, and the condition under which the category is worth skipping.

CategoryThe question it answersSignalMetric you can instrumentSkip it when
HappinessDo users like it?Users say they would be upset to lose itIn-product survey score, share choosing the top optionYou have no way to reach churned users, so the sample is survivors only
EngagementDo they use it voluntarily, and how deeply?Sessions and core actions per active user per weekCore actions per weekly active userThe product has a natural cycle, such as tax filing or payroll
AdoptionAre new users taking it up?New users reaching first successful useShare of new accounts completing the core action in 7 daysThe product is used once, which makes Adoption and Retention the same number
RetentionDo they come back?Cohort still active in a later periodShare of the week-0 cohort active in week 4You have not been live long enough for a week-4 cohort to exist
Task successCan they get the job done?Task completed without error or abandonmentCompletion rate and median time on taskThe surface is open-ended and there is no single defined task

Categories and signals follow Rodden, Hutchinson and Fu (CHI 2010). The skip conditions are the practical limits teams hit.

A worked HEART example: comments in a document editor

Take a comment feature inside a document tool. The numbers below are a worked set, chosen so the arithmetic is checkable rather than because they describe any real product.

The goal, written first: a reviewer can raise a point on a document and see it resolved without leaving the document. That sentence rules two categories out immediately. Happiness is not the question here, because nobody will tell you in a survey whether a specific comment thread worked. Adoption and Task success are the question.

Say there are 42,000 documents opened in a week, and 3,100 of them get at least one comment. Adoption is 3,100 / 42,000, which is 7.4%. Of the 8,900 comment threads opened that week, 5,430 are marked resolved, so task success is 61%.

Now the trap. Four weeks later the comment count has gone from 8,900 to 10,500, an 18% rise, and the team reports a win. But documents opened has gone from 42,000 to 56,000 while commented documents only went from 3,100 to 3,700, so adoption has fallen from 7.4% to 6.6%. Resolved threads went to 6,090 of 10,500, which is 58% rather than 61%. The raw count rose and both rates fell. The product got worse at the goal while the dashboard got greener.

This is why the paper insists on ratios. A raw count in a growing product measures the growth. It says nothing at all about the experience.

Where does the HEART framework break down?

  1. Happiness surveys sample survivors. An in-product survey reaches the people still in the product. The users whose experience was worst have already left and cannot answer. A rising happiness score during rising churn is a normal and very misleading pattern.
  2. Engagement punishes products with a natural cycle. A payroll tool, a tax filer, an annual benefits enrolment: low engagement in week three is correct behaviour. Nothing is wrong. Measuring engagement here creates pressure to add reasons to open an app that people should not need to open.
  3. Adoption and Retention collapse on single-use products. For a product someone uses once, the two categories measure the same event and reporting both makes the dashboard look more informative than it is.
  4. Task success needs a defined task. On an open-ended canvas, a search box or a feed, there is no single task to complete, so completion rate has no denominator that means anything. Define a narrow task or leave the category out.
  5. Five rows becomes a dashboard nobody reads. The paper's own advice is to choose. A team that tracks all five ends up with a report that has something reassuring on it every week, which is the same as having no metric at all.
  6. The categories describe a web application. HEART was written in 2010 for large web apps. For an AI product where the output quality varies per request, the missing category is output quality, and none of the five holds it. You need an eval score alongside HEART rather than inside it.

How do PM interviews ask about HEART?

The question is almost never "what is HEART." It is a metrics question, and HEART is one way to answer it. The shapes to expect:

  • "How would you measure the success of [feature]?" Name a goal first, then one or two categories, then the metric. A candidate who lists all five categories has described a framework instead of answering the question.
  • "Engagement is up and revenue is flat. What is happening?" This tests whether you know engagement is a ratio and whether you check the denominator.
  • "Pick one metric for this product and defend it." Here HEART is the wrong tool on its own, because the interviewer wants a single number. Pair it with a north star metric and use HEART for the supporting set.
  • "Your survey score went up and retention went down. Which do you trust?" The survivor-bias answer. Say it plainly: the survey only reached the people who stayed.

The strongest version of a HEART answer names the goal, picks two categories, says which three you are leaving out and why, and gives a ratio rather than a count. That sequence shows you have used the framework rather than read about it.

HEART framework FAQs

What does HEART stand for?

Happiness, Engagement, Adoption, Retention and Task success. The five categories come from Kerry Rodden, Hilary Hutchinson and Xin Fu's CHI 2010 paper at Google.

Do I have to use all five HEART categories?

No, and the original paper says so. Choose the categories where you can write a real goal sentence for users. Two well-instrumented categories beat five that nobody checks.

What is the difference between HEART and Goals-Signals-Metrics?

HEART is the list of five categories. Goals-Signals-Metrics is the three-step process for turning a category into a number: write the user goal, find the observable signal, then pick the metric you can instrument. They were published together and the process is the part teams skip.

Is HEART the same as a north star metric?

No. A north star metric is one number the whole company aligns on. HEART is a set of categories for building out a balanced view of user experience, usually several metrics at once. They work together: the north star sits above, and HEART shapes the supporting metrics beneath it.

How does HEART apply to an AI feature?

The five categories still work for usage and satisfaction, and they miss output quality entirely, which for an AI feature is often the thing that decides whether users come back. Run an eval score next to the HEART metrics rather than trying to fit quality into Task success.

Related reading