Back

Result

The problem wasn't the number, it was the invisible value

An end-to-end iteration of Result, TaxDown's most important screen: the last one before users decide whether to pay, and the only one that's always there.

TL;DR

01 · Problem

Users who saw their result didn't understand what TaxDown added over the Tax Agency's draft return: if the number looked similar, they saw no reason to pay. And users who saw no number at all converted 2.4 times worse.

02 · Solution

I turned the three barriers from research into principles and shifted from showing a figure to explaining its value: first with cohort experiments, then with a personalized, AI-generated explanation.

03 · Impact

When the result can't be calculated, explaining why and giving an estimate lifts payment by +17.8%, with significance. When the result can be calculated, the AI explanation trends positive and is already live for everyone.

+17.8%

Conversion to paid when the result can't be calculated · Exp 1, significant

+9.05%

Submission for review when the result can't be calculated · Exp 1

+6.1%

Payment with the AI explanation when the result can be calculated · relative, not significant (≈+20% before an external blocker)

+21%

Conversion to paid from Result, year over year · 8.6% → 10.4%

  • Role Product Design Lead
  • Team Squad Imagineering
  • Platform Desktop · iOS · Android · Responsive
  • Timeline Apr → Jul 2026

01 · Context

The screen every user goes through

The key moment

Result

It shows the result, the savings from deductions the user didn't know about (≈€400 on average) and the value compared with filing on your own.

Case 1

Can be calculated

We have all the information and can calculate reliably: the result, the savings and why it's worth continuing. ~80% see their result at the end of the flow.

Case 2

Can't be calculated

Nine complex tax situations rule out a reliable calculation. Rather than risk a wrong figure, we explain that an expert review is needed.

02 · Business problem

If users don't understand the screen, they don't convert

Many users come just to find out their result and leave without buying. And conversion only holds up when it's a refund: a scalability problem.

Refund

9.34%

Convert to paid. The only case where the number sells itself.

Amount owed

4.05%

Less than half. The bad news comes with no value attached.

Break-even (0) · hit hardest

2.15%

If the number matches the draft, users see no reason to pay.

The segment that comes out worse than the draft has the lowest NPS: 41.0, and 36.1 if they also owe money. Right before paying.

03 · User problem

Users don't see TaxDown's invisible work

Result "works" visually, but it doesn't communicate what we do under the hood. If the number looks like the draft's, users assume they gain nothing by paying. Renta Web, the government's official filing program, is the competitor in their heads.

Value · disappointment

"How is this different from the government's program?"

402 support conversations (≈10% of the season) about discrepancies with the draft. Over 60% of those handled by a tax expert saw no value.

Clarity · overwhelm

"I didn't know what to do"

Price is the most cited reason for not converting, closely tied to not knowing what the service includes or what to do next.

Trust · insecurity

"I trusted it, but with doubts"

Users who don't see their result have an NPS 8.3 points lower (66.7 vs 75.0) and convert 2.4 times worse (4.78% vs 11.58%).

Evidence · live metrics, with the season already underway

Survey + interviews intent and non-conversion

3 barriers

Insecurity, disappointment and overwhelm. Three arrival intents: convinced-but-stuck, explorer, and simulator who comes to compare.

Amplitude conversion by type

>4×

gap between the best and worst case: 9.34% for refunds vs 2.15% for break-even (0).

NPS matrix difference vs draft

41.0

NPS of those who come out worse. And a nuance: those with no changes (64.2) are more satisfied than those who get improvements (56.5).

Support pre-payment

402

conversations (≈10%) opened over discrepancies with the Tax Agency's draft.

Constraints · the limits I designed within
  • Depends on power of attorney We can only compare side by side with the Tax Agency when we have the user's data.
  • Insufficient instrumentation The savings property never flagged results worse than the draft: the priority segment was invisible.
  • Fixed season window Experiments ran from April 30 to July 6; Smart Result went from design to production in ~3 weeks.

04 · UX Thinking

From showing a number to explaining its value.

I first tackled when the result can't be calculated. What I learned there led to AI when the result can be calculated.

  1. Against overwhelm

    Guide

    Make the next step crystal clear.

  2. Against disappointment

    Prove value

    Explain the difference from the draft.

  3. Against insecurity

    Build trust

    Make the human review visible.

Step 1 · When the result can't be calculated

After: estimated result, why your return is complex and 3 steps with a tax expert
Before: 'Your advisor will review your case', with no figure or explanation

Explaining why turned waiting into trust

With no figure and no explanation, this was the worst-converting cohort. I added an estimated result, why their return is complex and 3 steps with a tax expert.

4.78% → 67.5%Submission for review without vs with an explanation (~14×)
+17.8%Payment with the estimated result · Exp 1, significant

Step 2 · When the result can be calculated: AI

A generic message didn't work (Exp 2, reverted). What worked was explaining each specific case, and with so many scenarios that can't be written by hand: we generate it with AI.

  1. 01

    Reads the scenario

    Tax profile and difference from the draft.

  2. 02

    An LLM processes it

    Returns a structure, not free text.

  3. 03

    Rendered in my design

    In the slots I designed.

Smart Result design with the slots the AI fills in

Scroll inside the screen

My role

I designed the mold, not the copy

The design is a placeholder: the content changes, but the AI always renders it in the slots I defined. That keeps the screen consistent and easy for users to understand.

  1. Result and color by type
  2. Difference vs the draft
  3. Generated explanation
  4. Report with labeled changes
  5. A single CTA

UI design · when the result can be calculated

The same structure for everyone, content tailored to each user.

"Congratulations!", "We're sorry" or "Even", with no explanation at all.

After: refund
Refund
After: amount owed
Amount owed
After: break-even (0)
Break-even (0)
Before: Congratulations! You're getting a refund
Refund
Before: We're sorry, you owe money
Amount owed
Before: Even, nothing to pay and no refund
Break-even (0)
Scenario matrix · what the AI absorbs

Each user sees a different combination of type × plan × savings.

01 · Result typeRefundAmount owed+ Not required to fileBreak-even (0)Can't be calculated
02 · PlanBasicPro / LiveFull + Validation check
03 · SavingsWith savingsNo savings
04 · Detail · 6 statesEmptyWarningSavingStuffWarning & SavingWarning & Stuff
Under the hood · instrumentation

Measure what matters

The savings property hid the users who came out worse than the draft. A signed difference vs the draft let us segment the message, feed the AI and measure properly. A ~1-day change that unlocked the entire analysis.

Design principles

3 principles · tap each one

“Every result, its own message.” The content adapts to the result type and the tax profile: there's no one-size-fits-all Result.

“Nothing without explaining where it comes from.” Every figure shows what we did and how it differs from the draft.

“One screen, one decision.” Prioritize what helps users decide and strip out noise that doesn't add to conversion.

Key decisions
D1

A personalized, AI-generated explanation against the draft

Exp 2 proved that a generic message doesn't move the needle. The prompt absorbs the combinations of scenarios and profiles instead of a hand-built message matrix.

D2

Communicate the invisible work

Deductions, adjustments and scenarios applied. The "worse" case is reframed as a better review: we added data the Tax Agency didn't have.

D3

Side-by-side comparison with the draft

It turns doubt into a visual answer. Scoped for the next iteration, since it depends on power of attorney.

D4

Highlight the human review

What the tax expert will do for you. "I trusted it, but with doubts" is the dominant response.

D5

Instrument with a signed difference

It complements the savings property to identify worse results, segment and measure.

Discarded alternatives
  • Detailed next steps when the result can be calculated. Tested as Exp 2 and reverted: submission for review dropped 4.75%. We simplified the content to avoid adding noise at the moment of decision.
Trade-offs
  • Partial comparison. Side by side is only possible with power of attorney; we prioritized explaining the difference on screen, which reaches every calculated result, over unreliable comparisons.

05 · Impact

The experiments validate the direction

Explaining moves the needle: strongly when the result can't be calculated, and when the result can be calculated only if the explanation is personalized, not with a generic message.

Three experiments · uplift vs control

Exp 1 · Can't be calculated
+9.05% · +17.8%✓ prod.
Exp 2 · Can be calculated
−4.75% · +1.51%↩ reverted
Smart Result · AI · can be calculated
+2.5% · +6.1%✓ 100%

Submission for review Payment

Where the number doesn't speak

+9.6%

in submissions for "same as the draft" and +4.4% in payment for "amount owed". For "refund", ~zero.

The precedent

67.5% vs 4.78%

submission for review with and without an explanation of why (~14×). Not an A/B test; Exp 1 confirmed it.

Year over year

8.6% → 10.4%

conversion to paid from Result: the combined effect of everything that changed on the screen.

In production

100%

The AI explanation has been available to everyone since July 6.

Smart Result didn't reach significance: the season ended with less than 10% of the target sample, and the payment lift held steady at ≈+20% until an external blocker before checkout contaminated the metric.

Before → after

Message"Congratulations!", "We're sorry" or "Even" + figure→A personalized, AI-generated explanation of what we did and why it differs
Invisible workNot shown→Deductions, adjustments and scenarios applied, made visible
Tax Agency comparisonNon-existent or generic→The difference is explained on screen; side by side is scoped
Human reviewImplicit→Visible: what the tax expert will do that the draft doesn't

What didn't work

01 · Generic message

Exp 2 was reverted. −4.75% in submission for review and +1.51% in payment, not significant. Without personalization, the needle doesn't move.

02 · Break-even (0)

Still the hardest hit (2.15% during the season). Smart Result shows its best relative signal there, but on a small sample.

03 · Restarts

Rose 1.31 pts, the only guardrail we missed. Still to be determined whether it's good or bad friction.

04 · Support

No data yet on the support impact: measuring the drop in discrepancy tickets requires a post-season window.

Learnings

Start

Personalizing the message by scenario and profile from the design stage, and instrumenting with a signed difference vs the draft.

Stop

Treating Result as a single message for everyone and assuming that showing the figure is enough.

Continue

Explaining before showing, the pattern that worked when the result can't be calculated, and making the human review visible at the moment of highest decision.

Next step · Generative Result with AI

01 · Decide

Choose the final AI system: the two we compared tied on every metric; a tax-quality review is still pending.

02 · Extend

Bring the personalized AI explanation to when the result can't be calculated and let users interact with it, using the same components.

03 · Modularize

A screen that adapts the value information to each situation: previous years, member get member, security and comparison with the draft.

Want to dig deeper?

Every detail, documented

Let's talk

I'm looking for products with more scale and complexity, where Design carries weight in decisions alongside Product and Tech: hard-to-structure problems, demanding teams and broader scope, without stepping away from building.