Back

Designing an AI Agent for High-Stakes Shopping Decisions

My role

Lead AI Product Designer

Company

Amazon

Focus

Agentic UX · Human-AI Interaction · Adaptive UI · Trust & Recovery · AI Evaluation

Background

Mattress returns are expensive. Once used, many cannot be resold, creating high costs for returns, disposal, and logistics.

At Amazon, I led the design of an AI shopping agent (Rufus, now Alexa) to help customers choose the right product before purchasing.

The work started with mattresses and expanded into a scalable framework for other teams for high-consideration products: expensive purchases like furniture, appliances, and electronics that require customers to weigh multiple factors before deciding.

The goal: increase purchase confidence while reducing costly returns and cancellations.

What I delivered
  • $20M reduction in operational costs
  • Reduced mattress returns
  • Deployed AI shopping agent
  • Working AI-powered prototypes
  • Guardrails, Error Recovery, Evaluation
  • Design system tokens & components
  • Reusable framework for high-consideration shopping

The Summary

Buying a mattress online is unusually difficult. Customers can’t touch or try the product before purchasing, yet comfort is highly personal.

  • They have to balance factors like firmness, sleeping position, temperature, partner preferences, price, allergies, and delivery needs.
  • Traditional search and filters help customers find mattresses. They don’t do much to help customers decide which mattress is right for them.
  • Conversational AI created an opportunity to understand those subjective needs and help customers work through the trade-offs before purchasing.
  • Previous implementation was designed around low cost products, handled high consideration purchases poorly
  • Previous implementation would hallucinate and claim Amazon Services such as Haul Away disposal or Installation did not exist, and even directed customers to competitors.

The Strategy

Helping customers make a decision, not just search

After examining the full process, I chose initial intake as the first place to apply AI.

Early AI shopping experiences often behaved like conversational search. Customers asked a question, got an answer, and still had to piece together the decision themselves.

I designed the agent to learn what mattered to each customer, remember preferences, weigh trade-offs, and recommend products based on their needs.

The opportunity was to use AI to help customers arrive at the right product before they purchased.

Designing Agent Autonomy

Deciding what the AI should control

The agent needed freedom to understand personal preferences, weigh trade-offs, and make useful recommendations.

But not every decision should be left to AI. Requirements like allergies, budget limits, compatibility, and service availability needed to be consistently respected.

The trade-off: The agent had less freedom, but customers could trust it to respect what mattered most to them.

Adaptive UI

Using the right interface for the moment

Conversation worked well for understanding what customers wanted, but not every answer belonged in a chat bubble.

I designed the experience to adapt as customers shopped. Recommendations could become product cards, comparisons could become tables, and simple questions could stay conversational.

This made complex information easier to understand and act on.

The trade-off: A more adaptive experience required more design and development work, but made the agent easier to use than a chat-only experience.

Uncertainty, failure, and recovery

When the AI doesn't know

Early testing showed the AI could confidently give customers incorrect information, including claiming Amazon delivery services didn’t exist and directing them to competitors.

I designed the experience to rely on verified product and service information instead of the AI’s general knowledge. When information couldn’t be verified, the agent acknowledged that rather than guessing.

The trade-off: The agent sometimes had to say “I don’t know,” but that was better than confidently giving customers the wrong answer.

Evaluating the Experience

Defining what “good” looks like

AI quality is subjective unless you define what success means.

I created separate criteria for system integrity and customer experience, then tested the agent against realistic shopping scenarios.

The evaluations uncovered issues that looked fine in individual demos, including problems that required changes to the experience itself rather than another prompt adjustment.

The eval framework could be adopted across teams for high consideration shopping.

The trade-off: Evaluation added time throughout the design process, but gave us evidence to improve the product instead of relying on demos and intuition.