
My role
Lead AI Product Designer
Company
Amazon
Focus
Agentic UX · Human-AI Interaction · Adaptive UI · Trust & Recovery · AI Evaluation
Mattress returns are expensive. Once used, many cannot be resold, creating high costs for returns, disposal, and logistics.
At Amazon, I led the design of an AI shopping agent (Rufus, now Alexa) to help customers choose the right product before purchasing.
The work started with mattresses and expanded into a scalable framework for other teams for high-consideration products: expensive purchases like furniture, appliances, and electronics that require customers to weigh multiple factors before deciding.
The goal: increase purchase confidence while reducing costly returns and cancellations.
Buying a mattress online is unusually difficult. Customers can’t touch or try the product before purchasing, yet comfort is highly personal.
After examining the full process, I chose initial intake as the first place to apply AI.
Early AI shopping experiences often behaved like conversational search. Customers asked a question, got an answer, and still had to piece together the decision themselves.
I designed the agent to learn what mattered to each customer, remember preferences, weigh trade-offs, and recommend products based on their needs.
The opportunity was to use AI to help customers arrive at the right product before they purchased.
Deciding what the AI should control
The agent needed freedom to understand personal preferences, weigh trade-offs, and make useful recommendations.
But not every decision should be left to AI. Requirements like allergies, budget limits, compatibility, and service availability needed to be consistently respected.
The trade-off: The agent had less freedom, but customers could trust it to respect what mattered most to them.
Using the right interface for the moment
Conversation worked well for understanding what customers wanted, but not every answer belonged in a chat bubble.
I designed the experience to adapt as customers shopped. Recommendations could become product cards, comparisons could become tables, and simple questions could stay conversational.
This made complex information easier to understand and act on.
The trade-off: A more adaptive experience required more design and development work, but made the agent easier to use than a chat-only experience.
When the AI doesn't know
Early testing showed the AI could confidently give customers incorrect information, including claiming Amazon delivery services didn’t exist and directing them to competitors.
I designed the experience to rely on verified product and service information instead of the AI’s general knowledge. When information couldn’t be verified, the agent acknowledged that rather than guessing.
The trade-off: The agent sometimes had to say “I don’t know,” but that was better than confidently giving customers the wrong answer.
Defining what “good” looks like
AI quality is subjective unless you define what success means.
I created separate criteria for system integrity and customer experience, then tested the agent against realistic shopping scenarios.
The evaluations uncovered issues that looked fine in individual demos, including problems that required changes to the experience itself rather than another prompt adjustment.
The eval framework could be adopted across teams for high consideration shopping.
The trade-off: Evaluation added time throughout the design process, but gave us evidence to improve the product instead of relying on demos and intuition.