Alia
Most of the support tickets were the same four questions. Answering them automatically was the easy part.
Try the prototypethe whole product in one pass: the entry, what Alia is, the terms, a guided first question, an answer you can rate.
What was wrong. Yellow Card’s support team was answering the same handful of questions on repeat: how do I deposit, how do I trade stablecoins, how do I verify, why is my deposit pending. Every one was a ticket, every ticket was a wait, and the people with genuinely hard problems were queued behind them.
Why it mattered. How do you put a language model inside a money app without letting its confidence stand in for the company's? A language model will answer anything in the same even voice, right or wrong. In a support article a wrong answer is a bad article. In a chat about someone’s pending deposit it is a financial event. So the design work was not the conversation. It was the frame around it.
What I designed. Alia, an assistant inside the app that answers those questions at the moment they come up. Not the model, and not the answers: an assistant that is allowed to be helpful only where it is also allowed to say it does not know, and that always has somewhere to send you when it can't.
What it looked like. An entry in the header, an introduction and a sheet, four questions, and every state after that. The parts that matter most are what happens when it goes wrong, further down; the assistant itself is at the end, and you can talk to it.
- role
- product designer. I owned the whole Alia experience: the conversation model, the entry and introduction, the disclaimer, the suggested questions, the input rules, every failure state, the feedback loop and the hand-off to a human.
- team
- Alia sat between three owners. The support team owned the answers and the ticket queue; engineering owned the model integration and the chat infrastructure; I owned everything the customer saw and every state the conversation could be in.
- status
- shipped in the Yellow Card consumer app.
scope, signals and what was not mine
- problem
- a support queue made of the same four questions, inside a money app where a confident wrong answer costs more than a ticket.
- scope
- entry from the home screen · introduction and disclaimer · the four suggested questions · free input and its limit · understood, not-understood, slow and offline states · thumbs feedback · the hand-off to support · the copy for every fixed line.
- not mine
- which questions the model could answer, the knowledge base behind it, the support agent's side of a ticket, and the model itself were owned elsewhere and are out of scope here.
- signals
- escalation to a human opened above 20% and settled around 9% within two weeks. what that number can and cannot say is defined below.
the tickets were repetitive. the risk was not.
The volume was not mysterious. Support could name the questions from memory, and the ticket data agreed: deposits, stablecoins, verification, pending. They had known answers. What they did not have was a place to be answered at the moment the customer was confused, which was inside the app, usually with a transfer in flight.
The obvious fix was a chatbot, and the obvious risk was that a chatbot about money is not a chatbot about shoes. A model will answer “why is my deposit pending” in the same steady voice whether it knows or not, and the customer has no way to tell the two apart. The product had to make that difference visible before the customer could be hurt by it.
a better help centre
accurate, cheap, already exists. also already where people were not looking; the tickets came from inside the app, at the moment of confusion.
a scripted bot with a menu
safe, because it can only say what it was given. and a dead end the moment a question is not on the menu, which is most questions.
an assistant that answers anything
helpful on the first try, and confidently wrong on the second. in a support article a wrong answer is a bad article. in a chat about a pending deposit it is a financial event.
a scoped assistant with designed exits
the model answers, but inside a frame: what it is, what it can help with, what it does when it does not understand, and where you go when it was not enough.
The four suggested questions on the first screen are the four questions from the ticket data, in order. That is not a design flourish; it is the product admitting what it is for. Anything else is typed, and typed questions get one of two things: an answer, or an honest sentence saying it did not understand.
Alia is allowed to be helpful only where it is also allowed to say it does not know.
Everything in the product follows from that sentence: the disclaimer, the guided questions, the clarification, the input limit, and the hand-off.

before the first message: what Alia is, and what it isn’t
Alia lives one tap from the home screen, as Ask Alia in the header. That entry says nothing about AI, on purpose: the introduction does. The first screen is a name, a mark, one sentence, I’m your AI-powered assistant, and a link to the disclaimer. Then a sheet rises over it and asks the customer to acknowledge what an AI answer is before the first one is sent.
The sheet is legal text, and I would not pretend anyone reads it. Its job is the tap. I understand is the moment the customer agrees that what follows is generated and may be wrong, and the product is allowed to hold them to that. What the sheet does not do is make the answers trustworthy. That takes the rest of the system, and it starts right after the sheet goes down.
The disclaimer is the beginning of the trust model, not the whole of it.

up close: the introduction and the sheet
a blank box is a bad first question
The worst state of a chat is the empty one: a cursor, a placeholder, and a person who does not know what this thing can do. So Alia asks first. Its opening message is one sentence and four cards, and tapping a card sends it as the customer’s own message. Most conversations never need the keyboard.
The keyboard is still there, because the four questions are the most common ones, not the only ones. But free input has a limit: 160 characters, with a counter, and at the limit the counter turns red and the field stops. Short questions map well to answers. Long stories do not, and long stories are exactly the conversations that needed a person anyway, so the limit pushes them toward one instead of toward a worse answer.
up close: the input, empty and full
the happy path was easy. the product is what happens when it isn’t.
A question and an answer took a week to design. The map of everything else took the rest of the project. I drew the conversation as a flow before any screen existed, and most of it is not the happy path: is the model on or off, is it slow, did it understand, was the answer any good, does the customer want a person, does the person turn up.
Each branch became a state the interface had to be able to show, and each state got a rule for what Alia must do and what it must not. The must-nots are the important half. A model left alone will guess, complete your sentence for you, and answer the nearest question it knows. In this product, every one of those is a bug.
the states
open the table of statesfive states, four questions each
| understood | not understood | slow | offline | rated down | |
|---|---|---|---|---|---|
| what came in | a question the model can map to something in the knowledge base | gibberish, a fragment, a language it was not given, a question outside the product | any question, on a bad day for the model | the customer opens Alia while the model is unavailable | a thumbs-down on an answer |
| what Alia knows | the topic, and an answer with reasonable confidence | that it does not know | that a reply is coming, later than usual | nothing, and that it knows nothing | that its answer was not enough for this person |
| what it must do | answer briefly and ask to be rated | say so, quote the customer's own words back, ask for more detail | show it is working, and say so if it takes long | say it is offline and offer something that works without it | apologise and offer a human, immediately |
| what it must not do | answer about the customer's own account or balance, which it cannot see | guess, or answer the nearest question it does know | sit silent, or time out into nothing | pretend, or show an empty chat | try again, ask what was wrong first, or leave it at the rating |
| what the customer sees | the answer, a timestamp, and thumbs up / down | “I don’t quite understand what you mean by …”. no thumbs: a clarification is not an answer | the typing dots, then a line that it is taking longer than usual | an offline line, two articles to read, and a disabled input | “Sorry I couldn’t help. Would you like me to open a ticket…?” with yes and no |


the rules
open the rulesif, then, because
Alia is about to speak for the first time
thenit says what it is, in its own words, and asks the customer to acknowledge what that means before the first message
becausethe introduction is the only moment everyone reads. after that the model is talking, and nothing it says can be trusted to explain itself.
the model cannot map a message to something it knows
thenAlia quotes the message back and asks for more, and the reply carries no thumbs
becausequoting proves it read the words. guessing the nearest answer would look identical to a right answer, and in a money app that is the dangerous case.
a customer rates an answer down
thenthe next thing on screen is an offer to hand over to a person, with a yes and a no
becausea thumbs-down is not feedback for a dashboard. it is someone saying this did not work, and the only honest response is a way out.
the customer has more than 160 characters to say
thenthe counter turns red and the input stops, before sending
becauseshort questions map well; long stories do not, and long stories are exactly the conversations that needed a person anyway.
the model is unavailable
thenAlia says so and offers articles; the input is disabled
becausean empty chat with a working input is a promise the product cannot keep.
edge cases
- gibberish
- “jfne9%%7&bcdskni$$&8hzbcxlcs” gets the clarification, with the string quoted back.
- a fragment
- “So how can for profits as to start money per” gets the same clarification. the model must not complete the sentence for the customer.
- a typo
- “How ling will it take for my M-PESA deposit to reflect?” is answered as if it were spelled right; the answer is the same one the correctly spelled question got.
- the limit
- at 160 characters the counter turns red and the field stops accepting input. the customer trims, or takes the story to a ticket.
- a question about their own account
- Alia can explain how deposits work; it cannot see this deposit. the answer says what usually happens and when to contact support.
- thumbs-down, then no
- declining the ticket closes the offer; the chat stays open.
- the agent does not connect
- an agent who never arrives gets a waiting state and, after a while, the ticket carries on without the chat; the customer is told either way.
- coming back
- the disclaimer is acknowledged once; a returning customer lands in the chat with the link to it still in the introduction.
a thumbs-down is a request, not a rating
Every answer ends in two buttons. Thumbs-up is a data point. Thumbs-down is not: it is a person saying this did not work, about their money, and the only honest reply is a way out. So the next thing on screen is not a second attempt or a survey. It is Sorry I couldn’t help. Would you like me to open a ticket for you so our support team can take over?, with a yes and a no.
That gives the conversation exactly two endings. Either the customer got what they came for, or a human has the thread. There is no third state where someone argues with a bot about a transfer. The rule cost us something: every thumbs-down opens a ticket offer, including the ones where the customer just disliked the tone, and support absorbs that. It was the right side to err on.
Every conversation resolves or hands off. “No dead ends” is a rule about what the last message can be.


what we measured, and what it can say
The number I watched was not how many people used Alia. It was how many asked for a human on the first try. If the suggested questions were getting people in the door without helping them, escalation would stay high, and the whole thing would be a slower way to open a ticket.
Window and cohort: the first two weeks after release, all conversations, measured as the share that ended in a hand-off to a human. No baseline existed: before Alia every one of these was a ticket.
how it was measured, and what it does not prove
- measured
- the share of Alia conversations that ended in a hand-off to a human
- against
- the first days after release, when it opened above 20%; there was no earlier baseline, because before Alia every one of these was a ticket
- why it mattered
- it was the number that said whether the guided questions were helping people or just getting them in the door. staying above 20% would have meant the suggestions were an entrance, not an answer
- does not prove
- that the answers were right. it counts the people who asked for a human, not the people who should have
how it was measured, and what it does not prove
- measured
- the same share, in the first days
- against
- itself, two weeks later
- why it mattered
- the early number was mostly people testing what Alia could handle; it had to fall for the product to be doing its job
- does not prove
- how much of the fall was the product improving and how much was the testers going away
what kind of number each one is
open the definitions
- behaviour
- what people did after an answer: rated it, asked for a person, or left. these are the only numbers here.
- not measured
- answer correctness, ticket deflection against the previous period, and whether a customer who stayed in the chat actually got what they came for. each would need a different instrument, and the reflection says which one I would build first.
Nine percent is low enough that the guided questions were doing their job, and not so low that people who needed a human were stuck convincing a bot first. That is the whole claim. It is a claim about the exits, and it holds.
What it is not is a claim about the answers. Fewer hand-offs can mean better answers, or quieter customers. Telling those apart needed a person reading conversations, and that instrument did not exist while I was there.
what this changed, and what would change it back
The assumption that changed. I went in thinking the work was the conversation: tone, the suggested questions, making a bot sound like Yellow Card. That was a fortnight. The work that took the rest of the project was deciding what Alia does in the moments the model cannot carry: a sentence it cannot parse, an answer that was wrong, an outage. The failure states are the product; the happy path is a demo.
What it taught me about AI in a money product. In a money product, a confident wrong answer is worse than no answer. So the design has to be more honest than the model: say what it is before it speaks, quote the customer's words back when it does not understand rather than guess, and treat a thumbs-down as a request for a person rather than a data point. None of that makes the model better. All of it makes being wrong survivable.
Not the same outcome. Escalation falling from above 20% to around 9% is not the same as Alia being right. It says fewer people asked for a human; it does not say the ones who stayed got a correct answer. A rate that low could also mean people stopped bothering to rate at all. Correctness needed a different measure, and we did not have one.
What I would measure earlier. Answer quality on a sample, from the first day: a support agent reading a hundred conversations a week and marking each answer right, partial or wrong. The escalation rate told us how the exits behaved. It said nothing about the answers, and the answers were the point.
What is still unclear. Whether the disclaimer changed anything. Everyone tapped I understand; nobody could have told you what it said. It did the legal job. Whether it did the trust job, or whether the trust came from the clarification and the hand-off, I cannot separate.
what would make me change the rules
open the list
- a sampled review finds Alia confidently wrong about money more than once in a hundred answers: the assistant is switched to clarification-only for that topic until the knowledge base is fixed.
- thumbs-down falls to near zero while tickets opened from inside the chat do not: people have stopped rating, not started being helped, and the hand-off needs to be offered without the rating.
- the 160-character limit shows up in support tickets as a complaint: the limit rises, or long messages route straight to a ticket with the text attached.
- customers ask about their own balances and transfers more than about how things work: Alia either gets read access to the account, with a different disclaimer, or stops pretending to be the place for those questions.
the frame is a bet that a model’s uncertainty can be made visible and cheap to act on. any of these would mean the bet is not paying.
try it
Ask it something it can’t answer.
Start on the home screen, open Alia, get through the sheet, and tap a card. Then type. Anything about deposits, verification, stablecoins or pending transfers gets an answer; anything else gets the clarification. Rate an answer down and watch what happens next.
things worth trying, in order
Things worth doing, in order:
- Ask Alia, then I understand
- tap How do I deposit to my wallet?
- type nonsense and send it
- thumbs-down an answer → Yes, open a ticket
- keep typing past 160 characters
- use the buttons below for the states you cannot reach by tapping





