LOOP Discover
LOOP already had both sides of a marketplace. They just weren’t connected.
Try the prototype
where it ends up: “find somewhere quiet for dinner tomorrow in westlands under 4k.”
LOOP is NCBA’s neobank. Discover is the marketplace inside its consumer app. About 300,000 merchants and 600,000 consumers were registered, but only a fraction used LOOP each month. Merchants managed offers in a separate deals platform. Customers had to find those offers for themselves.
I connected merchant products to Discover, built a recommendation engine, then designed a simpler way for customers to say what they wanted. The work started with merchant tools, moved into the consumer app and reached a controlled beta.
- role
- product design and design engineering lead, from marketplace strategy through implementation and beta.
- team
- LOOP merchant and consumer teams · one PM · four engineers · data/ML · merchant operations.
- status
- controlled beta, 120 participating merchants.
the problem, scope, timeline and early signals
- problem
- activate an under-connected two-sided marketplace and increase merchant payment volume.
- scope
- marketplace strategy · recommendations · product design · working prototype · frontend implementation · beta testing.
- timeline
- may 2026 to present.
- early signals
- payment volume linked to Discover +26% · Discover MAU +17%, first eight beta weeks compared with the previous eight. promising, but still early.
the shortest path was already inside LOOP
I led this from the merchant product team. Our team had become responsible for the merchant business, and the number we needed to move was payment volume.
The original Discover was a static deal shelf. Every customer saw almost the same offers. Merchants also had to recreate and update each offer by hand in a separate deals platform.
We considered three routes. I focused on the one that could turn LOOP’s existing merchants, products and customers into transactions fastest.
rebrand
changes how LOOP looks and how people feel about it. six to nine months. a long way from any actual transaction.
acquire more merchants
more supply, eventually. expensive, slow, and it adds to a base that was mostly dormant already.
use what LOOP already has
the merchants, their inventory, the consumers, and the payment rail sitting between them. the shortest path to a transaction.
I took three steps rather than trying to solve everything with one redesign.
First, connect the products. LOOP launched its merchant ERP in Q1 2026. About 56% of eligible merchants had adopted it, and that number was growing. I proposed linking the ERP to the deals platform. A merchant could choose a product they already managed, set a lower price for Discover and publish it without rebuilding the listing.
Second, personalise the marketplace. As more offers came in, I led the recommendation engine that decided what each customer should see first.
LOOP already had useful signals from merchant payments, card purchases, mobile-money activity and location when available. With the consent already held by LOOP, these signals gave each customer a more relevant starting point than the same static shelf for everyone.
BEFORE Q1 ERP
merchant · inventory, POS
✕ separate workflow
DISCOVER · deals uploaded by hand
✕ rarely in the way
consumer · 600k registered, 70k active
CONNECTED MODEL
ERP product → existing deals portal
↓
opt in · lower Discover price
↓
signals → recommendations → DISCOVER
↓
payment → volume

recommendations helped, but they could only guess
The recommendation engine made Discover more relevant by using payment patterns and nearby offers. But as more offers arrived, the feed became crowded. The same deal could appear more than once, and every card tried to show badges, price, savings, distance and expiry at the same time.

I simplified the hierarchy around three questions: where am I, what can I search for, and what is worth seeing now? Each deal appeared once, and secondary detail waited until it mattered.
I stripped each card back to a clear order: merchant and distance, offer, then price. Expiry appears only when it matters. The last line shows one useful reason, such as what the customer saves, who it serves or why it matches the request. The image keeps one discount badge and one save button.
That improved browsing, but recommendations were still guesses. They could infer that someone liked restaurants; they could not know they needed dinner for two tomorrow in Westlands, under KES 4,000, somewhere quiet.
Third, let customers say what they need right now. Filters could build the request one control at a time. A sentence could do it in one go. That is why I explored chat. It was not the original idea; it was the next step after recommendations reached their limit.
the obvious answer was chat
Once people could use natural language, I had to decide where it should live. I tested the obvious answer first: a chatbot in front of the marketplace. It greeted people, suggested questions and returned deals in message bubbles. After a few messages, earlier results were hard to find and the marketplace had disappeared behind the conversation.
It also forced people into chat when many only wanted to browse.


That gave us a clear rule: keep the marketplace first, make the agent available everywhere and never require chat. A request should update the results, filters or comparison in place, not become another message.
That ruled out an assistant home screen, conversation history and bubbles. One field instead had to support search, natural language and voice without looking like three different controls.

DISCOVER
│
┌──────────┴──────────┐
↓ ↓
browse express intent
↓ ↓
deals UI text · voice · filters · taps
└──────────┬──────────┘
↓
the same marketplace
then the happy path broke
Natural language worked quickly. You could type dinner for two in Westlands under 4k and get three deals with a one-line explanation. The first version worked beautifully, as long as you never changed your mind.
Then:
- actually, make that four
- somewhere quieter
- what about Kilimani?
- not Java House
- ignore the budget
A search box treats Enter as the end. Chat turns every change into a new message and can lose track of what came before. Neither worked. The real design problem was not how to accept a sentence, but how to let someone keep changing it.
The request is not a message. It is something people can change.
the sentence isn’t the state
A person speaks in sentences. The system stores the request as a small set of facts: the task, must-haves, preferences, exclusions and assumptions. It also keeps the original sentence so the person can edit it. Typing, speaking, choosing a filter or tapping a deal all update the same request.
Before Discover searches, it checks that those facts still make sense together. It removes contradictions, expired details and constraints that no longer fit the task. Without that check, someone could search for a TV, say “actually headphones,” and accidentally keep a 55-inch screen size. Search only receives the clean, current request.
utterance
↓
interpretation
↓
intent mutation
↓
validated current intent
↓
find and order offers
↓
record why each result appeared
↓
presentation
“find somewhere quiet for dinner tomorrow in Westlands under 4k”
- task
- Food & Drinks · dinner
- requirements
- Westlands · tomorrow · up to KES 4,000
- preference
- quiet
- also held
- exclusions · ignored slots · assumptions · scope · what changed last · history
TV under 50k in Westlands, then actually headphones. The budget and the area survive because they still make sense. The product words go. And because carrying constraints across a change like that can catch people out, the screen says what it kept and gives you a way to drop it.

a request you can see and change
Showing every part of the request as a chip made the interface feel like a debugging tool. I replaced it with a quieter header: what Discover thinks you want, which assumptions affect the results and an Edit action for the full detail.
People can change the same request in different ways. They can type, speak, use a filter, tap a deal or ask the agent to change the budget. The method can change from one step to the next without starting over.



four operations cover most changes
open the four operationschange, exclude, ignore, explore
- change
- under 6kup to 4k becomes up to 6k. never both.
- exclude
- not Java Housestays excluded through every later change.
- ignore
- don’t care about the budgetbudget becomes any. Discover does not guess a new one.
- explore
- what about Kilimani?kilimani results, with a way back to westlands. the user never learns the word branch.



what language is allowed to understand, change and do
open the ruleswhat language may understand, change and do
simple requests stay simple
Discover responds to what a request needs, not how long it is. Coffee filters the marketplace while you type. Coffee under 500 applies the category and budget directly. Only subjective, comparative or multi-step requests use the model. The orb appears only when the work takes noticeable time.
the agent has clear limits
- Explain. Answer using information already on screen.
- Change the view. Update results and offer undo.
- Save something. Complete the save, then offer undo.
- Prepare, but do not confirm. Check the latest price and availability, then show the customer exactly what will happen.
- Stop before payment. The customer always confirms and pays.
The higher the risk, the less the agent can do on its own.
start over, refine or explore
The main field starts a new request. The refine bar changes the results already on screen. What about Kilimani? briefly explores another option and keeps a route back. People only need to see what changed and how to undo it.
The system also prevents conflicting details from building up silently. A deliberate choice such as Westlands or Karen is valid. Two incompatible constraints are not.
Voice is another way to edit the same request, not a separate mode. The field expands, the words appear as they are spoken and the transcript stays editable for a moment. If Discover is unsure about an amount, it flags it instead of guessing. The request then follows the same path as typing.
Two small rules. If the field already has a draft in it, speaking adds to the draft and waits. If the field is empty, what you say becomes the request.

the smaller rules: voice, search area, history and the bottom bar
A short word such as pizza filters the results already on screen. If none match, Discover keeps those results and offers to search the whole marketplace. If the customer already asked to search all of Discover, it does so. The agent saves work without behaving unpredictably.
Discover keeps three things separate: what the customer asked, what it showed and what the customer did. They are separated because each one must undo differently. Undo that reverses the latest action. The original ones restores an earlier set of results. No chat transcript is needed.
A permanent message box at the bottom would make the marketplace feel like chat again. Instead, the bottom bar appears after results and says Refine these results. Voice sits beside it, with Cheaper · Closer · Open now behind a toggle. I also rejected a chat-first home, a chip for every detail, a rewritten sentence after every change, a message box that never leaves and an orb that never stops moving.
text ────────┐ voice ───────┤ filters ─────┤ controls ────┼──→ current intent ──→ results taps ────────┤ agent ───────┘



testing and validation
No single test could tell us whether Discover worked. We needed to know if people understood it, if the model understood them, if the results were accurate and if any change reached merchant payment volume.
So we tested in three stages: the prototype before development, the model and search results during development, then customer behaviour and business results in the controlled beta.
first, could people understand and control it?
I tested working prototypes with internal LOOP customers and a smaller group of external customers. They browsed deals, searched within a budget, changed areas, removed a condition, recovered from no results and returned to an interrupted request.
We watched where people finished, hesitated, went the wrong way or corrected themselves. We also asked why they thought the results had changed. Reaching the right screen was not enough if they did not understand how they got there.
To move quickly, we tested early versions internally first. We then ran A/B comparisons with internal participants and the controlled beta group. We compared task completion, refinements, deal opens and drop-off. These tests helped us choose what to build next, but they did not prove that Discover caused more transactions.
open the observationswhat we saw, and the change it caused
| chat first | browsing disappeared behind the conversation | the marketplace remained the default surface |
|---|---|---|
| every detail as a chip | clear, but technical and visually heavy | a short request summary, with detail behind Edit |
| summary without an action | people read it as a static heading | an explicit Edit action |
| details kept from an earlier search | people missed what was still active | kept details became visible and removable |
| no exact result | people assumed the system had failed | honest alternatives and one-change recovery |
| voice and location | unclear amounts and fallback areas could pass unnoticed | editable words, flagged uncertainty and a location note above results |
how the model was trained and testedthe model, the data, the automated suite
then, was the system correct?
We started with a multilingual open-weight model, which we could run in LOOP’s own environment. We trained it for one narrow job: turning a customer’s words into the request Discover understands. Each training example paired a sentence with the correct task, location, budget, time, preferences and exclusions.
The examples included Kenyan neighbourhoods and landmarks, KES amounts, English, Swahili and mixed-language phrasing. We added human-reviewed variations and anonymous beta corrections when our data rules allowed it.
Payment activity was not used to teach the language model how to understand sentences. It belonged to the separate recommendation and ranking system that decided which valid offers to show first.
We also tested language understanding separately from search quality. The model could understand a request correctly and still retrieve poor offers. A useful offer did not prove that the request had been understood correctly.
the tests ran automatically
We turned a held-back set of labelled requests into an automated test suite. Every change to the model, prompt or validation rules ran against the same set before release.
The suite checked whether the output used the required structure, captured every condition, avoided adding details the customer had not said, rejected contradictions and handled changes of mind correctly.
We tracked complete-request accuracy, accuracy for each field, precision, recall, unsupported details, contradictions, valid output and refinement accuracy. We also measured the time needed to produce a usable interpretation.
Results were split by English, Swahili, mixed-language requests, Kenyan locations and KES expressions. This stopped a strong overall score from hiding a local weakness. Search quality was evaluated separately against the expected catalogue results.
finally, did behaviour and commercial value change?
In the controlled beta, we measured the path from opening Discover to searching, refining, opening, saving, comparing and redeeming a deal. This separated whether people could use the product from whether they kept using it and whether it improved business results.
Some questions still need more evidence. Did Discover create new payment volume or move existing spending? Will people keep using it after the novelty wears off? Are merchants getting fair exposure? Will trust hold as the catalogue and model reach grow?
results people can trust
Not every part of a request has equal weight. Under 4k is a must-have. Preferably quiet is a preference. Must-haves remove results; preferences change their order. A preference should never hide an otherwise valid option. Exact matches and alternatives stay clearly separate.
When nothing fits, Discover shows which single change would unlock real options. A customer can also set a target-price alert and anonymously share that demand with the merchant.
The agent can find, compare, save and prepare, but it cannot pay. Price and availability are checked again before confirmation; any change is shown before the person decides.

why results appeared, old requests and location
Discover records why each result appeared: which must-haves and preferences it met and which it missed. Why these? reads from that record instead of asking the model to invent an explanation.
Near me can fail quietly when location is off. Discover therefore shows the last area it used above the results and gives the customer one tap to change it. It never applies a distance limit without a usable location.
Time-based details also expire at different speeds. Open now lasts about ten minutes, Near me about half an hour and Tonight until the day ends. A fixed area, budget, product or merchant does not expire simply because time passed. When someone returns, Discover removes expired details and says what changed.
Because this is a banking app, Discover does not show an old private request on an empty screen. After a short period, it offers Continue where you left off · Dinner · 2 h ago. The full request appears only if the customer chooses to continue.


designing for results that take time
The orb shows that Discover is working. It is not an AI mascot. It appears where the results will arrive, then disappears when they are ready. People can keep using the screen while it works.
Each request gets its own identifier. If someone changes their mind before the first request finishes, Discover keeps the newest request and ignores the older response. The same rule works for edits, filters, navigation and results that arrive out of order.


how loading works, and what can go wrong
The screen moves through searching, comparing and preparing. The first card appears in place, the rest follow, then the orb disappears. The prototype slows this down so the sequence is easy to see. The real product should show instant results immediately and use the orb only when work truly takes time.
The simple path was easy. The list below covers what can go wrong and separates what is working in the prototype from what is only designed.
edge-case coverage
| intent | edit · replace · refine · contradict · exclude · ignore · branch · undo · new task · task change with carry-over | prototype |
|---|---|---|
| scope | local vs global · zero local matches · ambiguous short nouns · explicit “search all” | prototype |
| voice | low-confidence amount · interruption · appending to a draft · half a sentence never submitted | prototype, transcription simulated |
| location & time | location off · stale “near me” · open now expiring · tonight expiring · resume after hours | prototype, per-slot lifetimes |
| results | no exact matches · partial matches · preference misses · duplicate recommendations across sections | prototype |
| async | slow request · instant request · out-of-order completion · interaction during loading | prototype, single simulated request |
| transaction | price changed · deal ended · availability changed | prototype states (?live); production must check the live catalogue |
| voice permission | microphone unavailable or denied | designed, not built |
taking it into the codebase
I led the work from the merchant team and took it end to end across the consumer product. The consumer team owned the wider app. Engineering, data and merchant operations owned their production systems. I held the product direction across those teams.
I built the interaction layer in the LOOP codebase: one input for search, language and voice; the editable request; refinement; progressive results; comparison; alerts; stale-price handling; confirmation states; and responsive layouts. It shipped in the beta behind a feature flag.
Engineering reviewed my implementation and connected it to the catalogue, search and redemption services. Data and ML owned model interpretation and evaluation. Together, we set a simple memory rule: keep a detail only while it still makes sense for the current request.
Once the interaction, safety limits and loading behaviour were clear, the work could move from prototype logic into the live marketplace.
from prototype to beta
Discover is in a controlled beta with 120 merchants. A multilingual open-weight model, which can run within LOOP’s own environment, turns each sentence into the structured request described above. Fixed rules check that request before any search begins. The model never makes up products, prices or availability.
We test English, Swahili and mixed-language shopping requests, Kenyan places, local product terms and KES amounts. Search uses live offers from the connected ERP and deals platform. The catalogue remains the source of truth for price, stock and validity.
Access to the catalogue did not answer every product question. Helping someone find a merchant is different from automatically steering them to a cheaper competitor. The beta therefore checks four things separately: whether the model works, whether customers and merchants trust it, whether it creates business value and whether it works reliably in production.
model qualitycan it understand local requests and find the right offers?
- problem
- local phrases, English mixed with Swahili, shillings and landmarks can confuse a general model.
- rule
- “it feels intelligent” is not a useful measure. we test whether it understands each detail, returns the right offers, avoids invented claims and handles corrections.
- measuring
- valid output, accuracy on local language, search accuracy, time to a useful result, recovery when nothing matches and corrections after results.
- still learning
- which failures need more training and which need better instructions, search or fixed checks.
marketplace trustcan it help customers without working against merchants?
- problem
- the agent can compare merchants and say who is cheaper. that helps the customer but can work against the merchant whose catalogue made the comparison possible.
- rule
- the agent follows what the customer asks. “show me something cheaper” starts a search, but a product page does not automatically suggest a competitor. a clear request matters more than a guess about the customer.
- measuring
- whether new and small merchants are seen, whether comparison leads to redemptions and whether merchants keep participating.
- still learning
- how much ranking should depend on relevance, distance, freshness, merchant quality or commercial value.
business resultsdoes better discovery create new transactions, or move existing ones?
- problem
- a transaction that starts in Discover might have happened elsewhere in LOOP anyway.
- rule
- Discover redemptions have their own transaction tag. We count them when the customer redeems the offer and compare the beta with the same period before it.
- measuring
- Discover payment volume, active users and the share of opened deals that are redeemed.
- still learning
- how much payment volume is truly new. the tag proves Discover took part, but not that the transaction would otherwise have never happened.
production reliabilitydoes it still work when data is late, wrong or missing?
- problem
- a smart model showing an old price or the same deal twice is still a bad marketplace.
- rule
- Discover checks stock, price, opening hours and deal validity again before the customer confirms. If the model fails, the ordinary marketplace still works.
- measuring
- old information found at confirmation, catalogue errors and how often customers need the basic fallback.
- still learning
- the merchant controls: what information is connected, why a merchant appears, how the data is used and how the merchant can reverse it. not designed yet.
early signals from the beta
Across the 120 participating merchants, payment volume linked to Discover was 26% higher in the first eight beta weeks than in the previous eight. Discover MAU was 17% higher over the same period. This counts people who opened Discover and completed at least one meaningful action.
These are promising early results, not final proof. The transaction tag shows that Discover played a part in the payment. It does not prove the payment was completely new or separate the effect of seasons, promotions and catalogue changes.
what this changed
I started with a common belief: if recommendations became accurate enough, the marketplace would feel personal. The work changed my mind. Past behaviour can improve the starting point, but it cannot explain what someone needs right now. The better answer was to combine useful guesses with a request the customer could see and change.
If I started again, I would define the model test set and the payment-volume comparison before designing the final interaction. We added both during the work, but earlier definitions would have made weak language cases and the limits of the beta results visible sooner.
The beta has supported the value of better discovery, but it has also challenged the idea that attributed payment volume is enough. Discover can show that it took part in a transaction; it cannot yet show how much of that spending was completely new.
I deliberately stopped the agent before payment and left merchant controls for comparison outside this beta. If stronger testing shows that payment volume is only moving around, or that comparison makes smaller merchants harder to find, I would revisit the ranking and comparison model before widening the product.
try it
Enough explaining. Try it.
This is the whole prototype. Type, tap the mic, scroll, open a deal, come back.
things worth saying to it, in order
Some things worth saying to it, in order:
- dinner for two in Westlands under 4k, preferably somewhere quiet
- what about Kilimani?
- don’t care about the budget
- not Java House
- compare these with the original
- undo that
what is real and what is simulated in the prototype
The interaction itself is real. The prototype stores and validates the request, handles edits and keeps track of why each result appeared. That is the code running when you use it.
The embedded prototype understands a limited set of phrases through fixed rules, not a language model. The beta uses the model for that step, while the same request structure and safety checks remain underneath.
Voice uses prepared phrases that appear word by word, and its confidence level is simulated. Location uses a fixed previous area. Price and availability changes are designed prototype states. Only the two beta results above are measured outcomes; everything else demonstrates behaviour, not performance.







