The problem
RelayPay grew, and its support team couldn't keep up.
RelayPay sells cross-border payments and invoicing to African startups and small businesses. The questions pile up: onboarding, pricing and fees, payout timelines, failed or delayed transactions, invoicing and compliance. Most are repetitive and already documented. The rest touch someone's account, a payment's status, a dispute or a frustrated customer, and need real care.
They needed a first-line support agent a customer can talk to. It answers from approved knowledge, checks a customer, transaction or payout when it's safe, opens a ticket or hands the case to a person when it isn't, and logs every step for the team to review.
- Approved answers only
- Sensitive details stay private
What I asked first
The brief left these open. I answered each one before writing any code.
- Where does the agent sit in a voice call? My server is the voice platform's brain, not a tool it calls. Two models in every turn would double the wait and let the second one rephrase what the first held back.
- "Check my account", on a line that shouldn't share account details? Both hold. A caller needs two facts that match one record, and a miss sounds exactly like a partial match. The plan and the status are said. The email and the internal notes never are.
- What about records that are weeks old? A payment due in August and still processing in late September is a problem to ticket, not a date to repeat. Code works out that the date has passed.
- What does a handoff need? A booked callback in a real calendar and an email to the team with everything: who called, what they asked, what they were told and what was checked. The customer never repeats themselves.
- Which model? Measured, not assumed. I wrote the rule before running anything, then ran both models on the same scenarios three times.
- What counts as resolved? Worked out in code when the call ends: answered, ticket opened, escalated, abandoned or failed. Nobody types it.
How it runs
- Person
- Tool
- AI
- Stored
- Check
- Result
Hear
01, Human:
A customer speaks or types
Caller
On the web page, by voice or by typing.
02, Step:
Speech to text
Vapi
And back to speech at the end.
Think
03, AI:
Decide what to do
Sonnet 5.5, Agent SDK
Answer, ask, look up, open a ticket or hand off.
04, Data:
Use the support tools
MCP server
Help content, records, tickets, callbacks. Every call logged.
Check
05, Logic:
Check before speaking
Code
Every number, reference and claim against what the tools returned.
Then: step 3 (fails? repair once); step 6.
Speak
06, Output:
Spoken back
Vapi
References and booked times are written by code.
07, Data:
On record for the team
Supabase, console
Each turn, tool call and check. Handoffs booked and emailed.
Enforced, not requested
- Nothing is spoken unchecked Every reply is checked against what this turn's tools actually returned, before the voice says a word.
- Numbers come from evidence Every number must appear in what was searched, looked up or said by the caller.
- Code writes the references Ticket numbers, booked times and dates are written by code from the tool result, never by the model.
- Two facts to verify Identity needs two details that agree on one record. A miss and a partial match get the same reply.
- No contact details, ever No tool can return an email or list customers, so a trick question fails at the tool.
- After a handoff, lookups stop The tool server enforces it, not the prompt.
- Safe to retry Voice calls repeat requests. A ticket or a handoff can't be created twice.
- Every tool call is logged By the tool server itself, refusals and errors included.
The thing that nearly got past me
Incident
A payment 41 days late, called "within the normal window"
In the first real runs, every answer about one payment said it was "processing within the normal expected window". Its estimate was 19 August, 41 days earlier. The record's summary was written when the record was, and the model read it out as if it were current.
The fixThe lookup now flags a summary whose date has passed. Code removes any sentence that repeats it, and writes the true one itself: "The record showed an estimated arrival of 19 August, which has passed, and it's still processing."
The build, in numbers
22 / 24
test scenarios passed end to end, through the real agent
59 / 60
benchmark runs passed on Sonnet 5.5, against 39 on Haiku 4.5
4.6 s
average turn on a live voice call, measured by Vapi
8
support tools, in an MCP server anyone can run without an account
82
questions used to measure the search threshold, not guess it
53
failures met in the build, each written down with its cause and fix
From the test runs and a live call.
What it refuses to do
- Guess a status or a date It says what the record shows, and when it was due.
- Read out contact details No tool returns them in the first place.
- Promise a timeline It explains what payout timing depends on, and offers to check.
- Explain a compliance review It routes the case to a specialist, never the reason.
- Say "booked" without a booking Only when the calendar has confirmed one.
- Change anything on an account Every customer and payment record is read-only to it.
What I took from it
On a support line, the model can talk. Code decides what it's allowed to say.
The limit I'd fix first
if I did it again
- The turn that books a callback is the slowest, and can pass the 14 second limit. I'd run the agent on an always-on server with a warm process, instead of starting one every turn.
- Identity is two matching facts, not real authentication. A live deployment would send a one-time code to the email on file.
Check it yourself
- Live app (opens in a new tab)
- Walkthrough video (opens in a new tab)
- Repository (opens in a new tab)
- Support console (opens in a new tab)For the support team. Password protected.
