Back to home
Banking
Conversational AI
Enterprise UX
AI Knowledge Assistant for
Banking Operations
A trustworthy GenAI assistant that helps bank employees find accurate answers
from complex operational documents — where every response is cited, evaluated and safe to act on.
Role
Lead Product & Strategy Designer
Domain
Banking · Enterprise AI · Knowledge Management
Users
Branch & Banking Operations employees
Approach
Discovery · Design Thinking · UX Strategy · Prototyping · Validation
Product
Internal GenAI Knowledge Assistant
Core AI pattern
RAG — Retrieval-Augmented Generation
Knowledge Base
PDFs · SOPs · circulars · scanned documents · tables
Note on this case study: Due to NDA and confidentiality, some details have been changed. The numbers shown are realistic estimates, not actual production data. This case study is intended to show the problem, my design approach, and the potential impact of the solution.
Success criteria — Outcomes
01
Business context
Employees didn't have an information
problem. They had a retrieval and trust problem
Banking operations depend heavily on internal knowledge. An employee
dealing with an account closure, KYC requirement, transaction issue or
operational exception may need to consult multiple internal documents before
taking action.
Traditional enterprise search helped employees find documents, but
employees still had to find and interpret the answer themselves.
Policies
SOPs
Circulars
Product documents
FAQs
Scanned PDFs
Tables
The existing experience
01
Question
02
Search internal systems
03
Try different keywords
04
Open several documents
05
Search within PDFs
06
Interpret policy language
07
Cross-check information
08
Take action
For routine questions, this could turn a simple information request into
several minutes of searching and verification.
Business consequence
This wasn't only an inconvenience. Slow or incorrect
information could contribute to:
longer customer handling times
dependency on experienced employees
repeated queries to support teams
inconsistent interpretation of policies
slower operational decisions
compliance and operational risk
02
The initial brief
The request sounded simple: "Can we build a chatbot over our banking knowledge?”
The initial solution direction was a GenAI chatbot capable of answering
employee questions from internal documents. But that framing immediately
raised a more important question.
From
How should we design the chatbot?
To
How should employees retrieve, understand and verify
AI-generated knowledge?
This reframing became the foundation of the product.
03
Problem-framing workshop
Before designing the interface, I aligned
the team on the problem.
I facilitated a structured Design Thinking workshop with business, operations,
technology and AI stakeholders. The goal wasn't to brainstorm chatbot features
— it was to understand where knowledge breaks down, what employees
actually need, what could go wrong with AI, and what the product must prove
before employees can trust it.
01
Understand the business
20 min
02
Map the current experience
30 min
03
Identify pain points
30 min
04
Map users and scenarios
25 min
05
Explore AI opportunities and risks
30 min
06
Prioritise
25 min
07
Define success
20 min
Go deeper — what each workshop stage asked (Coming soon)
04
What we learned
Four findings that changed the product.
Rather than publishing dozens of sticky notes, the workshop compressed into
four insights — each with a direct design implication.
Insight 01
Finding the document wasn't the real job
Search could return relevant documents, but employees still had
to open them and interpret the information.
The real employee goal was: “Give me the answer that applies to
my situation.”
Design implication
Document retrieval
→
Answer retrieval + evidence
Insight 02
Speed without trust had limited value
An AI answer could theoretically reduce search time from
minutes to seconds. But employees were reluctant to act on
information they couldn't verify.
A fast unsupported answer could therefore create more risk than
a slower traditional search.
Design implication
Question → Search → Documents → Read → Verify
→
Question → Answer → Verify source
Go deeper — Why we changed the success measure (Coming soon)
Insight 03
Banking knowledge wasn't clean data
A large part of the knowledge wasn't available as perfectly
structured text. Relevant information could appear inside
scanned PDFs, tables, annexures, circulars, long policy
documents and differently formatted legacy documents.
This meant the UX couldn't assume “one question = one
paragraph from one document.”
Design implication
One question, one paragraph
→
Extraction → retrieval → synthesis → structure
Go deeper — Response format follows information type (Coming soon)
Insight 04
“I don't know” was a valid product response
Traditional conversational products often optimise towards
always providing a response. That behaviour was inappropriate
here.
If the knowledge base doesn't contain sufficient evidence, the
model could produce a plausible answer using general model
knowledge. In banking operations, plausible isn't sufficient.
Design implication
Always answer
→
No evidence → no generated answer
Go deeper — Refusal as a safety feature (Coming soon)
05
From insights to the actual problem
Three tensions defined the product: Speed, trust, safety.
Speed
Employees need information quickly.
Trust
Employees need evidence before
acting.
Safety
The system shouldn't answer beyond
what it knows.
The opportunity wasn't
Build faster search using GenAI.
It became
Design a knowledge experience that gives employees
fast answers while preserving the ability to verify where
those answers came from — and safely withholding
answers when supporting evidence isn't available.
The design challenge
06
Who should we design for?
One behavioural persona, not five job
titles.
The potential audience included branch employees, operations managers,
relationship managers, service teams and compliance teams. Their job titles
were different, but many of their information behaviours overlapped. Creating
five separate personas would have produced artificial complexity without
materially changing the core interaction.
Instead, I prioritised personas against query frequency × urgency × consequence of incorrect information.
User
Frequency
Urgency
Wrong-answer risk
Priority
Branch Operations Manager
High
High
High
Primary
Operations Specialist
Medium
High
High
Secondary
Relationship Manager
Medium
Medium
Medium
Later
Teller
Medium
High
Medium
Later
Compliance
Low
Medium
Very high
Stakeholder
Primary user
Branch Operations Manager
The role combined frequent information needs, time pressure
and significant consequences from incorrect information.
Job to be done
When I need to make an operational decision, help me
quickly find the policy or procedure that applies to my
situation, so I can confidently take the correct action
without searching through multiple documents.
Find
Get relevant information quickly.
Understand
Convert complex policy language into something actionable.
Verify
Understand where the answer came from.
Trust
Know whether sufficient evidence supports the response.
Recover
Know what to do when the system cannot answer.
Their needs
07
Mapping the journey
Collapse search, find, understand and
interpret — but keep verify visible.
I mapped the existing knowledge journey to identify where the experience was
failing. The distinction mattered: we wanted AI to remove unnecessary work, not
remove employee oversight.
Stage
Employee behaviour
Friction
Opportunity
Recognise need
Encounters operational question
Often urgent
Capture natural-language intent
Search
Searches portal / knowledge base
Keyword dependency
Semantic retrieval
Find
Opens multiple documents
Too many results
Retrieve relevant evidence
Understand
Reads policies
Long, complex language
Summarisation
Interpret
Connects information
Cognitive load
Structured answers
Verify
Cross-checks documents
Time-consuming
Source attribution
Act
Makes decision
Uncertainty remains
Confidence & evidence
Escalate
Contacts support
Additional delay
Guided escalation
The target experience
01
Ask
Employee asks naturally.
02
Retrieve
System searches only approved
enterprise knowledge.
03
Understand
Information is extracted from text,
scans and tables.
04
Evaluate
Retrieved context and generated
response are evaluated.
05
Answer
Information is presented concisely.
06
Verify
Sources remain immediately
accessible.
07
Act
Employee proceeds confidently.
08
Feedback
Feedback improves future quality.
Search → Read → Interpret → Verify
became
Ask → Understand → Verify
08
Experience principles
Five principles instead of a feature
checklist.
01
Answer first
Employees shouldn't need to read several paragraphs to
understand the response.
UX response
Direct answer first → supporting detail second.
02
Evidence over authority
The assistant shouldn't expect users to trust it because “AI said
so.”
UX response
Every material answer exposes its supporting source.
03
Make uncertainty visible
AI confidence shouldn't be hidden inside the technical system.
UX response
High confidence — strong supporting evidence.
Review recommended — relevant but incomplete
evidence.
Unable to verify — insufficient supporting evidence.
04
Structure information around the task
Response format should follow information type rather than
forcing everything into chat paragraphs.
UX response
Process → steps
Requirements → bullets
Comparison → table
Long content → summary
Critical exception → warning
05
Safe failure over plausible answers
A believable incorrect answer is worse than an explicit limitation.
UX response
“I couldn't find enough verified information to answer this
safely.”
Then provide relevant next actions.
Go deeper — the response hierarchy behind every answer
+
09
Product decisions
What entered the first release — and
what trust required.
Not every AI capability needed to enter the first release. I prioritised capabilities
according to user value × trust impact × risk reduction × technical feasibility.
P0
Trust foundation
natural-language questions
RAG-based knowledge retrieval
concise answers
source attribution
source preview
scanned-document extraction
table understanding
no-answer guardrail
P1
Experience maturity
contextual follow-up
confidence states
feedback
structured summaries
conflicting-source handling
P2
Future capability
proactive recommendations
personalised knowledge
saved queries
trending questions
advanced analytics
The trust model
Trust became a product architecture decision — four questions the experience needed to answer for
every employee.
1
Did you understand me?
Context relevance
2
Did you find the right
information?
Retrieval quality
3
Is your answer supported by
that information?
Faithfulness
4
Can I verify it myself?
Source attribution
Relevant context + faithful answer + visible evidence + safe uncertainty = trustworthy AI experience
Go deeper — why the failure state mattered most
+
10
The experience
Ten product states, one coherent
system.
The case study doesn't need dozens of screens — it needs the states that prove
the product thinking. A hiring manager should see one coherent product
system, not ten unrelated chatbot screens.
01
Ask
02
Answer
03
Trust
04
Verify
05
Complex knowledge
06
Structure
07
Context
08
Uncertainty
09
Conflict
10
Feedback
01
Ask
Knowledge Assistant · Banking Operations
What documents do we need to close a deceased
customer's savings account?
Searching approved enterprise knowledge only.
Policies
SOPs
Circulars
Product documents
← Previous
1 / 10
Next →
11
AI governance
Model evaluation became an
interaction-design input.
Governance wasn't treated as something that happened only behind the
scenes. Evaluation directly influenced the experience employees received.
Evaluation
Question
Context relevance
Did we retrieve information relevant to the employee's question?
Faithfulness
Is the generated answer supported by retrieved evidence?
Answer relevance
Does the response actually answer the question?
Hallucination
Did the model introduce unsupported claims?
Evaluation → UX behaviour
Relevant + supported
Answer normally
Show concise answer + evidence + sources.
Relevant but ambiguous
Review recommended
Surface uncertainty and supporting sources.
Conflicting evidence
Don't hide the conflict
Present both sources and encourage verification.
Insufficient evidence
Don't answer
Explain the limitation and provide escalation.
Poor context relevance
Clarify
Ask the employee to refine the question.
Closing the feedback loop
User feedback wasn't treated simply as 👍 / 👎. Negative feedback was
categorised, which distinguished retrieval problems from generation
problems and content gaps — each requiring a different product or
technical response.
Incorrect answer
Missing information
Source not relevant
Answer unclear
Other
12
How we defined success
Success wasn't “the chatbot
answered.”
We defined success across efficiency, trust, quality, safety and adoption. The
figures below are targets and success criteria set during the initial evaluation —
not claimed outcomes.
Dimension
Metric
Target
Efficiency
Time to verified answer
< 1 min
Self-service
Queries resolved without escalation
↑
Trust
Answers with traceable evidence
100%
Usefulness
Helpful-response rate
↑
Quality
Answer relevance
70–75%+
Quality
Context relevance
70–75%+
Safety
Unsupported generated answers
Zero
Adoption
Repeat usage
↑
Operations
Support dependency
↓
Business impact model
Faster retrieval
↓ search time
Faster employee decisions
↓ handling time
Higher self-service
↓ dependency on support teams
Consistent information
↓ operational variation
Grounded answers
↓ risk of unsupported decisions
13
My contribution & reflection
From knowledge search to
trusted decision
support.
The final concept transformed the employee experience from “find the right
document and figure out the answer” to “ask a question, understand the answer
and verify the evidence.” When sufficient supporting knowledge wasn't
available, it didn't invent.
My role
As the lead designer, I worked across problem framing, product
definition and experience design:
facilitating discovery and problem-framing workshops
mapping the existing knowledge-retrieval journey
identifying and prioritising primary users
translating business and AI requirements into experience
principles
defining trust and verification patterns
defining answer, uncertainty and failure states
collaborating with AI/engineering teams around RAG
behaviour
connecting AI evaluation metrics with user-facing behaviours
defining feedback mechanisms
designing and validating the end-to-end experience
establishing UX success criteria
My role wasn't to decide how the model worked
internally; it was to ensure the capabilities and
limitations of the AI were translated into
understandable, safe and useful employee
experiences.
Evidence before confidence
Users trust what they can verify.
Uncertainty needs design
“No reliable answer found” is a legitimate product state.
AI governance affects UX
Faithfulness, context relevance and hallucination determine what
the product should show, hide, warn about or refuse to answer.
Key product decisions
Instead of
I chose
Because
Multiple personas
One primary behavioural persona
Core information needs overlapped
Document search
Answer + evidence
Employees needed information, not files
Long AI responses
Structured responses
Reduced cognitive load
Hidden citations
Visible evidence
Trust required verification
Raw confidence %
Understandable confidence states
Scores lacked actionable meaning
AI always answering
Explicit refusal
Reduced unsupported responses
Generic error
Guided escalation
Preserved task continuity
👍 / 👎 only
Categorised feedback
Identified quality failure type
Governance after launch
Governance within UX
AI quality directly affects user trust
In high-stakes enterprise environments, trust comes from transparency and predictable behaviour — not from how intelligent the interface appears. The result was not simply a chatbot interface. It was a framework for responsible knowledge retrieval in a regulated environment.