Back to home

Banking

Conversational AI

Enterprise UX

AI Knowledge Assistant for

Banking Operations

A trustworthy GenAI assistant that helps bank employees find accurate answers

from complex operational documents — where every response is cited, evaluated and safe to act on.


Role

Lead Product & Strategy Designer

Domain

Banking · Enterprise AI · Knowledge Management

Users

Branch & Banking Operations employees

Approach

Discovery · Design Thinking · UX Strategy · Prototyping · Validation

Product

Internal GenAI Knowledge Assistant

Core AI pattern

RAG — Retrieval-Augmented Generation

Knowledge Base

PDFs · SOPs · circulars · scanned documents · tables

Note on this case study: Due to NDA and confidentiality, some details have been changed. The numbers shown are realistic estimates, not actual production data. This case study is intended to show the problem, my design approach, and the potential impact of the solution.

Success criteria — Outcomes

< 1 min

target time to find and verify an answer

< 1 min

target time to find and verify an answer

70–75%+

target answer accuracy during initial evaluation

70–75%+

target answer accuracy during initial evaluation

100%

of material answers traceable to retrieved sources

100%

of material answers traceable to retrieved sources

0

unsupported answers by designwhen evidence is missing

0

unsupported answers by designwhen evidence is missing

01

Business context

Employees didn't have an information

problem. They had a retrieval and trust problem

Banking operations depend heavily on internal knowledge. An employee

dealing with an account closure, KYC requirement, transaction issue or

operational exception may need to consult multiple internal documents before

taking action.

Traditional enterprise search helped employees find documents, but

employees still had to find and interpret the answer themselves.

Policies

SOPs

Circulars

Product documents

FAQs

Scanned PDFs

Tables

The existing experience

01

Question

02

Search internal systems

03

Try different keywords

04

Open several documents

05

Search within PDFs

06

Interpret policy language

07

Cross-check information

08

Take action

For routine questions, this could turn a simple information request into

several minutes of searching and verification.

Business consequence

This wasn't only an inconvenience. Slow or incorrect

information could contribute to:

longer customer handling times

dependency on experienced employees

repeated queries to support teams

inconsistent interpretation of policies

slower operational decisions

compliance and operational risk

02

The initial brief

The request sounded simple: "Can we build a chatbot over our banking knowledge?”

The initial solution direction was a GenAI chatbot capable of answering

employee questions from internal documents. But that framing immediately

raised a more important question.

If an employee is expected to act on an AI-generated

answer, what would make that answer trustworthy enough

to use?

If an employee is expected to act on an AI-generated

answer, what would make that answer trustworthy enough to use?

From

How should we design the chatbot?

To

How should employees retrieve, understand and verify

AI-generated knowledge?

This reframing became the foundation of the product.

03

Problem-framing workshop

Before designing the interface, I aligned

the team on the problem.

I facilitated a structured Design Thinking workshop with business, operations,

technology and AI stakeholders. The goal wasn't to brainstorm chatbot features

— it was to understand where knowledge breaks down, what employees

actually need, what could go wrong with AI, and what the product must prove

before employees can trust it.

01

Understand the business

20 min

02

Map the current experience

30 min

03

Identify pain points

30 min

04

Map users and scenarios

25 min

05

Explore AI opportunities and risks

30 min

06

Prioritise

25 min

07

Define success

20 min

Go deeper — what each workshop stage asked (Coming soon)

04

What we learned

Four findings that changed the product.

Rather than publishing dozens of sticky notes, the workshop compressed into

four insights — each with a direct design implication.

Insight 01

Finding the document wasn't the real job

Search could return relevant documents, but employees still had

to open them and interpret the information.

The real employee goal was: “Give me the answer that applies to

my situation.”

Design implication

Document retrieval

Answer retrieval + evidence

Insight 02

Speed without trust had limited value

An AI answer could theoretically reduce search time from

minutes to seconds. But employees were reluctant to act on

information they couldn't verify.

A fast unsupported answer could therefore create more risk than

a slower traditional search.

Design implication

Question → Search → Documents → Read → Verify

Question → Answer → Verify source

Go deeper — Why we changed the success measure (Coming soon)

Insight 03

Banking knowledge wasn't clean data

A large part of the knowledge wasn't available as perfectly

structured text. Relevant information could appear inside

scanned PDFs, tables, annexures, circulars, long policy

documents and differently formatted legacy documents.

This meant the UX couldn't assume “one question = one

paragraph from one document.”

Design implication

One question, one paragraph

Extraction → retrieval → synthesis → structure

Go deeper — Response format follows information type (Coming soon)

Insight 04

“I don't know” was a valid product response

Traditional conversational products often optimise towards

always providing a response. That behaviour was inappropriate

here.

If the knowledge base doesn't contain sufficient evidence, the

model could produce a plausible answer using general model

knowledge. In banking operations, plausible isn't sufficient.

Design implication

Always answer

No evidence → no generated answer

Go deeper — Refusal as a safety feature (Coming soon)

05

From insights to the actual problem

Three tensions defined the product: Speed, trust, safety.

Speed

Employees need information quickly.

Trust

Employees need evidence before

acting.

Safety

The system shouldn't answer beyond

what it knows.

The opportunity wasn't

Build faster search using GenAI.

It became

Design a knowledge experience that gives employees

fast answers while preserving the ability to verify where

those answers came from — and safely withholding

answers when supporting evidence isn't available.

The design challenge

How might we help banking employees find and

understand operational information in seconds, while

ensuring every generated answer is grounded,

verifiable and safe to act on?

How might we help banking employees find and understand operational information in seconds, while ensuring every generated answer is grounded, verifiable and safe to act on?

06

Who should we design for?

One behavioural persona, not five job

titles.

The potential audience included branch employees, operations managers,

relationship managers, service teams and compliance teams. Their job titles

were different, but many of their information behaviours overlapped. Creating

five separate personas would have produced artificial complexity without

materially changing the core interaction.

Instead, I prioritised personas against query frequency × urgency × consequence of incorrect information.

User

Frequency

Urgency

Wrong-answer risk

Priority

Branch Operations Manager

High

High

High

Primary

Operations Specialist

Medium

High

High

Secondary

Relationship Manager

Medium

Medium

Medium

Later

Teller

Medium

High

Medium

Later

Compliance

Low

Medium

Very high

Stakeholder

Primary user

Branch Operations Manager

The role combined frequent information needs, time pressure

and significant consequences from incorrect information.

Job to be done

When I need to make an operational decision, help me

quickly find the policy or procedure that applies to my

situation, so I can confidently take the correct action

without searching through multiple documents.

Find

Get relevant information quickly.

Understand

Convert complex policy language into something actionable.

Verify

Understand where the answer came from.

Trust

Know whether sufficient evidence supports the response.

Recover

Know what to do when the system cannot answer.

Their needs

07

Mapping the journey

Collapse search, find, understand and

interpret — but keep verify visible.

I mapped the existing knowledge journey to identify where the experience was

failing. The distinction mattered: we wanted AI to remove unnecessary work, not

remove employee oversight.

Stage

Employee behaviour

Friction

Opportunity

Recognise need

Encounters operational question

Often urgent

Capture natural-language intent

Search

Searches portal / knowledge base

Keyword dependency

Semantic retrieval

Find

Opens multiple documents

Too many results

Retrieve relevant evidence

Understand

Reads policies

Long, complex language

Summarisation

Interpret

Connects information

Cognitive load

Structured answers

Verify

Cross-checks documents

Time-consuming

Source attribution

Act

Makes decision

Uncertainty remains

Confidence & evidence

Escalate

Contacts support

Additional delay

Guided escalation

The target experience

01

Ask

Employee asks naturally.

02

Retrieve

System searches only approved

enterprise knowledge.

03

Understand

Information is extracted from text,

scans and tables.

04

Evaluate

Retrieved context and generated

response are evaluated.

05

Answer

Information is presented concisely.

06

Verify

Sources remain immediately

accessible.

07

Act

Employee proceeds confidently.

08

Feedback

Feedback improves future quality.

Search → Read → Interpret → Verify

became

Ask → Understand → Verify

08

Experience principles

Five principles instead of a feature

checklist.

01

Answer first

Employees shouldn't need to read several paragraphs to

understand the response.

UX response

Direct answer first → supporting detail second.

02

Evidence over authority

The assistant shouldn't expect users to trust it because “AI said

so.”

UX response

Every material answer exposes its supporting source.

03

Make uncertainty visible

AI confidence shouldn't be hidden inside the technical system.

UX response

High confidence — strong supporting evidence.

Review recommended — relevant but incomplete

evidence.

Unable to verify — insufficient supporting evidence.

04

Structure information around the task

Response format should follow information type rather than

forcing everything into chat paragraphs.

UX response

Process → steps

Requirements → bullets

Comparison → table

Long content → summary

Critical exception → warning

05

Safe failure over plausible answers

A believable incorrect answer is worse than an explicit limitation.

UX response

“I couldn't find enough verified information to answer this

safely.”

Then provide relevant next actions.

Go deeper — the response hierarchy behind every answer

+

09

Product decisions

What entered the first release — and

what trust required.

Not every AI capability needed to enter the first release. I prioritised capabilities

according to user value × trust impact × risk reduction × technical feasibility.

P0

Trust foundation

natural-language questions

RAG-based knowledge retrieval

concise answers

source attribution

source preview

scanned-document extraction

table understanding

no-answer guardrail

P1

Experience maturity

contextual follow-up

confidence states

feedback

structured summaries

conflicting-source handling

P2

Future capability

proactive recommendations

personalised knowledge

saved queries

trending questions

advanced analytics

The trust model

Trust became a product architecture decision — four questions the experience needed to answer for

every employee.

1

Did you understand me?

Context relevance

2

Did you find the right

information?

Retrieval quality

3

Is your answer supported by

that information?

Faithfulness

4

Can I verify it myself?

Source attribution

Relevant context + faithful answer + visible evidence + safe uncertainty = trustworthy AI experience

Go deeper — why the failure state mattered most

+

10

The experience

Ten product states, one coherent

system.

The case study doesn't need dozens of screens — it needs the states that prove

the product thinking. A hiring manager should see one coherent product

system, not ten unrelated chatbot screens.

01

Ask

02

Answer

03

Trust

04

Verify

05

Complex knowledge

06

Structure

07

Context

08

Uncertainty

09

Conflict

10

Feedback

01

Ask

Natural-language knowledge retrieval. No keyword syntax,

no filters to configure.

Natural-language knowledge retrieval. No keyword syntax, no filters to configure.

Knowledge Assistant · Banking Operations

What documents do we need to close a deceased

customer's savings account?

Searching approved enterprise knowledge only.

Policies

SOPs

Circulars

Product documents

← Previous

1 / 10

Next →

11

AI governance

Model evaluation became an

interaction-design input.

Governance wasn't treated as something that happened only behind the

scenes. Evaluation directly influenced the experience employees received.

Evaluation

Question

Context relevance

Did we retrieve information relevant to the employee's question?

Faithfulness

Is the generated answer supported by retrieved evidence?

Answer relevance

Does the response actually answer the question?

Hallucination

Did the model introduce unsupported claims?

Evaluation → UX behaviour

Relevant + supported

Answer normally

Show concise answer + evidence + sources.

Relevant but ambiguous

Review recommended

Surface uncertainty and supporting sources.

Conflicting evidence

Don't hide the conflict

Present both sources and encourage verification.

Insufficient evidence

Don't answer

Explain the limitation and provide escalation.

Poor context relevance

Clarify

Ask the employee to refine the question.

Closing the feedback loop

User feedback wasn't treated simply as 👍 / 👎. Negative feedback was

categorised, which distinguished retrieval problems from generation

problems and content gaps — each requiring a different product or

technical response.

Incorrect answer

Missing information

Source not relevant

Answer unclear

Other

12

How we defined success

Success wasn't “the chatbot

answered.”

We defined success across efficiency, trust, quality, safety and adoption. The

figures below are targets and success criteria set during the initial evaluation —

not claimed outcomes.

Dimension

Metric

Target

Efficiency

Time to verified answer

< 1 min

Self-service

Queries resolved without escalation

Trust

Answers with traceable evidence

100%

Usefulness

Helpful-response rate

Quality

Answer relevance

70–75%+

Quality

Context relevance

70–75%+

Safety

Unsupported generated answers

Zero

Adoption

Repeat usage

Operations

Support dependency

Business impact model

Faster retrieval

↓ search time

Faster employee decisions

↓ handling time

Higher self-service

↓ dependency on support teams

Consistent information

↓ operational variation

Grounded answers

↓ risk of unsupported decisions

The business value isn't “employees liked the chatbot.” It is that faster access to verified knowledge

can improve employee productivity while reducing operational dependency and the risk of

inconsistent information.

The business value isn't “employees liked the chatbot.” It is that faster access to verified knowledge can improve employee productivity while reducing operational dependency and the risk of inconsistent information.

13

My contribution & reflection

From knowledge search to

trusted decision

support.

The final concept transformed the employee experience from “find the right

document and figure out the answer” to “ask a question, understand the answer

and verify the evidence.” When sufficient supporting knowledge wasn't

available, it didn't invent.

My role

As the lead designer, I worked across problem framing, product

definition and experience design:

facilitating discovery and problem-framing workshops

mapping the existing knowledge-retrieval journey

identifying and prioritising primary users

translating business and AI requirements into experience

principles

defining trust and verification patterns

defining answer, uncertainty and failure states

collaborating with AI/engineering teams around RAG

behaviour

connecting AI evaluation metrics with user-facing behaviours

defining feedback mechanisms

designing and validating the end-to-end experience

establishing UX success criteria

My role wasn't to decide how the model worked

internally; it was to ensure the capabilities and

limitations of the AI were translated into

understandable, safe and useful employee

experiences.

Evidence before confidence

Users trust what they can verify.

Uncertainty needs design

“No reliable answer found” is a legitimate product state.

AI governance affects UX

Faithfulness, context relevance and hallucination determine what

the product should show, hide, warn about or refuse to answer.

Key product decisions

Instead of

I chose

Because

Multiple personas

One primary behavioural persona

Core information needs overlapped

Document search

Answer + evidence

Employees needed information, not files

Long AI responses

Structured responses

Reduced cognitive load

Hidden citations

Visible evidence

Trust required verification

Raw confidence %

Understandable confidence states

Scores lacked actionable meaning

AI always answering

Explicit refusal

Reduced unsupported responses

Generic error

Guided escalation

Preserved task continuity

👍 / 👎 only

Categorised feedback

Identified quality failure type

Governance after launch

Governance within UX

AI quality directly affects user trust

In high-stakes enterprise environments, trust comes from transparency and predictable behaviour — not from how intelligent the interface appears. The result was not simply a chatbot interface. It was a framework for responsible knowledge retrieval in a regulated environment.