Introduction
Controllers at GS review large volumes of data quality exceptions before signing off on the accuracy of different financial datasets. The review process was almost entirely manual: reviewers investigated discrepancies one by one and decided on remediation measures. I led the design of an AI-driven platform that changes this model, where AI investigates exceptions proactively and controllers review, adjust, and approve its suggestions.
My role
Lead UI/UX Designer on the engagement, working with the working group lead, AI engineering lead and SMEs. I owned the end-to-end design from problem definition through POC and MVP, ran stakeholder and user group walkthroughs with the different user groups in the firm to pressure-test decisions against real workflows.
Problem statement
Each close cycle, a single reporting deliverable could surface hundreds of exceptions, and resolving them took 4 to 7 hours of manual investigation. Around 70% of exceptions traced back to recurring root causes, yet reviewers still worked through them individually. Beyond the time cost, there was a higher bar to clear: controllers personally attest to the accuracy of financial data, so any AI assistance had to be trustworthy enough to act on and auditable enough to defend. The brief: design a review experience where AI does the investigation while the controller stays firmly in control of the sign-off.
Discovery through POC and wireframes
The engagement started with an introductory scope from the project lead. Rather than a long research phase upfront, we built a proof of concept early and used it to anchor discussions with reporting SMEs, user representatives, and senior stakeholders. Reacting to something concrete made it easier to align on scope and validate the need.
Once the POC was approved, I moved into wireframes and iterated within the working group. Sessions with SMEs and walkthroughs of real exception reports showed what the data looked like, how exceptions were categorized, and the volumes reviewers handled in a cycle. The goal was wireframes solid enough to take to key user groups. The findings shaped the design decisions:
• Exceptions trace back to recurring root causes, and reviewers already think in clusters, so group-level resolution matched their reasoning.
• Trust takes more than a confidence score. Controllers wanted the reasoning, materiality, and source before acting.
• The three user groups shared a core workflow but differed in data and cadence, pointing to one adaptable system.

1. POC demo build
Early POC screen used to demo the use case to senior stakeholders. Not everything followed the right conventions yet, but it proved the concept and got the go-ahead.

2. Initial working wireframes
Informed by sessions on users' current tooling: AI suggestions surfaced as current vs. suggested value, grouped by action type, with confidence and impact.

3. Later wireframes
Sharpened the overview datapoints, fine-tuned suggestion grouping, and introduced the AI rationale side panel with reasoning and sources.

UX objectives
What the design had to achieve, distilled from discovery and stakeholder sessions.​​​​​​​
Validating across user groups and iterations
With the objectives set, we took the mockups to the key user groups: Revenue, Risk, Regulatory Reporting, and others. Each round of walkthroughs surfaced where their workflows differed and the types of exceptions each group resolved. The feedback fed back into the designs iteratively, and the goal shifted from a single screen that worked to a UX framework that fit all user groups.

The user sessions drove concrete changes. Risk's workflow, for example, includes deciding whether an exception is blocking or non-blocking. That became an AI suggestion for them. Sessions also sharpened how sources are shown: the type, recency, and liveness of a source directly affected users' confidence to adopt a suggestion.
Key Design Decisions
Group-level actions. Exceptions are actioned in clusters with a shared rationale. Drill-down allows overrides on individual rows. Following mental modal of the reviewers.
Hybrid interaction. Action buttons are the primary surface, with a context-scoped chat alongside. This settled the chat-versus-buttons debate in the working group. Reviewers can make use of the AI chat to solve for edge cases.
Attention-based quick filters. The AI scores each exception on confidence and materiality, then groups them into High, Medium, and Low Attention. These sit above the table as quick filters. The default view stays unfiltered so controllers see the full set before narrowing.
Rejection as a forward path. Rejecting a suggestion opens Add Context & Revalidate. Controllers can attach uploads, sources, or free text. The AI re-runs scoped to that group.
Extensible action vocabulary. Action types adapt per user group across Revenue, Risk, and Reporting. The IA stays the same underneath. New groups onboard without structural changes.
Impact
The engagement concluded at MVP handoff, before production release, so the outcomes below separate what was delivered from what the design was built to achieve.
Delivered
Approved and taken to MVP. End-to-end experience validated with three key user groups and signed off by key stakeholders.
One framework, three workflows. A single adaptable system across Revenue, Risk, and Regulatory Reporting, with an extensible action vocabulary for onboarding new groups.
Auditable human-in-the-loop. Suggest-and-confirm with inline reasoning and cited sources, meeting the trust bar manual review was built around.
Designed to achieve
4 to 7 hours to under an hour per cycle. The validated design target, not a measured production result, with the architecture built toward minutes as AI confidence coverage grows.
The root-cause 70%. Group-level resolution aimed at the recurring clusters behind most exceptions.

You may also like

Back to Top