Introduction
During financial close, controllers at GS review large volumes of data quality exceptions before signing off on the accuracy of the datasets. The review process was almost entirely manual: reviewers investigated discrepancies one by one and decided on remediation measures, then documented their findings for audit. I led the design of an AI-driven platform that changes this model, where AI investigates exceptions proactively and controllers review, adjust, and approve its suggestions.
My role
Lead UI/UX Designer on the engagement, working with the working group lead, AI engineering lead, SMEs, and the internal UX lead. I owned the end-to-end design from problem definition through POC and MVP, including the information architecture, the interaction model for reviewing AI suggestions, and alignment with the bank's internal design system. I ran stakeholder and user group walkthroughs with controllers across Revenue, Risk, and Regulatory Reporting to pressure-test decisions against real workflows.
Problem statement
Each close cycle, a single reporting deliverable could surface hundreds of exceptions, and resolving them took 4 to 7 hours of manual investigation. Around 70% of exceptions traced back to recurring root causes, yet reviewers still worked through them individually. Beyond the time cost, there was a higher bar to clear: controllers personally attest to the accuracy of financial data, so any AI assistance had to be trustworthy enough to act on and auditable enough to defend. The brief: design a review experience where AI does the investigation while the controller stays firmly in control of the sign-off.
Discovery through POC and wireframes
The engagement started with an introductory project scope put together by the project lead. Rather than a long upfront research phase, we built a proof of concept early and used it as the anchor for discussions with reporting SMEs, user representatives, and senior stakeholders. Reacting to something concrete made it far easier for stakeholders to align on the scope and validate the need itself than an abstract brief would have.
Once the POC got the go-ahead, I moved into wireframes, iterating within the working group and deepening my understanding of the domain and the users with every session. Working sessions with SMEs and walkthroughs of real exception reports revealed what the data actually looked like, how exceptions were categorized, and what volumes reviewers handled in a typical cycle.
The goal of this phase was to get the wireframes to a point where they could be taken to key user groups. Along the way, the findings shaped the design decisions:
• Controllers personally attest to the accuracy of a dataset, so they need to see the full picture. Anything that hides part of the workload works against that instinct.
• Most exceptions trace back to recurring root causes, and reviewers already think in those clusters. Group-level resolution matched their reasoning.
• Trust in an AI suggestion takes more than a confidence score. Controllers wanted the reasoning, the materiality, and the data source before acting.
• The three user groups shared a core workflow but differed in data and cadence, pointing to one adaptable system over three tools.
1. POC demo build
Early POC screen used to demo the use case to senior stakeholders. Not everything followed the right conventions yet, but it proved the concept and got the go-ahead.
2. Initial working wireframes
Informed by sessions on users' current tooling: AI suggestions surfaced as current vs. suggested value, grouped by action type, with confidence and impact.
3. Later wireframes
Sharpened the overview datapoints, fine-tuned suggestion grouping, and introduced the AI rationale side panel with reasoning and sources.
UX objectives
What the design had to achieve, distilled from discovery and stakeholder sessions.
Validating across user groups and iterations
With the objectives set, we took the mockups to the key user groups: Revenue, Risk, Regulatory Reporting, and others. Each round of walkthroughs surfaced where their workflows differed and the types of exceptions each group resolved. The feedback fed back into the designs iteratively, and the goal shifted from a single screen that worked to a UX framework that fit all user groups.
The user sessions drove concrete changes. Risk's workflow, for example, includes deciding whether an exception is blocking or non-blocking. That became an AI suggestion for them. Sessions also sharpened how sources are shown: the type, recency, and liveness of a source directly affected users' confidence to adopt a suggestion.