← Back to projects Product Design Case Study

Trustworthy AI Feedback

An AI tool flagged bad audit evidence without explaining why, so frustrated customers turned to overloaded auditors for help. I redesigned it to explain itself, as lead designer, cutting audit time by roughly 12%.

Role Lead Product Designer
Timeline 2 months, Q2 2025
Team 1 PM, 6 engineers, 1 other designer
Tools Figma, Claude Design, user interviews, usability testing
The redesigned evidence request modal, showing Verify Assist Results that flag an incomplete evidence upload and a discrepancy in submitted evidence
TL;DR

Imagine submitting evidence for a complicated audit and an AI tool tells you "something's missing" without saying what. Get it wrong and you could fail the audit, so you call your auditor for help, and your auditor is already juggling many other customers and audits at once. That was the daily reality for our customers and auditors. I redesigned the tool so it actually explains itself: what's missing, and how to fix it, before anything is submitted. Audit times are already down about 12% since launch, and customers finally know what they need.

The Problem

An AI tool that flagged problems but never explained them

What was broken

The company was a fast-paced startup with an AI tool that checked audit evidence before it reached our auditors. The problem: it often got things wrong, and even when it was right, it never explained what was actually missing. Customers had no way to fix the gap themselves, so they'd end up going back and forth with auditors by email or phone. That slowed audits down and cost the company money.

Starting from behind

I walked into this project with two strikes against me. The customer success team didn't want to give me access to customers because of bad experiences with past design projects. And the auditors didn't trust anyone from Product, because of tension between their team and ours before I ever joined. I had to earn both relationships before I could even start designing.

Constraints

  • Two-month timeline (Q2 2025) with a lean team: 1 PM, 6 engineers, and 1 other designer.
  • No customer access at the start: customer success withheld it due to friction from earlier design projects.
  • No auditor trust at the start: preexisting tension between the auditor team and Product meant beginning from zero credibility.
Original evidence request modal showing a single vague flag: Some of the evidence provided doesn't fully satisfy this request yet

Before

Verify Assist Results only surfaced a general message that something was missing, with no indication of what or where.

Redesigned evidence request modal showing itemized Verify Assist Results that explain exactly what's missing and why

After

Verify Assist now has context on which systems are in scope for the customer's audit, so it can give accurate, specific results on what's missing, system by system.

Approach

Earning trust before earning access

Building trust through consistency

I worked hard to build trust by actively listening and asking more questions before showing solutions. Every time someone gave me feedback, I brought back a design that clearly showed I'd listened, and checked in between meetings, not just during them. Once a few auditors trusted me, that trust opened doors with the rest of the team. I used the same approach with customers: instead of asking every customer success manager for help, I built one strong relationship, showed him he could trust me not to waste his or his customers' time, and he introduced me to several of his customers.

Finding common ground between auditors and leadership

Leadership was focused on Verify Assist and making it better. Auditors didn't care about Verify Assist at all: it had shown no real value before, and they didn't trust the AI to help them or their customers. In my meetings with auditors, I worked on understanding what they did care about and where their pain points actually were, so I could find the common ground between what they wanted and what leadership wanted and design one solution that made (almost) everyone happy. That common ground was simple: both sides wanted audits to be faster and easier. Once I found it, I could show auditors through the designs themselves that my focus was making their lives easier, and that despite their past experience with Verify Assist, we wouldn't move forward with a solution that didn't actually lessen their pain.

Moving fast with AI-assisted design

I used Claude Design alongside our company's design system to iterate quickly. When someone gave me feedback, I could show up the next day with a new version instead of making them wait a week. That speed was a big part of why skeptical auditors started trusting me: it showed I was actually listening and acting, fast.

Talking to the people actually using the tool

I interviewed auditors during a meeting they already had every week, so I wasn't asking for extra time. I talked to customers and ran usability tests with them through the one relationship I'd built. And I looked at real data: how often the tool got things wrong, and what it was telling people when it did.

Key insights

  1. Customers didn't need a diagnosis, they needed to know exactly what to go find.
  2. Auditors already reviewed every piece of evidence by hand, so a detailed customer-style view was noise to them, not help.
  3. Customers with multiple products got confused by one unified view, since they usually only understood their own product's requirements.
  4. Customers were willing to edit the system mapping themselves when they thought it was wrong, and getting them to confirm that mapping before running Verify Assist turned out to be key to the whole feature's success.

Mapping the flow before designing

Before jumping into real designs, I mapped out the full user flow, every step from a customer uploading evidence to an auditor reviewing it, and walked it with stakeholders to confirm we agreed on what we were actually building. That flow became the reference point the team kept coming back to whenever priorities got fuzzy, and it's what's mapped out below.

The Solution

A tool that explains itself

Now, when someone uploads evidence, the AI figures out exactly what part of the audit it belongs to, instead of just saying "something's missing" with no explanation. If something doesn't match the audit's plan, both the customer and the auditor get a heads-up before it's even submitted, so problems get caught early instead of after the fact.

  • Automatic matching: evidence is matched to the relevant system upon upload.
  • Self-service correction: if the AI gets something wrong, customers can fix it themselves in a couple of clicks.
  • Early mismatch alerts: both customer and auditor are notified when submitted evidence doesn't match the audit's plan, before it's even submitted.
  • Itemized checklist: based on which systems are determined in-scope for their audit, customers know exactly what's missing, not just that something is.
SOC 2 Evidence Request · CC7.1
Vulnerability Scanning & Remediation — Q2 Sample

Upload vulnerability scan reports for each sampled month, and remediation documentation for any Critical or High findings.

Verify Assist Results
Run on 07/24/2026
Core Application (GitHub)

June 2026 has a High-severity finding (CVE-2026-2094) that still needs a documented remediation plan and its current status (e.g., Remediated, In Progress) to fully satisfy this request.

Corporate Identity Provider

No scan results have been uploaded for this system yet. At least one sampled month is required before this request can be submitted.

Are there discrepancies in my evidence? Critical Issue Identified
Is my evidence outside the review period? No Issues Found
Systems In Scope

Evidence for this control must come from the following systems. Each system needs at least one piece of evidence.

AWS Production Account 3 attachments
Core Application (GitHub) 2 attachments
Corporate Identity Provider 0 attachments
Marketing Website (WordPress) — not in scope 1 attachment
Vulnerability Scan & Remediation Evidence

Provide vulnerability scan results for each sampled month, plus remediation documentation for any Critical or High severity findings (e.g., Remediated, In Progress).

April 2026 Vulnerability Scan Report.pdf 05/03
May 2026 Vulnerability Scan Report.pdf 06/02
June 2026 Vulnerability Scan Report.pdf 07/04
CVE-2026-1188 - Remediation Plan.pdf 05/10
CVE-2026-2094 - Remediation Plan.pdf 06/14
Marketing Site Scan Export.pdf 06/20
Notes (Optional)
Assignee
A
A. Morgan
Status
In Progress 2/2 complete
Evidence Format

Exported documents such as .docx, .csv, or .pdf, or screenshots as .jpg or .png.

Additional Guidance

Depending on how your environment and controls are configured, the evidence you provide may vary. Common scenarios include:

Infrastructure scanning — scans are run against in-scope production infrastructure on a defined cadence (e.g., monthly), with remediation tracked according to internal SLA policy.

Application scanning — scans are run against application code prior to release, with all Critical and High findings resolved before code reaches production.

Related Controls
CC7.1 - Vulnerability Detection CC7.2 - Security Monitoring CC8.1 - Change Management

A working recreation of the redesigned evidence request modal — click "Is this evidence complete?" to expand it.

Key decisions and trade-offs

  • Scoped the first version to single-product customers only, deferring a multi-product filtered view, to keep momentum instead of designing for every case at once.
  • Cut the auditor-facing view down to a single plan-mismatch alert after testing showed auditors ignored a detailed view, since they already review evidence manually.
  • Partnered with another designer to ship the mismatch alert inside a feature she already had in flight, rather than building new UI from scratch.
  • Made AI correction a must-have requirement, not a nice-to-have, so no one felt stuck with whatever the model decided.
Collaboration & Iteration

Leaving a mark beyond this project

While working on this, I noticed our AI design elements across the company were outdated and relied too heavily on icons, which didn't hold up for this project. I redesigned them and shared the new versions with the design systems team, and they were adopted. That means this work is now part of how every team at the company builds AI features going forward.

Outcomes

What this project demonstrates

Even though a full audit cycle hasn't completed since this feature launched, adoption of Verify Assist has already increased, and audits are moving roughly 12% faster.

  • A real trust deficit, still closing: it isn't fully resolved, but I now have a better working relationship with some of the auditors who distrusted Product before the project began, which makes collaboration and user feedback easier going forward.
  • A scoped rollout that shipped: chose a simpler single-product version first rather than stalling on an all-customer-types design, then came back for the harder multi-product case.
  • Influence beyond the project: redesigned outdated, icon-heavy AI design patterns and got them adopted by the design systems team, so the work now shapes how every team at the company builds AI features.

"A lot of the time when I'm uploading evidence I don't really know what I am uploading, it's just what was given to me, so it's been really nice that the tool now flags exactly what is missing, so when I reach out for more evidence I can tell them exactly what I am looking for."

A customer, after the redesign.

Reflection

Looking back

What I'd do differently

I wish I'd worked on closing the gap between what leadership wanted and what auditors wanted much earlier. That mismatch caused avoidable friction later on.

What I learned

When I showed auditors another designer's work to prove I was addressing their concerns, it confused them more than it helped, even though I had permission to share it. People inside a company respond very differently than outside customers do, and I had to adjust for that. This was my first time designing for people inside my own company instead of external customers, and it taught me just how much more sensitive and political those relationships can be.

Next Steps

Catching mismatches earlier

Right now, we only catch a mismatch with the audit's plan after a customer has already submitted evidence, later than the ideal. In my conversations with auditors, I learned that scope actually gets captured many times, through many different avenues, over the course of an audit, and the sooner we catch a mismatch against any of those, the more friction we remove for everyone involved. This project surfaces scope mismatches earlier than before, but there are upstream places we could be catching them even earlier still, so customers never waste time collecting evidence for something that was never part of their audit, or miss out on automated evidence collection for something that should have been. That's the next step.

Let's connect

Get in touch for opportunities or just to say hi!