In validation Enterprise procurement Foodservice packaging Trust-first design

Helping procurement teams make high-stakes RFQ decisions, with confidence, speed, and accountability.

Sourcing managers carry every RFQ alone: finding suppliers, validating certs, comparing risk, defending the choice to finance, then signing the send. I led the design of an experience that gives them team-scale support without taking away ownership. AI assists. The human decides.

Role
Lead Product Designer
Solo · end-to-end
Timeline
2025–2026
Concept → validation
Partners
Sourcing advisors
Stakeholder reviews
Focus
Workflow · trust
Supervised assistance
Projected
TIME TO DEFENSIBLE
RFQ DRAFT
In test
CONFIDENCE EXPLAINING
VENDOR CHOICE
In test
TRUST WITH
SCEPTICAL BUYERS
Delivered
FULL PROTOTYPE
+ TRUST LAYER

No production metrics yet. Prototype UI numbers are illustrative. Not measured results.

Supervised RFQ experience: vendor comparison, draft document, and human approval gate
What we built An AI-assisted RFQ experience designed to help procurement teams review complex vendor submissions with greater confidence and transparency, before they sign the send.
My role

What I owned, and what I didn't.

I was the sole design lead on this exploration, from problem framing through high-fidelity prototype and validation plan. No ambiguity about scope: I shaped the experience, the trust model, and the narrative we used to align stakeholders around supervision over automation.

What I owned

  • Discovery and problem framing, grounded in shipped HikeOn ERP work
  • Stakeholder discussions and alignment on supervision vs. automation
  • Procurement workflow redesign, requirement through approve & send
  • AI interaction patterns, supervision view, insight cards, explainability drawer
  • Trust and approval experiences, confidence, sources, human send gate
  • High-fidelity prototype and validation framework
  • Design documentation for trust mechanisms and iteration rationale

What I did not own

  • Model training or prompt engineering
  • Backend infrastructure or agent implementation
  • Production AI engineering or deployment
  • Supplier data sourcing or compliance rule authoring
01 · The Challenge

One person carries the RFQ: and everyone asks why later.

I started here because I'd watched this break down inside HikeOn. Sourcing managers in foodservice packaging don't fail from lack of effort. The work is fragmented, serial, and impossible to defend after the fact. The design problem isn't "add AI." It's help one accountable person make a defensible call under time and margin pressure.

1
Person owns the full RFQ end-to-end
No shared queue
6+
Parallel workstreams on every send
Shortlist · certs · compare · spec · doc · send
0
Built-in audit trail in the workflow
Evidence lives elsewhere
Days
Lost to cert chasing & manual compare
Same HikeOn bottlenecks
Who carries the problem
Primary user · sourcing & purchasing

One manager. Every decision. Full accountability.

They own supplier shortlist, compliance checks, side-by-side comparison, the RFQ document, and the send, often under deadline with finance and compliance reviewing only after the fact.

Foodservice packaging Margin pressure High-stakes send
The gap · goal vs. reality

What success looks like

  • Right suppliers on the shortlist. Not whoever was used last time
  • Verified certs before send, not discovered after delivery
  • Comparable quotes on the same spec dimensions
  • A spec that reflects business need. Not tribal memory
  • Evidence in the RFQ when someone asks "why this vendor?"

What gets in the way

  • Work runs through email, spreadsheets, and tribal knowledge
  • Purchase history and compliance rules live in separate systems
  • Comparison happens manually. If it happens at all
  • The RFQ document doesn't capture the reasoning behind choices
  • Post-send reviews surface gaps the manager couldn't see forming
Where the evidence scatters
No single source of truth · inputs never meet in one place
Email threads Spreadsheets ERP purchase history Compliance rules Tribal knowledge Supplier portals
"Why this vendor?"

The answer is scattered. Not in the RFQ. That's the moment trust breaks down.

Impact at three levels
Business impact

Weaker position at the table

Slower cycles, missed suppliers, and RFQs that don't reflect full purchase intelligence, weaker negotiating leverage and higher compliance exposure downstream.

Emotional impact

Accountable for what you can't see

Managers feel responsible for outcomes they can't fully inspect forming. Every RFQ is a high-stakes judgment call built from fragmented inputs, with their name on the send.

Operational impact

Days on work that shouldn't scale

Certificate chasing and supplier comparison. The same bottlenecks I saw inside the HikeOn ERP, consume days that should go toward negotiation and relationship work.

This became the design brief: give one accountable person a supervised path from requirement to send, with evidence attached, not reconstructed after the fact. Speed without defensibility wasn't an option.

02 · Context & Discovery

Grounded in real procurement pain, with honest limits on formal research.

No formal field study. I did not run a multi-week ethnography. I grounded this in shipped ERP work, desk research, artifact review, and informal conversations with sourcing professionals. I state that plainly: the assumptions below guided every decision until validation could test them.

Stakeholders & user groups

Sourcing / purchasing

Own the RFQ end-to-end, supplier list, compliance, comparison, and send.

Compliance & finance

Review supplier choices after the fact, need evidence, not gut feel.

Operations

Bear the cost when certs slip, lead times miss, or the wrong supplier ships.

Existing workflow · observed & described
1 Requirement Product, qty, certs, delivery, target price
2 Shortlist Memory, spreadsheets, past orders
3 Validate certs Async, often over email
4 Compare Price, risk, lead time, separate tools
5 Assemble & send Manual RFQ when confident enough

Serial today, one person coordinates every handoff across email, spreadsheets, and tribal knowledge.

Research methods · tap to expand

Domain immersion · HikeOn ERP

Shipped work · primary signal

I shipped vendor-cert validation and PO drafting inside HikeOn. I noticed where managers lost time, cert chasing, context switching, manual RFQ assembly, and used that as the primary signal for this exploration.

Desk research

Patterns · benchmarks

Agentic UX patterns, enterprise trust models, and procurement product benchmarks, to understand what "supervision" looks like when AI is involved.

Informal conversations

Not structured ethnography

Informal conversations with sourcing professionals. Not structured interviews. I asked how RFQs run today, where trust breaks, and what they'd need to defend a supplier choice. No participant count to cite; these shaped assumptions, not conclusions.

Artifact review

Templates · legacy flows

RFQ templates, cert requirement lists, and the legacy supplier-comparison workflow inside the ERP, what managers already trusted vs. worked around.

Constraints & unknowns at the start

Fixed constraints

  • The irreversible send must stay human, non-negotiable for trust.
  • Third-party supplier data, not a closed catalogue.

Open unknowns

  • Would a supervision metaphor land with sceptical buyers?
  • How much AI reasoning before transparency becomes noise?
What we learned · synthesised. Not verbatim quotes

Speed isn't the only job

Managers need to defend a supplier choice to finance and compliance, speed without explainability is a liability.

Parallel work, in one head

Discovery, validation, risk, and drafting run concurrently in theory but serially in practice because one person coordinates everything.

Trust breaks in the handoff

When reasoning lives in email threads, the RFQ document doesn't carry the why, and reviewers push back late.

Automation without oversight fails

In a function where a wrong order costs real money, "the AI sent it" is not an acceptable outcome.

I changed the framing early: not "how do we automate sourcing?" but "how do we make one person's parallel work supervisable, so they can run a team-sized process without losing accountability?" (The question that drove every design decision)

03 · Problem Definition

Jobs to be done: and where to focus first.

Primary user · Sourcing manager

Owns the RFQ end-to-end, and carries accountability when finance or compliance asks why.

Jobs to be done
Functional job

Get the RFQ right

Find qualified suppliers, validate compliance, compare options, and issue a complete RFQ, without missing a cert or a spec field.

Emotional job

Sign with confidence

Feel confident hitting send. Not anxious that something was missed in the supplier list, the certs, or the comparison.

Opportunity areas · ranked by severity · tap to expand
P1

Accountability gap

Why it ranks first. Without explainability and a human send gate, AI assistance won't be adopted, regardless of how fast it runs. Trust is the adoption gate.

P2

Serial bottleneck

Why it ranks second. One person coordinates discovery, validation, risk, and drafting. The core time sink. Parallel work exists in theory but runs serially in practice.

P3

Context switching

Why it ranks third. Supplier intel, certs, and purchase history live in different places, email, spreadsheets, ERP modules. Cognitive load is high and errors hide in the gaps.

P4

Late spec gaps

Why it ranks fourth. Missing requirements surface after the RFQ is drafted, costly to fix downstream and damaging to supplier relationships.

I decided to design for P1 and P2 first. Accountability and the serial bottleneck shaped the core bet: supervised parallel assistance with a trust layer and human send gate. P3 drove the single-view supervision model. P4 drove early requirement validation before drafting.

04 · Defining Success

Metrics we defined upfront, and how we label them honestly.

I define success upfront, and label it honestly. Nothing here is a production result. Validation is in progress. Every metric is tagged projected, estimated, or in test so we don't overclaim.

2
Projected
2
Estimated
3
In test
0
Measured
Projected · validation target Estimated · directional In test · sessions running Measured · none yet
Metrics by category · select a tab
Projected

Procurement cycle time

Requirement captured → RFQ sent. Compare manual baseline vs. prototype in validation sessions.

Method · timed task runs with realistic packaging requirements

Estimated

RFQ throughput per manager

RFQs completed per week if adoption holds, depends on trust and workflow fit, not measured yet.

Method · directional estimate post-validation; not a committed target

In test

Task completion without external tools

Can the manager finish the full RFQ run inside the prototype, no email, spreadsheets, or side tabs?

Method · task-based sessions with success / partial / fail criteria

In test

Confidence explaining vendor choice

After the session, can they explain why the top vendor was ranked first, without reading the UI?

Method · comprehension check post-session

In test

Willingness to trust draft recommendations

Accept vs. heavy-edit behaviour on vendor ranking and insight cards, signal for automation bias.

Method · observe accept / reject / override patterns in prototype

Projected

Exception rate

Errors and cert gaps caught before send, measured via error-injection tasks in validation.

Method · deliberately wrong agent findings; did the user catch them?

Estimated

Escalation to manual workflow

How often users abandon AI output and edit the RFQ from scratch, fallback path usage; not measured yet.

Method · track manual override / full-edit behaviour in sessions

Prototype UI figures elsewhere on this page are illustrative design-spec data. Not results from these metrics.

05 · Design Principles

Principles that governed every trade-off.

I wrote these before opening Figma. When stakeholders pushed for a faster, flashier demo, these principles were how I held the line, especially the human send gate and explainability before automation.

PRINCIPLE

The human remains accountable. AI proposes; the manager approves the send. In procurement, accountability can't be delegated to a model, it has to stay visible in the UI.

PRINCIPLE

AI should explain itself. Every recommendation links to evidence, pricing, certs, history, risk. Not a black-box score. Reviewers need the same story the manager saw.

PRINCIPLE

Reduce cognitive switching. Parallel work should surface in one supervisable view. Not seven tabs, threads, and spreadsheets.

PRINCIPLE

Progressive disclosure. Show status and findings first; open reasoning on demand. Experts need depth without drowning novices.

PRINCIPLE

Design for exceptions. Certs pending, conflicting risk signals, missing spec fields. The happy path is rare; the UI must make exceptions legible.

PRINCIPLE

Support expert workflows. Managers should be able to override, edit, and reject without fighting the system, autonomy is the product for this audience.

06 · Exploring Solutions

Three directions, and why only one earned the investment.

The obvious pitch was full automation, "watch AI build the RFQ." I explored three directions and rejected two because they didn't solve the real job: helping a sceptical manager inspect, challenge, and still sign.

OPTION A · COPILOT IN RFQ

What we tried on paper. Familiar chat pattern inside the RFQ flow.
Why I rejected it. Hides parallel state. A manager can't supervise seven workstreams or catch errors they can't see. Trust erodes fast in high-stakes sourcing.
Trade-off accepted. Faster to build, wrong shape for this audience.

OPTION B · DASHBOARD-CENTRIC ANALYSIS

What we tried on paper. Comparison and reporting after analysis completes.
Why I rejected it. Still serial. The manager drives each step. Improves visibility, not throughput. Feels like another BI tool bolted on.
Trade-off accepted. Easier stakeholder sell, doesn't fix the coordination bottleneck.

OPTION C · SUPERVISED PARALLEL ASSISTANCE

Why I selected it. Distributes work like a team; keeps the human in command; reasoning stays inspectable. Maps to how managers already think about sourcing, just not how tools support it today.
Trade-offs I accepted. More design surface: status language, error states, latency perception, explainability at scale. Worth it for accountability + speed together.

I pushed for the direction that wasn't the flashiest demo. The one a sceptical sourcing manager could inspect, challenge, and still sign. (How I evaluated options)

Prototype evidence for the selected direction lives in Section 07, where each screen appears only after its rationale, not before.

07 · The Final Experience

Decisions first, screens as evidence.

Every screen below earned its place. I don't show UI until the problem, options, and rationale are clear. Select a step to walk through the decision, then see the prototype evidence and why that choice mattered.

Select a step · rationale before screenshot

Step 1 · Enter the requirement

One card, not a wizard, intent before automation.

User goal

Capture what they're sourcing, product, quantity, certifications, delivery, target price, in one place.

System response

Validates completeness; surfaces known cert requirements for the category.

Design rationale

I chose one card, not a wizard: managers need to own the brief before any assistance runs. Front-loading automation felt wrong for an audience that signs what they send.

Expected outcome

A requirement complete enough to start supervised analysis.

NEW RFQ · REQUIREMENT Compostable food containers · 50,000 units BPI cert required Delivery · 6 weeks Target price · $0.08/unit Region · North America Start analysis
Early concept Managers capture intent in one card before any analysis runs, ownership of the brief comes first.
Why this design

Layout: Single card, not a multi-step wizard, reduces abandonment and keeps the manager accountable for what they asked for.
Trade-off: Less "magic" on first screen. I accepted that because this audience won't trust output they didn't define.

Step 2 · Supervise parallel analysis

Supervision, not spectacle, catch issues before they propagate.

User goal

See discovery, validation, cert checks, and risk assessment progress without switching contexts.

System response

Live activity view and findings feed, suppliers surfaced, certs verified or flagged, risk tiered, as work completes.

Design rationale

I put status and findings in one view. Not seven tabs. The layout prioritises "what needs attention" over "what looks impressive." Intervention has to happen before the RFQ is drafted.

Expected outcome

Shared situational awareness; early catch on cert gaps or spec misses.

Live supervision view with findings feed and parallel analysis status
Evidence Procurement specialists could monitor parallel work in one place, catching cert gaps and spec misses before they reached the draft.
Why this design

Information order: Findings feed above agent status, exceptions surface before progress theatre.
Trade-off: More visual complexity than a chat copilot. I chose inspectability over demo simplicity because hidden state kills trust with this audience.

Step 3 · Review recommendations and adjust

Recommendations are proposals, not decisions.

User goal

Compare vendors, understand why one is ranked first, accept or reject spec suggestions.

System response

Comparison table with evidence-backed ranking; "Why this?" drawer; accept/reject insight cards that re-run affected analysis when accepted.

Design rationale

I made the "Why this?" drawer first-class after early reviews, footnotes failed. Defence is a job equal to discovery; a tooltip doesn't survive a finance review.

Expected outcome

A defensible shortlist the manager can explain to reviewers.

Vendor comparison table with evidence-backed ranking and Why this drawer
Evidence Evidence drawers reduced blind trust by exposing the reasoning behind recommendations, one click from any ranked vendor.
Why this design

Interaction: Comparison table for scanability; drawer for depth on demand, progressive disclosure, not information dump.
Trade-off: Accept/reject on insight cards adds friction. Visible consent matters more than fewer clicks when AI suggested the spec change.

Step 4 · Approve the RFQ and send

The irreversible action stays human.

User goal

Review the assembled RFQ, edit if needed, and send with confidence.

System response

Draft RFQ from requirement, history, supplier intel, and compliance rules; editable document; hard approval gate on send.

Design rationale

I hard-gated send behind explicit approval. Stakeholders wanted one-click launch; I pushed back. Irreversible spend stays human, less automation, more adoption with this audience.

Expected outcome

A complete RFQ the manager owns end-to-end.

Generated RFQ document with Approve and Launch human approval gate
Evidence Approval checkpoints ensured humans remained accountable for final decisions: send is never an autonomous AI action.
Why this design

Governance: Review and send are separate steps. The irreversible action is unmistakable.
Trade-off: Slower than one-click launch. In procurement, clarity on the final step beats demo speed.

08 · Trust & Governance

Designing trust into AI decisions.

I designed every pattern here because someone will ask "why this vendor?". The sourcing manager, a compliance reviewer, or finance. If the product can't answer, it won't get used.

AI recommendation proposal, not decision TRANSPARENCY LAYER confidence · sources · "why this?" override · escalation · fallback Human approves Send
Trust mechanisms · filter by layer · tap to expand

Confidence indicators

Explain

Every output shows estimated certainty, never presented as fact when it's inference.

Why necessary. Prevents false precision; gives managers a signal for where to dig deeper.

Source transparency

Explain

Findings link to purchase history, cert records, supplier intel. Not orphaned scores.

Why necessary. Reviewers need an audit trail, not a ranking.

"Why this?" drawer

Explain

One click from any recommendation to evidence-linked reasoning.

Why necessary. Defence is a job; reasoning can't live in a footnote or a separate doc.

See prototype evidence · Step 3 · Review

Human approval gate

Control

Send RFQ is always the manager's action. AI never sends autonomously.

Why necessary. Irreversible spend requires human accountability.

See prototype evidence · Step 4 · Approve

Override & edit

Control

Every draft field and recommendation is editable; insight cards are accept/reject.

Why necessary. Experts must be able to disagree with the system without leaving the flow.

Fallback

Control

Manual path always available, edit the RFQ from scratch and send without accepting AI output.

Why necessary. Trust grows when users know they're not locked in.

Escalation paths

Safety

Warning states when certs are pending, risk is elevated, or confidence is low.

Why necessary. Exceptions are the norm in packaging compliance. The UI must foreground them.

Pattern carried from HikeOn · the AI cell

Suggestion Confidence "Why this?" Accept Edit Dismiss

Shipped first on HikeOn, where I learned that suggestion without confidence, reasoning, and accept/edit/dismiss doesn't survive enterprise review. I carried that pattern forward here because the trust layer is the product.

09 · Behind the Experience

How the work gets done, without the manager switching tools.

The system divides complex analysis into focused tasks, supplier discovery, cert validation, risk assessment, RFQ drafting, so the manager receives recommendations in one place. They see outcomes and reasoning, not the wiring. That was the design goal: team-scale help, one person's accountability.

Discovery Validation Risk · Draft One supervised view manager stays in command Human approves send
Model Specialised tasks, one supervisable surface, human send gate.

Trade-off I accepted: This costs more UI than a chatbot, status language, error states, latency. I chose it because supervision fit the job; a simpler demo wouldn't have earned trust.

10 · Testing & Iteration

How we're evaluating the bet, honestly, while it's still open.

Sessions in progress, no user metrics yet. This section is my validation plan plus iterations from stakeholder reviews. I'll add session findings when they exist, not before.

Plan Test Learn Iterate REPEAT
Validation framework · select a view

Prototype walkthroughs

With sourcing and purchasing professionals, guided first pass, then self-directed exploration.

Task-based sessions

Realistic foodservice packaging requirements with compliance certs, timed, success-criteria scored.

Comprehension checks

Post-session, can the user explain why a vendor was ranked first without reading the UI?

Error-injection

Deliberately wrong agent findings, do users catch them when reasoning is visible?

Stakeholder reviews

Orchestration view, explainability drawer, and approval gate flows, reviewed with internal stakeholders before user sessions.

Visible parallel work builds trust faster than a copilot-style chat.

Supervisory framing raises willingness to use AI among sceptical buyers.

Evidence-linked reasoning reduces blind acceptance of recommendations.

Automation bias

Rubber-stamping because the UI looks confident, accepting recommendations without reading reasoning.

Transparency overload

Too much agent narration becomes noise: managers stop reading instead of engaging.

Latency breaks "live" feel

Frozen streams read as broken. Not work-in-progress. Perception problem as much as engineering.

Plausible but wrong explainability

Reassuring copy that doesn't hold up under scrutiny, worse than no explanation at all.

What changed after feedback

Explainability placement

Initial approach
Vendor reasoning in footnotes and expandable rows, kept the comparison table clean.
What stakeholders said
Internal reviews: reasoning felt buried. "I'd miss this under time pressure."
What changed
I promoted a first-class "Why this?" drawer, one click from any recommendation.
Why it mattered
Defence is a job, not a nice-to-have. If reasoning isn't obvious, the tool won't survive a finance review.
Before · footnotes in table

Reasoning hidden in expandable rows, easy to miss under deadline.

Vendor A · Rank 1 · ⓘ see note 3
Vendor B · Rank 2 · ⓘ see note 7
After · first-class drawer Why this drawer exposing evidence-linked vendor reasoning
Evidence surfaced on demand, one click, not a scavenger hunt through footnotes.

Insight card acceptance

Initial approach
Spec suggestions applied silently when the manager moved forward, fewer clicks.
What I challenged
Silent acceptance looked like automation bias waiting to happen. No signal of what the human actually agreed to.
What changed
Explicit accept / reject on every insight card; accepted changes re-run affected analysis.
Why it mattered
Accountability requires visible consent, especially when AI suggested the spec change.

Send action

Initial approach
"Launch RFQ" combined draft review and send in one ambiguous step.
What stakeholders wanted
Faster path to send, fewer confirmation steps for the demo.
What changed
I separated review from send and hard-gated approval. The irreversible action is unmistakable.
Why it mattered
In procurement, one wrong send costs real money. Clarity beats speed on the final step.

Direction: chat vs. supervision

Initial pressure
Full-automation narrative, copilot that "does the RFQ for you."
What I pushed back on
Sourcing managers need to see work forming, not read a finished answer. Hidden state kills trust.
What changed
Committed to the supervision view: live activity, findings feed, inspectable streams.
Why it mattered
Adoption depends on inspectability. I traded demo flash for something a sceptical buyer could challenge.
Before · copilot pattern

Chat returns a finished answer, manager can't see parallel work forming or catch errors mid-stream.

"Here are your top 3 vendors and a draft RFQ."
No visibility into cert checks, risk flags, or missing spec fields.
After · supervision view Live supervision view with findings feed replacing hidden copilot output
Parallel work visible as it completes, intervention possible before the draft is assembled.

Confidence display

Initial approach
Numeric confidence scores on every output, precise, comparable.
What I worried about
Scores that look precise but aren't calibrated read as false authority. Stakeholder reviews flagged confusion risk.
What changed
Reframed as estimated certainty with plain-language labels, never presented as fact when it's inference.
Why it mattered
Still open in validation, but I'd rather signal uncertainty honestly than train rubber-stamping.
11 · What Was Hard

Constraints, resistance, and how decisions got made.

Design tension

Trust before speed

Stakeholders wanted full automation, "watch AI do it all." I held for inspectability: parallel work visible, send stays human. The harder design problem was calibrating how much to show without overwhelming a busy manager.

Product tension

Latency vs. live feel

Parallel analysis only works if the UI feels alive. I worked with engineering partners on perceived progress, lag reads as broken, not working. A design problem as much as a technical one.

Organisational

Scepticism toward AI

Buyers have seen automation over-promise. I didn't assume trust, I designed for scepticism: evidence, override, fallback, and a human send gate.

Organisational

Competing simplicity narratives

A chatbot is easier to sell internally. I held the line on supervisability because the user's job includes defence, not just speed.

Research gap

Thin formal discovery

Domain depth from HikeOn carried a lot of weight; validation had to compensate for the absence of a formal field programme. I was explicit about that trade-off.

Leadership here wasn't a title, it was holding a trust-first frame when faster, simpler alternatives were on the table, and making the validation plan honest about what we didn't yet know.

12 · Outcomes

What I can claim today, and what I can't yet.

I separate outcomes deliberately. This product hasn't shipped, there are no production metrics. What follows is measured (nothing yet), qualitative (what I delivered and aligned on), and projected (what validation may confirm).

Measured outcomes

Production data · 0
0

Not shipped. No cycle-time, adoption, or error-rate data exists. Prototype UI numbers are illustrative. Not results.

Qualitative outcomes

Delivered · 4

High-fidelity prototype of the full RFQ run, requirement through approve & launch, ready for validation sessions.

Trust layer designed and documented, confidence, explainability, approval gate, override paths.

Validation framework with hypotheses, methods, and error-injection tasks, honest about what we don't know yet.

Stakeholder alignment on supervision vs. automation: shared language that shaped what got built.

Projected outcomes

If validation confirms · 3

Reduced time from requirement to defensible RFQ draft.

Higher manager confidence explaining vendor selection to reviewers.

Fewer late-stage spec gaps caught only after send.

Projected outcomes are directional hypotheses. Not reported as achieved. User session findings will be added when they exist.

13 · Looking Back

Surprises, wrong assumptions, and what I'd do differently.

Surprise

The hardest problem wasn't layout, it was calibrating how much reasoning to show. Too little feels like a black box; too much feels like noise. I'm still tuning that dial in validation.

Wrong assumption

I assumed speed was the primary job. Conversations and domain work corrected me, defence and accountability rank equal or higher. I changed the success metrics and the trust layer because of that.

Do differently

I'd start validation earlier, even lightweight, instead of relying on HikeOn immersion alone. Real friction in the loop would have shifted the prototype sooner.

Version 2 focus

A unified exception queue, one "needs your attention" surface across all streams, so managers don't scan seven statuses to find the one blocker.

Leadership lesson

My job was often to slow the room down, insist on trust mechanics before automation mechanics. In enterprise AI, that's influence, not illustration. The hiring test isn't whether you can demo agents; it's whether you can make high-stakes decisions feel safe.

Autonomy is only acceptable when it's supervisable and the irreversible step stays human. The trust layer is the product. Not the technology behind it.
What I'd carry to the next enterprise initiative