Sourcing managers carry every RFQ alone: finding suppliers, validating certs, comparing risk, defending the choice to finance, then signing the send. I led the design of an experience that gives them team-scale support without taking away ownership. AI assists. The human decides.
No production metrics yet. Prototype UI numbers are illustrative. Not measured results.
I was the sole design lead on this exploration, from problem framing through high-fidelity prototype and validation plan. No ambiguity about scope: I shaped the experience, the trust model, and the narrative we used to align stakeholders around supervision over automation.
I started here because I'd watched this break down inside HikeOn. Sourcing managers in foodservice packaging don't fail from lack of effort. The work is fragmented, serial, and impossible to defend after the fact. The design problem isn't "add AI." It's help one accountable person make a defensible call under time and margin pressure.
They own supplier shortlist, compliance checks, side-by-side comparison, the RFQ document, and the send, often under deadline with finance and compliance reviewing only after the fact.
The answer is scattered. Not in the RFQ. That's the moment trust breaks down.
Slower cycles, missed suppliers, and RFQs that don't reflect full purchase intelligence, weaker negotiating leverage and higher compliance exposure downstream.
Managers feel responsible for outcomes they can't fully inspect forming. Every RFQ is a high-stakes judgment call built from fragmented inputs, with their name on the send.
Certificate chasing and supplier comparison. The same bottlenecks I saw inside the HikeOn ERP, consume days that should go toward negotiation and relationship work.
This became the design brief: give one accountable person a supervised path from requirement to send, with evidence attached, not reconstructed after the fact. Speed without defensibility wasn't an option.
No formal field study. I did not run a multi-week ethnography. I grounded this in shipped ERP work, desk research, artifact review, and informal conversations with sourcing professionals. I state that plainly: the assumptions below guided every decision until validation could test them.
Own the RFQ end-to-end, supplier list, compliance, comparison, and send.
Review supplier choices after the fact, need evidence, not gut feel.
Bear the cost when certs slip, lead times miss, or the wrong supplier ships.
Serial today, one person coordinates every handoff across email, spreadsheets, and tribal knowledge.
I shipped vendor-cert validation and PO drafting inside HikeOn. I noticed where managers lost time, cert chasing, context switching, manual RFQ assembly, and used that as the primary signal for this exploration.
Agentic UX patterns, enterprise trust models, and procurement product benchmarks, to understand what "supervision" looks like when AI is involved.
Informal conversations with sourcing professionals. Not structured interviews. I asked how RFQs run today, where trust breaks, and what they'd need to defend a supplier choice. No participant count to cite; these shaped assumptions, not conclusions.
RFQ templates, cert requirement lists, and the legacy supplier-comparison workflow inside the ERP, what managers already trusted vs. worked around.
Managers need to defend a supplier choice to finance and compliance, speed without explainability is a liability.
Discovery, validation, risk, and drafting run concurrently in theory but serially in practice because one person coordinates everything.
When reasoning lives in email threads, the RFQ document doesn't carry the why, and reviewers push back late.
In a function where a wrong order costs real money, "the AI sent it" is not an acceptable outcome.
I changed the framing early: not "how do we automate sourcing?" but "how do we make one person's parallel work supervisable, so they can run a team-sized process without losing accountability?" (The question that drove every design decision)
Owns the RFQ end-to-end, and carries accountability when finance or compliance asks why.
Find qualified suppliers, validate compliance, compare options, and issue a complete RFQ, without missing a cert or a spec field.
Feel confident hitting send. Not anxious that something was missed in the supplier list, the certs, or the comparison.
Why it ranks first. Without explainability and a human send gate, AI assistance won't be adopted, regardless of how fast it runs. Trust is the adoption gate.
Why it ranks second. One person coordinates discovery, validation, risk, and drafting. The core time sink. Parallel work exists in theory but runs serially in practice.
Why it ranks third. Supplier intel, certs, and purchase history live in different places, email, spreadsheets, ERP modules. Cognitive load is high and errors hide in the gaps.
Why it ranks fourth. Missing requirements surface after the RFQ is drafted, costly to fix downstream and damaging to supplier relationships.
I decided to design for P1 and P2 first. Accountability and the serial bottleneck shaped the core bet: supervised parallel assistance with a trust layer and human send gate. P3 drove the single-view supervision model. P4 drove early requirement validation before drafting.
I define success upfront, and label it honestly. Nothing here is a production result. Validation is in progress. Every metric is tagged projected, estimated, or in test so we don't overclaim.
Requirement captured → RFQ sent. Compare manual baseline vs. prototype in validation sessions.
Method · timed task runs with realistic packaging requirements
RFQs completed per week if adoption holds, depends on trust and workflow fit, not measured yet.
Method · directional estimate post-validation; not a committed target
Can the manager finish the full RFQ run inside the prototype, no email, spreadsheets, or side tabs?
Method · task-based sessions with success / partial / fail criteria
After the session, can they explain why the top vendor was ranked first, without reading the UI?
Method · comprehension check post-session
Accept vs. heavy-edit behaviour on vendor ranking and insight cards, signal for automation bias.
Method · observe accept / reject / override patterns in prototype
Errors and cert gaps caught before send, measured via error-injection tasks in validation.
Method · deliberately wrong agent findings; did the user catch them?
How often users abandon AI output and edit the RFQ from scratch, fallback path usage; not measured yet.
Method · track manual override / full-edit behaviour in sessions
Prototype UI figures elsewhere on this page are illustrative design-spec data. Not results from these metrics.
I wrote these before opening Figma. When stakeholders pushed for a faster, flashier demo, these principles were how I held the line, especially the human send gate and explainability before automation.
The human remains accountable. AI proposes; the manager approves the send. In procurement, accountability can't be delegated to a model, it has to stay visible in the UI.
AI should explain itself. Every recommendation links to evidence, pricing, certs, history, risk. Not a black-box score. Reviewers need the same story the manager saw.
Reduce cognitive switching. Parallel work should surface in one supervisable view. Not seven tabs, threads, and spreadsheets.
Progressive disclosure. Show status and findings first; open reasoning on demand. Experts need depth without drowning novices.
Design for exceptions. Certs pending, conflicting risk signals, missing spec fields. The happy path is rare; the UI must make exceptions legible.
Support expert workflows. Managers should be able to override, edit, and reject without fighting the system, autonomy is the product for this audience.
The obvious pitch was full automation, "watch AI build the RFQ." I explored three directions and rejected two because they didn't solve the real job: helping a sceptical manager inspect, challenge, and still sign.
What we tried on paper. Familiar chat pattern inside the RFQ flow.
Why I rejected it. Hides parallel state. A manager can't supervise seven workstreams or catch errors
they can't see. Trust erodes fast in high-stakes sourcing.
Trade-off accepted. Faster to build, wrong shape for this audience.
What we tried on paper. Comparison and reporting after analysis completes.
Why I rejected it. Still serial. The manager drives each step. Improves visibility, not throughput.
Feels like another BI tool bolted on.
Trade-off accepted. Easier stakeholder sell, doesn't fix the coordination bottleneck.
Why I selected it. Distributes work like a team; keeps the human in command; reasoning stays
inspectable. Maps to how managers already think about sourcing, just not how tools support it today.
Trade-offs I accepted. More design surface: status language, error states, latency perception,
explainability at scale. Worth it for accountability + speed together.
I pushed for the direction that wasn't the flashiest demo. The one a sceptical sourcing manager could inspect, challenge, and still sign. (How I evaluated options)
Prototype evidence for the selected direction lives in Section 07, where each screen appears only after its rationale, not before.
Every screen below earned its place. I don't show UI until the problem, options, and rationale are clear. Select a step to walk through the decision, then see the prototype evidence and why that choice mattered.
Select a step · rationale before screenshot
One card, not a wizard, intent before automation.
Capture what they're sourcing, product, quantity, certifications, delivery, target price, in one place.
Validates completeness; surfaces known cert requirements for the category.
I chose one card, not a wizard: managers need to own the brief before any assistance runs. Front-loading automation felt wrong for an audience that signs what they send.
A requirement complete enough to start supervised analysis.
Layout: Single card, not a multi-step wizard, reduces abandonment and keeps the manager accountable for what they asked for.
Trade-off: Less "magic" on first screen. I accepted that because this audience won't trust output they didn't define.
Supervision, not spectacle, catch issues before they propagate.
See discovery, validation, cert checks, and risk assessment progress without switching contexts.
Live activity view and findings feed, suppliers surfaced, certs verified or flagged, risk tiered, as work completes.
I put status and findings in one view. Not seven tabs. The layout prioritises "what needs attention" over "what looks impressive." Intervention has to happen before the RFQ is drafted.
Shared situational awareness; early catch on cert gaps or spec misses.
Information order: Findings feed above agent status, exceptions surface before progress theatre.
Trade-off: More visual complexity than a chat copilot. I chose inspectability over demo simplicity because hidden state kills trust with this audience.
Recommendations are proposals, not decisions.
Compare vendors, understand why one is ranked first, accept or reject spec suggestions.
Comparison table with evidence-backed ranking; "Why this?" drawer; accept/reject insight cards that re-run affected analysis when accepted.
I made the "Why this?" drawer first-class after early reviews, footnotes failed. Defence is a job equal to discovery; a tooltip doesn't survive a finance review.
A defensible shortlist the manager can explain to reviewers.
Interaction: Comparison table for scanability; drawer for depth on demand, progressive disclosure, not information dump.
Trade-off: Accept/reject on insight cards adds friction. Visible consent matters more than fewer clicks when AI suggested the spec change.
The irreversible action stays human.
Review the assembled RFQ, edit if needed, and send with confidence.
Draft RFQ from requirement, history, supplier intel, and compliance rules; editable document; hard approval gate on send.
I hard-gated send behind explicit approval. Stakeholders wanted one-click launch; I pushed back. Irreversible spend stays human, less automation, more adoption with this audience.
A complete RFQ the manager owns end-to-end.
Governance: Review and send are separate steps. The irreversible action is unmistakable.
Trade-off: Slower than one-click launch. In procurement, clarity on the final step beats demo speed.
I designed every pattern here because someone will ask "why this vendor?". The sourcing manager, a compliance reviewer, or finance. If the product can't answer, it won't get used.
Every output shows estimated certainty, never presented as fact when it's inference.
Why necessary. Prevents false precision; gives managers a signal for where to dig deeper.
Findings link to purchase history, cert records, supplier intel. Not orphaned scores.
Why necessary. Reviewers need an audit trail, not a ranking.
One click from any recommendation to evidence-linked reasoning.
Why necessary. Defence is a job; reasoning can't live in a footnote or a separate doc.
See prototype evidence · Step 3 · Review
Send RFQ is always the manager's action. AI never sends autonomously.
Why necessary. Irreversible spend requires human accountability.
See prototype evidence · Step 4 · Approve
Every draft field and recommendation is editable; insight cards are accept/reject.
Why necessary. Experts must be able to disagree with the system without leaving the flow.
Manual path always available, edit the RFQ from scratch and send without accepting AI output.
Why necessary. Trust grows when users know they're not locked in.
Warning states when certs are pending, risk is elevated, or confidence is low.
Why necessary. Exceptions are the norm in packaging compliance. The UI must foreground them.
Pattern carried from HikeOn · the AI cell
Shipped first on HikeOn, where I learned that suggestion without confidence, reasoning, and accept/edit/dismiss doesn't survive enterprise review. I carried that pattern forward here because the trust layer is the product.
The system divides complex analysis into focused tasks, supplier discovery, cert validation, risk assessment, RFQ drafting, so the manager receives recommendations in one place. They see outcomes and reasoning, not the wiring. That was the design goal: team-scale help, one person's accountability.
Trade-off I accepted: This costs more UI than a chatbot, status language, error states, latency. I chose it because supervision fit the job; a simpler demo wouldn't have earned trust.
Sessions in progress, no user metrics yet. This section is my validation plan plus iterations from stakeholder reviews. I'll add session findings when they exist, not before.
With sourcing and purchasing professionals, guided first pass, then self-directed exploration.
Realistic foodservice packaging requirements with compliance certs, timed, success-criteria scored.
Post-session, can the user explain why a vendor was ranked first without reading the UI?
Deliberately wrong agent findings, do users catch them when reasoning is visible?
Orchestration view, explainability drawer, and approval gate flows, reviewed with internal stakeholders before user sessions.
Visible parallel work builds trust faster than a copilot-style chat.
Supervisory framing raises willingness to use AI among sceptical buyers.
Evidence-linked reasoning reduces blind acceptance of recommendations.
Rubber-stamping because the UI looks confident, accepting recommendations without reading reasoning.
Too much agent narration becomes noise: managers stop reading instead of engaging.
Frozen streams read as broken. Not work-in-progress. Perception problem as much as engineering.
Reassuring copy that doesn't hold up under scrutiny, worse than no explanation at all.
What changed after feedback
Reasoning hidden in expandable rows, easy to miss under deadline.
Chat returns a finished answer, manager can't see parallel work forming or catch errors mid-stream.
Stakeholders wanted full automation, "watch AI do it all." I held for inspectability: parallel work visible, send stays human. The harder design problem was calibrating how much to show without overwhelming a busy manager.
Parallel analysis only works if the UI feels alive. I worked with engineering partners on perceived progress, lag reads as broken, not working. A design problem as much as a technical one.
Buyers have seen automation over-promise. I didn't assume trust, I designed for scepticism: evidence, override, fallback, and a human send gate.
A chatbot is easier to sell internally. I held the line on supervisability because the user's job includes defence, not just speed.
Domain depth from HikeOn carried a lot of weight; validation had to compensate for the absence of a formal field programme. I was explicit about that trade-off.
Leadership here wasn't a title, it was holding a trust-first frame when faster, simpler alternatives were on the table, and making the validation plan honest about what we didn't yet know.
I separate outcomes deliberately. This product hasn't shipped, there are no production metrics. What follows is measured (nothing yet), qualitative (what I delivered and aligned on), and projected (what validation may confirm).
Not shipped. No cycle-time, adoption, or error-rate data exists. Prototype UI numbers are illustrative. Not results.
High-fidelity prototype of the full RFQ run, requirement through approve & launch, ready for validation sessions.
Trust layer designed and documented, confidence, explainability, approval gate, override paths.
Validation framework with hypotheses, methods, and error-injection tasks, honest about what we don't know yet.
Stakeholder alignment on supervision vs. automation: shared language that shaped what got built.
Reduced time from requirement to defensible RFQ draft.
Higher manager confidence explaining vendor selection to reviewers.
Fewer late-stage spec gaps caught only after send.
Projected outcomes are directional hypotheses. Not reported as achieved. User session findings will be added when they exist.
The hardest problem wasn't layout, it was calibrating how much reasoning to show. Too little feels like a black box; too much feels like noise. I'm still tuning that dial in validation.
I assumed speed was the primary job. Conversations and domain work corrected me, defence and accountability rank equal or higher. I changed the success metrics and the trust layer because of that.
I'd start validation earlier, even lightweight, instead of relying on HikeOn immersion alone. Real friction in the loop would have shifted the prototype sooner.
A unified exception queue, one "needs your attention" surface across all streams, so managers don't scan seven statuses to find the one blocker.
My job was often to slow the room down, insist on trust mechanics before automation mechanics. In enterprise AI, that's influence, not illustration. The hiring test isn't whether you can demo agents; it's whether you can make high-stakes decisions feel safe.
Autonomy is only acceptable when it's supervisable and the irreversible step stays human. The trust layer is the product. Not the technology behind it.What I'd carry to the next enterprise initiative