An industrial sourcing pilot can look successful long before it proves anything useful. Users log in. The agent runs searches. Supplier names accumulate. Draft messages appear. A dashboard counts activity and the team reports engagement. Yet the same buyer may still be unable to tell which option is real, whether it meets the deadline or who approved contact. Activity has increased while the decision has not improved.
The pilot should test a narrower proposition: can this workflow turn difficult, incomplete demand into a smaller set of evidenced options, in less operator time, while preserving the authority needed for consequential actions? That proposition requires measures across input quality, research, evidence, decision speed, human correction, control failures and physical outcomes. No single metric can carry it.
Set the baseline before the tool changes the work
Choose a bounded task class: for example, obsolete automation spares for one plant, or repair and exchange routes for a defined equipment family. Record how those cases work today. Sample enough completed and failed requests to see the distribution: exact stock found, repair chosen, migration required, requirement withdrawn, no safe option and outcome unknown. Preserve the start time, operator effort, number of handoffs, evidence available at decision and eventual result where it can be verified.
Do not remove inconvenient cases from the baseline. A pilot tested only on clean catalogue numbers will measure search, not industrial sourcing. Include missing suffixes, uncertain diagnoses, impossible deadlines, duplicate listings, condition constraints, alternative approvals and correct no-answer outcomes. Declare exclusions before the run begins.
1. Requirement readiness rate
This is the share of submitted requests that reach a defined minimum state before broad supplier search or outreach: literal identity or an explicit identity gap, machine context, destination, useful-by time, condition boundary and substitution authority. The denominator is all requests accepted into the pilot, not only those that later succeed.
A low rate can mean the intake is badly designed, the source records are weak or the organisation does not know who owns technical approval. A high rate is not automatically good if users are forced to guess. Audit a sample for invented certainty and track the time needed to reach readiness. The desired change is faster clarification with fewer unsupported fields.
2. Credible-route yield
Count how many independently credible routes the workflow produces per ready requirement, and the share of ready requirements with at least one. A route is not a search result. It needs an identified company, a relevant capability or product claim, a current contact path or research route, provenance and no known disqualifying conflict. Several pages copied from one upstream listing count as one route until independence is established.
This measure reveals whether the system broadens the option set across authorised distribution, surplus, repair and migration instead of returning more versions of the same lead. More is not always better. Once the operator has enough distinct, qualified routes to make a decision, extra weak leads add review cost. Report the distribution, not only the mean.
3. Evidence completeness at decision
Define the mandatory claims for the decision: exact identity, quantity, condition, possession, location, dispatch or delivery timing, price and terms, plus the evidence date and source for each. Then measure the share present when an option is presented for approval. Score ‘unknown’ as incomplete but acceptable when it is visible; score an unsupported assertion as a separate quality failure.
This is the counterweight to speed. A workflow that produces an option in ten minutes by omitting condition and possession has not outperformed a slower process. The useful comparison is time to a decision-ready evidence threshold. Different route types may need different fields: a repair offer needs test scope and turnaround; a migration needs engineering and outage implications.
4. Time to comparable options
Measure elapsed time from a ready requirement to the first operator view containing the agreed number of eligible routes against the same fields. Pause or label time waiting for buyer clarification separately from search time and supplier-response time. Otherwise the metric will punish the workflow for an approval delay or hide poor intake inside an aggregate duration.
Use medians and percentiles. Averages conceal the long tail that matters in obsolete supply. Segment by result: exact stock, repair, migration, no safe option. The target is not instant output. It is a meaningful reduction in the time required to reach a comparable choice without lowering the evidence threshold.
5. Operator correction rate
Record material corrections per case and their type: object identity, evidence interpretation, duplicate route, condition, timing, technical compatibility, company identity, authority or recommendation. A correction is material when it changes eligibility, confidence, required action or the decision. Stylistic edits belong elsewhere.
At the beginning, a high correction rate is useful evidence. It exposes where the agent or workflow is weak. The important question is whether repeated correction types fall after prompts, tools or policy code change. Do not drive the number down by discouraging review. Pair it with review coverage and periodically double-check apparently clean cases.
6. Control failure rate
Count any attempt to cross a declared boundary without the required state: external contact without approval, use of a suppressed route, private-data leakage, a CRM write without read-back, a purchase or commitment implied by research, or a product claim promoted without release evidence. The numerator includes blocked attempts as well as actions that escaped, because both reveal pressure on the control.
This metric should have hard thresholds. A pilot does not earn the right to expand permissions because it saved time elsewhere. Review near misses and distinguish a model instruction failure from a missing server-side policy gate. Consequential controls should be enforced by the system, with the model providing context rather than being the only barrier.
7. Verified outcome rate
The final measure is the share of eligible requests with a verified physical or commercial outcome: delivered and accepted, repaired and returned to service, migrated, cancelled by the buyer, or closed with no suitable route. Define the observation window. Separate provider acceptance, sent messages, quotes and orders from the outcome itself. Unknown stays unknown.
Procurement bodies commonly use supplier lead time, on-time delivery, quality, cost and process efficiency as operating measures. Those become relevant after a route turns into a supplier commitment and delivery. The pilot should connect its research-stage measures to those downstream outcomes without claiming causality too early. A faster shortlist is valuable only if it does not produce worse deliveries, quality or total cost.
An illustrative scorecard
Suppose a twelve-week pilot receives 40 requests. Thirty reach the declared readiness threshold; 22 produce at least one credible route; 14 reach a comparable, decision-ready option set; eight lead to an approved order or repair; five are verified as delivered or returned to service within the observation window; three are still in progress; and several close with no safe option. Operators make 19 material corrections, concentrated in identity suffixes and duplicate stock routes. Two unauthorised outreach attempts are blocked by policy before sending.
The honest readout is not ‘five successes from 40’ or ‘hundreds of suppliers discovered.’ It is a funnel with identifiable loss points. Intake loses ten cases. Research fails to produce credible routes for eight ready requirements. Evidence or comparison loses another eight. The correction pattern points to identity parsing and route independence. The blocked outreach attempts show that the control worked and that the workflow still tried to cross the boundary. The next pilot change should address those causes, then rerun the same definitions.
Metrics that mislead on their own
- Logins and active users show access or curiosity, not a better sourcing decision.
- Searches run can rise because the agent is inefficient or the requirement is unresolved.
- Suppliers found rewards duplicates, weak matches and directory size.
- Messages sent rewards premature outreach and ignores consent, suppression and reply quality.
- Quotes received treats incomparable condition, lead time and terms as equivalent.
- Model acceptance or user satisfaction can be useful feedback but does not verify a physical outcome.
- Estimated savings invite selection bias unless the counterfactual, scope and downstream costs are declared.
Decide the stop and expansion rules in advance
The pilot needs a written decision at the end: stop, repair the workflow, extend the observation window or expand scope. Set conditions before results arrive. A critical privacy or authority breach may stop the pilot immediately. Repeated identity errors may block external action while allowing read-only research to continue. Strong time reduction with weak outcome coverage may justify a longer measurement window, not a success claim.
Harvey's adoption-metrics writing is useful because it treats adoption as a set of observable behaviours and outcomes rather than one usage number. Middleman needs the same seriousness, adapted to a workflow where physical state, supplier evidence and operator authority are central. The scorecard should make it harder to tell a flattering story and easier to decide what to improve.