Key takeaways

  • Measure leading indicators the pilot can influence: relevant opportunities surfaced, source quality, decision speed, actionability, stakeholder coverage, decision throughput, and adoption.
  • Capture a baseline before the pilot, keep the denominator beside every rate, and define each metric precisely enough that two people calculate it the same way.
  • Let the customer set acceptance thresholds from its present workflow and business needs; vendor-selected universal pass marks are not credible.
  • Treat every surfaced item as a decision object that should end in pursue, validate, watch, or pass, with a reason and next review date where applicable.
  • End with a decision memo that separates measured results, user judgment, limitations, and the conditions for expanding, extending, or stopping.

A public-sector intelligence pilot should not be judged by whether your company wins a contract in 30 or 45 days. That is an attractive metric because it sounds close to revenue. It is also a bad experimental design. Government buying cycles do not begin when a software trial begins, and an award that lands during the trial usually reflects months of prior capture work, relationship building, procurement decisions, and proposal effort.

A useful pilot asks a narrower and more consequential question: did this workflow help the team notice the right buyer movement, trust the supporting evidence, decide what deserved attention, and take the next action faster than it did before? Those are leading indicators. They occur inside the evaluation window, can be tied to the workflow under test, and reveal whether the product deserves a place in the team’s operating system.

This guide gives you a practical scorecard, a six-week evaluation design, and a final decision memo. The point is not to make every pilot pass. It is to make the decision at the end harder to argue with.

What should a 30–45 day public-sector pilot actually prove?

It should prove that a repeatable workflow changed the quality or speed of decisions available to the team during the pilot. That is different from proving market demand, winning a procurement, or forecasting revenue.

The distinction matters because public-sector selling contains long, externally controlled intervals. Agencies conduct market research, refine requirements, secure funding, choose acquisition paths, and publish solicitations on their own schedules. The Federal Acquisition Regulation encourages exchanges with industry from the earliest identification of a requirement, including RFIs, draft RFPs, conferences, and market research. That confirms why early signals and stakeholder context can matter. It does not mean a vendor can force a requirement to become an award within six weeks.

The pilot’s unit of value is therefore not the closed contract. It is the decision object: an agency, buyer event, pre-RFP signal, or posted RFP that the team must evaluate. A strong workflow helps each object move to one of four explicit account-triage states:

  • Pursue: the fit and timing justify a named action and owner.
  • Validate: the fit may be real, but the team names the uncertainty and a credible way to resolve it.
  • Watch: the evidence is promising but incomplete, so the team records the next condition or review date.
  • Pass: the team declines for a recorded reason it can learn from.

If the pilot produces more alerts but leaves them undecided, it has increased workload rather than improved the operation. The goal is not another feed to scroll. It is a dependable path from source to judgment to action. For the broader daily process, see the government sales operating cadence.

Five rules for a scorecard you can trust

1. Capture the baseline before anyone sees the new workflow

Measure the current process over a representative period or reconstruct it from timestamps and records. How many items did the team review? How many were confirmed relevant? How long did a decision take? What proportion ended with an owner and next step? A before-and-after claim without a baseline is only an impression.

Keep the comparison fair. Use the same territories, fit criteria, participants, and opportunity types where possible. Record known differences. The baseline does not have to be perfect; it has to make the uncertainty visible.

2. Put the denominator next to every percentage

“Eighty percent relevant” is meaningless unless the reader knows whether the team reviewed five items or five hundred. Every rate should show its numerator and denominator: 16 of 20 reviewed items, 9 of 12 priority accounts, or 14 of 18 accepted signals. Also define exclusions before launch. Otherwise, difficult cases quietly disappear from the calculation.

3. Set acceptance thresholds before results exist

A threshold should describe the minimum change that would make the workflow worth adopting for this customer. It is not an industry benchmark. Let the customer set thresholds from its baseline and business constraint, with the vendor documenting the formula and confirming that the required data can be observed.

4. Tie each metric to a decision

The GAO Agile Assessment Guide says metrics should be established early, aligned with incentives, tailored to the work, and designed to support specific decisions. Apply the same discipline here. If nobody can say what they would do differently when a metric rises or falls, remove it from the scorecard.

5. Preserve the evidence, not just the score

Keep the source link, recorded verification state, timestamps, fit rationale, user decision, and next action behind each result. Aggregate scores tell you what happened. The underlying records tell you whether it happened for the right reasons. This is especially important when a small number of high-value opportunities can be hidden by averages.

The public-sector pilot scorecard

Use the following dimensions as a menu. Select the five to seven that match the pilot hypothesis, define the formula in writing, and add a customer-set acceptance threshold. The examples illustrate the form of a threshold; they are not Settle benchmarks or universal standards.

A scorecard for a 30–45 day public-sector opportunity-intelligence pilot. Thresholds must be set by the customer before launch.
DimensionMetric and denominatorEvidence to retainExample customer-set threshold
Relevant opportunities surfacedCount confirmed relevant, plus confirmed-relevance rate = relevant items ÷ all items reviewed by the customerFit criteria, territory, reviewer decision, and reason for acceptance or rejectionSurface at least the agreed count in target markets while maintaining the agreed relevance rate
Source quality and verificationReproducible-source rate = items whose cited source and recorded check state can be independently reviewed ÷ items sampledPrimary-source URL, publication date when available, last-check timestamp, and any caveat or stale stateEvery sampled high-priority item has a reviewable source; exceptions are visibly labeled rather than presented as verified
Time-to-decisionMedian elapsed time from surfacing to pursue, validate, watch, or pass; report unresolved items separatelySurfaced timestamp, decision timestamp, decision maker, and queue statusReduce median decision time by the customer’s agreed amount without lowering relevance quality
ActionabilityUsable-next-action rate = accepted items with a next action the owner judges usable ÷ accepted items reviewedRecommended action, owner judgment, assigned owner, due date, and completion stateThe agreed share of accepted items leaves review with a usable next action and owner
Stakeholder coverageCoverage rate = prioritized accounts with the agreed stakeholder roles identified ÷ prioritized accounts reviewedRole, person or office, source, confidence caveat, and relationship statusPriority accounts cover program and mission, procurement and control, and ecosystem and influence—or explicitly record which lane is unknown
Decision throughputDisposition rate = items assigned pursue, validate, watch, or pass ÷ all items requiring review; include reason completenessDisposition, reason, owner, next review date for watch, and unresolved ageReach the agreed disposition rate by the review SLA, with no silent backlog older than the agreed limit
Workflow adoptionWeekly active decision makers and assigned-action completion; use invited eligible users as the denominatorMeaningful review, decision, assignment, and completion events—not page views or logins aloneEach required role completes the actions expected in the operating design for the agreed number of weeks

Relevant opportunity count and relevance rate belong together. A system can inflate precision by showing almost nothing, or inflate volume by sending noise. Read both numbers with the customer’s target-market coverage. The account tiering behind that evaluation should come from a documented model such as this SLED account prioritization template, not a list rewritten after the pilot starts.

Source quality also needs careful wording. A recorded verification state means the workflow preserves what was checked, when it was checked, and what source supported it. It does not mean every fact on the internet is universally true or current forever. A useful test is reproducibility: can a reviewer open the cited public record and understand why the item was surfaced? Learn more about the source types in Settle’s guide to pre-RFP signals.

How this framework was built: This scorecard is an editorial evaluation template. Define the baseline, metric formulas, thresholds, systems of record, and evidence rules with the customer before the pilot begins.

Where Settle is not the system of record for ownership or completion, record those fields in the customer’s CRM or evaluation worksheet; do not attribute CRM-recorded work to the pilot product.

A week-by-week evaluation design

A 30-day pilot can run in four working weeks plus setup and readout. A 45-day pilot gives you two additional weeks to test consistency. Do not spend half the window configuring the tool. Complete scope, users, criteria, and data access before day one whenever possible.

The pilot should resemble the operating rhythm the team would keep after the vendor’s implementation team steps away.
TimingObjectiveWorkDecision gate
Before day 1Freeze the evaluation designName the hypothesis, territories, opportunity types, participants, baseline period, metric formulas, data owner, and acceptance thresholdsDo not launch until both sides approve the one-page measurement plan
Week 1Calibrate fit and evidenceReview an initial sample together; correct fit criteria; test source links and recorded verification states; document exclusionsConfirm that the workflow can represent the customer’s actual pursue and pass logic
Week 2Run the real decision cadenceRoute live items to named users; record pursue, validate, watch, or pass; measure decision time, backlog age, and reason completenessResolve workflow friction before interpreting adoption numbers
Week 3Test actionability and stakeholder contextFor accepted items, review recommended next actions and map roles across program and mission, procurement and control, and ecosystem and influence, with source caveatsConfirm whether context changes who acts, what they do, or when they do it
Week 4Test the handoff into capture and responseMove selected pursue items into the existing CRM or response process; confirm the source record, fit rationale, owner, and next action survive the handoffDecide whether the workflow closes a real operational gap or creates duplicate work
Weeks 5–6, if usedTest durabilityRepeat without vendor-led meetings; sample rejected items; check for drift, unresolved queues, role gaps, and user drop-offSeparate a temporary onboarding effect from a workflow the team can sustain
Final readoutMake the adoption decisionLock the dataset; compare baseline with pilot; review misses and limitations; produce the decision memoExpand, extend for one named uncertainty, or stop

A concierge pilot can look excellent because the vendor runs every review. That help may be appropriate during onboarding, but it must be disclosed. By the final weeks, the customer’s own operating cadence should carry the workflow.

How to define the baseline and acceptance thresholds

Start with the business problem, not the product dashboard. If the problem is late discovery, measure when the team first saw a relevant event and whether the source was actionable. If the problem is feed overload, measure decision time, backlog age, and pass reasons. If the problem is shallow capture, measure stakeholder-role coverage and whether the context changed an action.

  1. Write the hypothesis. Example: “For our named SLED territories, the workflow will help account owners make documented decisions on relevant buyer movement faster than the current shared-inbox process.”
  2. Select the comparison period. Use a recent period with similar staffing, territories, and opportunity types. Record important differences.
  3. Define every field. Specify when the timer begins, what counts as reviewed, who can confirm relevance, and how unresolved items are handled.
  4. Choose the minimum useful change. The customer should state what improvement would justify switching cost, training, and ongoing spend.
  5. Name guardrails. A faster decision is not progress if relevance quality collapses; more actions are not progress if owners judge them unusable.
  6. Freeze the plan. Changes are allowed only when documented with the reason and the effect on comparability.

This is the same logic acquisition teams use when they connect required capability, performance standards, risks, and decision milestones. FAR 7.105 requires acquisition plans to identify the milestones at which decisions should be made. Your internal software pilot is not a federal acquisition plan, but the discipline is useful: define the need, the evidence, the risks, and the decision points before work begins.

Metrics that look persuasive but should not decide the pilot

  • Contract wins or revenue: too slow, too sparse, and too confounded for a short evaluation. Track downstream outcomes later as longitudinal evidence, not as the 45-day acceptance test.
  • Raw opportunity volume: rewards noise. Pair count with confirmed relevance, market scope, and review capacity.
  • Logins, page views, or alerts sent: show exposure, not value. Measure meaningful review, decisions, assignments, and completed actions.
  • Estimated hours saved without a baseline: invites optimistic arithmetic. Observe a bounded task before and during the pilot, or label the estimate as a user-reported perception.
  • A single composite score: can hide a fatal weakness. Strong adoption cannot compensate for untrustworthy sources, and fast decisions cannot compensate for poor relevance.
  • Win probability: should not be treated as calibrated unless the system and dataset have actually established calibration. Use transparent fit criteria and observed evidence instead.

The final decision memo

Do not end with a celebratory slide deck and a vague promise to circle back. Produce a short decision memo that someone who did not attend the pilot can audit. It should contain:

  1. Decision requested: expand, extend, or stop, including scope and owner.
  2. Hypothesis and scope: markets, users, opportunity types, dates, and excluded work.
  3. Baseline and results: every metric with numerator, denominator, threshold, result, and data-quality note.
  4. Representative evidence: a small set of wins, misses, false positives, and unresolved cases with source records.
  5. User judgment: where the workflow changed a decision or action, and where it created friction.
  6. Limitations: seasonality, small samples, vendor assistance, missing integrations, or events that predated the pilot.
  7. Operating design: cadence, owners, governance, and systems of record if adopted.
  8. Next checkpoint: the date and downstream outcomes to review after a longer operating period.

An extension is justified only when it resolves a named uncertainty. “We need more time” is not enough. A valid extension might test whether relevance holds in a second territory, whether the team sustains adoption without vendor facilitation, or whether a required CRM handoff works. Give it its own threshold and end date.

What Settle should be evaluated on

Settle is designed to help public-sector teams move from source-backed buyer signals and relevant RFPs to a grounded decision and next action. A Settle pilot should therefore be evaluated on the quality of the fit criteria, the ability to review the original source and recorded verification state, the usefulness of buyer and stakeholder context, the speed and completeness of human-recorded pursue, validate, watch, or pass decisions, and the handoff into response work.

It should not be evaluated as an autonomous sales representative, a promise that every source is correct forever, or a guarantee of government revenue. It should also not require you to abandon the tools that already work. The practical question is whether it becomes the action layer between the sources your team monitors and the CRM, capture, and response systems where work continues. See how Settle’s workflow works.

A good pilot gives both sides permission to learn and permission to stop. If the evidence clears the customer’s pre-agreed thresholds, the adoption case is concrete. If it does not, the team knows exactly which hypothesis failed. Either result is more valuable than declaring success because a dashboard looked busy for six weeks.

Frequently asked questions

Can a 30–45 day pilot measure government contract wins?

Usually not in a useful or attributable way. Government opportunities often mature across months or years, and an award inside a short pilot was likely influenced by work that began before the pilot. Measure the earlier decisions and actions the workflow can affect during the evaluation window.

What is the most important metric in a public-sector intelligence pilot?

There is no universal single metric. Start with the customer’s binding problem. For a team missing relevant opportunities, confirmed relevance and source reproducibility may matter most. For a team buried in feeds, time-to-decision and decision throughput may be more important.

Who should set the pilot acceptance thresholds?

The customer should set them with the vendor before launch, using the current-state baseline, staffing reality, and desired operating change. The vendor can help define formulas and feasibility, but should not invent a generic pass mark after seeing the results.

How many users should participate in the pilot?

Include the smallest group that represents the real workflow: usually the public-sector leader, one or more account owners, and whoever validates procurement or proposal work. A large audience creates noise; a single champion can hide adoption and handoff problems.

What should the final pilot decision memo contain?

It should restate the hypothesis and scope, show baseline-versus-pilot results with denominators, summarize qualitative evidence and limitations, document unresolved risks, and recommend one of three outcomes: expand, extend for a named uncertainty, or stop.