How we ran these studies

By
Jeremy Nash
Study completed:
2026-07-19 to 2026-07-20
Published:
2026-10-01

How we ran these studies

This page covers the method behind all three research studies published here: name-only search vs. group search, the first-result-wrong study, and finding the heirs a skip trace can't. Each has its own population and result; this page is where the shared rules live, so they don't have to repeat on every page.

Pre-registration

Each study has a blueprint document, written and committed before any measurement call was made. The blueprint fixes the population, the arms being compared, the exact metrics, and the confidence-interval method in advance. Any deviation discovered during execution gets logged in the blueprint's own deviation section rather than applied silently. The vendor-rank study went through one full revision, caught by an internal four-lens audit, before its first measurement call; that revision and the reasoning behind it are recorded in its own deviation log.

The point of committing the design first is that a result can't be shaped after the fact by choosing a more favorable cut of the data. Where a study reports a number that moved after a design revision (vendor-rank's earlier 71% pilot, for instance), the earlier number is marked superseded rather than quietly dropped.

What "returned nothing" and "delivered" mean

These two terms recur across the studies and mean the same thing each time:

  • Returned nothing: the search produced zero records, full stop. This requires no judgment about whether any of the records were correct.
  • Target present: the correct record appears somewhere in the response, at any position.
  • Target delivered: for a traditional single-name search, the correct record is entry #1. For group search, the system affirmatively identified that record as the match. These are deliberately different bars, and the difference favors the traditional method on its face, since being named first is an easier test than being affirmatively identified.

"Correct" always means matched against a stated standard, documented on each study's own page: either a vendor-issued relationship ID that neither method produced, or this platform's own prior production identification. Which standard applies, and what that implies about the claim, is stated on each page because it changes what the number can support.

Cluster bootstrap over properties

None of these studies treats its observations as independent draws. Multiple people associated with the same property share a geography and, in several cases, the same underlying search response, so a plain binomial confidence interval would understate the true uncertainty. All intervals here are cluster bootstraps resampled over properties, not over individual trios or pairs, with 5,000 resamples. Where a study reports "n=200" or "n=69," the accompanying property count is the number that actually determines the width of the interval, which is why every quotation on these pages carries both numbers.

Sample provenance

All three studies draw from one tenant's real production history on this platform, covering jobs completed between June 20 and July 17, 2026, with search activity run on July 19–20, 2026. That population is predominantly Texas. None of these studies claims to generalize beyond one tenant's book, one geography, or the time window measured. Jobs affected by a known matching defect, fixed prior to these studies, are excluded from every population built here.

Cost

Vendor-rank measurement consumed 309 data-provider calls against a 500-call ceiling set in advance. The original A/B study consumed 80 calls against the same ceiling. The large-n trio follow-up consumed 282 calls for its group-search arm alone, also against a 500-call ceiling. Each study's ceiling and actual spend are recorded in its own blueprint and results file.

Standing limitations, across all three

  • One tenant, one book, heavily Texas. Nothing here establishes geographic or industry generality.
  • Where "delivered" or "correct" is scored against this platform's own prior output rather than independent field verification, that is disclosed on the relevant page, because it changes what the figure can support.
  • The traditional-search arm, in every study, is a single query shape run once. It is not a simulation of a human investigator trying multiple queries, tools, or follow-ups.
  • No study here makes a claim about how any data provider's ranking algorithm works internally; where position in a result list is discussed, the claim is about position only.
  • No data provider is named in any of these studies pending legal review of the applicable data agreements.

The full blueprint, results file, and underlying data for each study are available on request.

Cite this: GroupSkip Research (2026). How we ran these studies. groupskip.com/research/methodology.