Skip to main content
Bel Covo

Research Methodology

The rules every published research stat satisfies before it ships, and the canonical reference for the Bel Covo quote-records dataset cited across our research pillars.

Bel Covo is an Oklahoma concrete flooring contractor. We polish slabs, install epoxy and polyaspartic coatings, and pour flatwork across the state, with launch coverage extending into Arkansas, Tennessee, Mississippi, Missouri, and Texas. Every published number on /research/ derives from our own quote and project records, not third-party industry surveys.

The dataset behind this methodology hub is 245 anonymized concrete flooring quotes, spanning 2021 through April 2026, scoped to Oklahoma plus the five neighboring launch states. Two populations sit inside that 245: the consented forward-going population, where quote requesters explicitly opt in to anonymized aggregates, and the historical population, where pre-policy quotes are aggregated under tighter cell-level controls. Each pillar page declares which population the cited number comes from. Aggregates from the consented population publish only when the cell carries 30 or more records. Historical aggregates publish at any non-empty cell so long as the geographic floor (state, or CBSA-metro for high-volume areas) holds.

Numbers refresh under two triggers. Every pillar gets a calendar refresh at least once per year so cited values do not drift past their reference period. Off-cycle refreshes fire when an aggregate moves more than 15 percent between any two consecutive snapshots, on the assumption that a sustained move that large reflects a real market shift our readers should see immediately rather than a year later.

Cite this methodology and any pillar that links here as: Bel Covo Concrete Craftsmen, "Research Methodology," https://belcovo.com/research/methodology/, accessed [date]. Pillar-level citations follow the same form, swapping the URL for the pillar page. The dataset license is Creative Commons Attribution 4.0 International.

Research methodology

In Plain English

This document is the rulebook every number on a Bel Covo /research/<topic>/ page satisfies before we publish it. Pillar pages are stats-driven articles built to be cited by journalists, by AI engines like Perplexity and Google AI Overviews, and by competing publishers fact-checking the work. A page is only as defensible as the rules behind its numbers, so we publish the rules in full and we hold every number against them.

The core idea is simple. We publish a stat only when (a) it comes from a sample large enough that the number is not noise, (b) the people behind that sample cannot be re-identified from the published cut, (c) the source is named and the source’s URL is verified, and (d) the cut, the date, and the methodology version are locked into a frozen reference that does not silently change after a journalist links to it. Anything that fails one of those four checks does not get published, full stop.

We treat these rules as load-bearing. Our team checks each one against every published cut before the page goes live. A reviewer challenging a number on a Bel Covo research page can start here: every threshold, every floor, and every disclosure standard is described below. If the rules ever change, the methodology version changes with them, and prior pages stay frozen at the version under which they were published.

Why this document exists

Made-up numbers from a contractor brand are catastrophic to trust. “Fresh” stats that silently change after a journalist cites them are almost as bad. The pillar-page strategy only works if the numbers are defensible to a hostile reader: a fact-checker, a competitor, an AI engine ranking authority. This document is the artifact a hostile reader is meant to read.

Two audiences use this doc. Our team consults it to decide whether a draft cut is publishable. External readers (journalists, AI citation engines, auditing publishers) consult it to decide whether to trust a number we published. The rules below are written to satisfy both.

Sample size floor

The default minimum sample size for any published cut is n >= 30.

This floor applies to the consented population. Methodology v2 (added 2026-04-26) admits a separate historical population that publishes without an n floor; see “Historical records (added under v2)” below for the v2 carve-out and the principled reason for it.

Cuts with fewer than 30 underlying quotes are omitted from publication, not zeroed. A zero is a claim. An omission is a refusal to claim. A reader who notices a regional or service cut missing from a published table should assume we withheld it deliberately because n was below the floor.

Why 30: it is the standard rule of thumb where central limit theorem normality assumptions stabilize for sample medians and means. The same threshold is used in the published methodologies of NerdWallet, Apartment List Research, and Zillow Research, three peer publishers in adjacent verticals (consumer finance, rental housing, real estate). Citing the same floor as those peers makes the choice unsurprising to a reader auditing it.

Per-cut overrides are allowed but must be documented in the cut’s method label and explained in the page’s methodology section. Two patterns:

  • Loosened below 30: only allowed when a strong Bayesian prior makes a smaller sample informative. Example: a cut where the metric is the proportion of customers who selected a binary option, and prior pillars on the same topic have already established the population variance. A loosened cut must be flagged in body copy as preliminary and must show its n.
  • Tightened above 30: encouraged for any cut where variance is visibly high (e.g., commercial vs. residential mixes within a single ZIP3). The default n-floor is a floor, not a target. If a tighter floor produces more defensible numbers, we use it.

Cuts must never be tightened in a way that quietly removes inconvenient data while leaving favorable cuts at the lower threshold. The n-floor decision is a methodology choice and bumps the methodology version (see below).

Anonymization rules

The published research must not re-identify any individual customer. The rules:

Geographic floor: ZIP3, never ZIP5

The smallest geographic unit that may appear in a published region cut is ZIP3, the first three digits of a US ZIP code. ZIP5 is forbidden. ZIP5 cuts can re-identify rural customers in low-density areas where a single household carries the cut.

ZIP3 cuts must always declare the state alongside the ZIP3. ZIP3 prefixes are not unique to a state. The 730 prefix spans Oklahoma and Texas. The 671 prefix spans Kansas and Missouri. A ZIP3 cut without a state code is ambiguous, so we publish the two together or not at all.

State cuts and metro cuts (defined as Census-designated Core Based Statistical Areas, CBSAs) are allowed unrestricted. National cuts are allowed for any aggregate that combines all regions.

What counts as a Bel Covo quote

A “Bel Covo quote” is any submitted quote in our quote records tagged with research consent. Three conditions all hold:

  1. The quote was submitted through our public quote calculator (not internal test traffic, not manual office entries).
  2. The customer’s research-consent flag is set to true at submission time.
  3. The quote was submitted successfully with the required ZIP and service selections. Incomplete or malformed attempts are excluded from the population entirely.

The consent flag is opt-in, not opt-out. A customer who does not actively grant research consent is excluded from every aggregate. Changes to the consent wording do not bump the methodology version unless they materially change which submissions qualify.

What never leaves our records

The following fields exist in our internal records but are stripped before any aggregation step and never appear in any published cut:

  • Customer name (first, last, business)
  • Street address (the address fields used for routing only)
  • Phone number
  • Email address
  • Free-text notes

Aggregations operate on the stripped projection. Any handling of the original PII fields (routing a quote to a branch, sending follow-up email, scheduling a site visit) is handled separately from the research records and is firewalled from anything we publish.

Historical records (added under v2)

Bel Covo’s quote records go back further than the date research consent became an explicit opt-in field. We treat those older records as a separate population called historical records. Historical records carry the same anonymization rules as consented records (no street address, no name, no phone, no email leave our records), but because consent was not asked at submission time, we publish historical aggregates under additional constraints designed to prevent re-identification and to keep the historical claim distinct from the consented claim.

Historical aggregates publish only at the CBSA-metro level or coarser. ZIP3 cuts on historical records are not published. Every historical aggregate is clearly labeled with the date range it covers and the historical population tag.

The v2-specific rules:

  • Geographic floor for historical records: CBSA-metro level minimum. ZIP3 cuts are forbidden for historical aggregates. The principled reason is that customers in 2021 did not opt into research participation, so we cannot publish them at the granularity that an opt-in customer in 2026 sees. Metro-or-coarser is the level where re-identification is structurally impossible (CBSAs are large, MSAs are larger).
  • Mixed-population aggregates are allowed when explicitly composed: a pillar may publish a metro-level “consented (2025 to 2026): n=X, median $Y” aggregate alongside a “historical (2021 to 2025): n=Z, median $W” aggregate. It MAY NOT silently combine the two into a single aggregate without the population labels.
  • Sample-size floor does not apply to historical aggregates (v2 amendment, 2026-04-26). The default n >= 30 floor applies to the consented population only. Historical aggregates publish without a sample-size floor, and every published historical cut documents its n in plain sight so the reader can weigh it. The principled reason: the floor exists to keep ordinary noise out of consented aggregates that compete for trust against zero-data publishers; for the historical population the alternative to publishing a small-n cut is publishing nothing about a real cohort of past customers, and an n that is visible is more honest than an aggregate that quietly disappears. The CBSA-metro+ geographic floor still applies to historical aggregates and is not relaxed.
  • Body copy disclosure: any pillar using historical aggregates explains, in plain prose near the headline number, why the historical population exists and how it differs from the consented population. The disclosure is not buried in a footnote; it sits where the reader sees it.
  • Trust strip surfaces methodology version: every pillar’s trust strip reads methodology v2 when historical records appear in the snapshot, v1 when only consented records are used. A reader scanning the trust strip can tell at a glance which methodology produced the numbers below it.

A pillar author choosing between v1 and v2 follows a simple rule. If the pillar’s claims are best supported by consented data alone (or if consented data already meets the n>=30 floor at the desired granularity), publish at v1. Use v2 when the historical population materially improves the claim: more data depth, longer time window, fills a metro cut that consented data alone cannot reach, or the consented data is too thin to publish at all and the historical record is the honest fallback. The integrity guarantee of v1 (consented-only, ZIP3-allowed, n >= 30 enforced) stays intact for v1 pillars.

Permitted and forbidden cut shapes

Not every aggregate cut shape that snapshot.mjs CAN compute is one we publish. The cut-shape allowlist is set by docs/comms-rules.md Rule 17 (council ruling 2026-04-26 on contractor-specific cost decomposition).

Permitted cut shapes (publish freely):

  • cost_per_sqft per region (state, CBSA-metro, ZIP3 for consented; state and CBSA-metro for historical). Median, IQR (p25/p75), CI bounds.
  • cost_per_sqft per project-size band (small/medium/large/very-large) within a region.
  • project_size_sqft per region. Median, IQR.
  • labor_hours_per_1000_sqft per service code per region. Median, IQR. This is the council-approved primary-research citation hook; it is the dominant labor-intensity signal that journalists, AI engines, and competing publishers want, without exposing Bel Covo’s internal cost structure.
  • count_per_service_code and share_per_service_code per region. Counts of quoted records per service code within a region cut, and the share each code represents of the region’s total. This is system-distribution (what people choose), not cost decomposition (Rule 17 stays clean). Permitted because the metric exposes nothing about Bel Covo’s internal cost structure: a share of 22% epoxy / 28% polished is the same kind of fact as the project-size-band shares already shipped on epoxy-floor-cost-2026.mdx (line 119: “76.8% of accepted Oklahoma quotes are under 2,000 square feet”). Pair with cost_per_sqft in the same row when the dual-axis insight is the point.
  • 10-year cost-of-ownership totals computed from permitted cut-shape inputs, when (a) every numeric input is itself a permitted cut, (b) every failure-rate, lifespan, recoat-frequency, or warranty-claim assumption cites an external authority (ASTM, ICRI, manufacturer doc), and (c) the council pre-clears the pillar before drafting. The output is a multi-year total cost, not a decomposition; Rule 17 stays clean. The recoat-only multiplier in any cost-of-ownership formula must come from external manufacturer documentation, not Bel Covo records, since the cleaned legacy dataset does not distinguish full-install from recoat jobs.

Forbidden cut shapes (do NOT publish, regardless of how the request is framed):

  • Any cost-component share of the total quoted price: labor_share, materials_share, supplies_share, profit_share, cost_breakdown, markup_share, margin_share, or any synonymous metric an agent might invent (cost_component_share, direct_cost_share, overhead_share, etc.).
  • Per-region or per-service-code cost decomposition tables.
  • Comparative cost decomposition across service codes.
  • Raw per-record cost-component fields (materials_quoted, supplies_quoted, labor_quoted, cost_quoted, profit_quoted) in any committed JSON. The cleaned legacy dataset omits these by design (scripts/research/ingest-legacy-quotes.mjs ALLOWED_CLEANED_KEYS).

The forbidden list applies to all populations (consented and historical), all geographic granularities, and all surfaces (research pillar, cost page, blog post, JSON-LD structured data, downloadable dataset). Read the canonical rationale, prohibited tokens, permitted alternatives, and council reference in docs/comms-rules.md Rule 17. The mechanical enforcement gate is security-review.sh Rule 11.

If a future research question genuinely needs a cut shape outside the permitted list, route the question through /council before drafting; the council ruling is precedent-setting for every pillar (Prioritizer flag, 2026-04-26).

Citation format

Citation style for research pages inherits from the Bel Covo house style. Concretely, every external authority cited on a pillar page must satisfy three requirements:

  1. The authority is registered in our internal authority registry. This is the canonical list of trusted sources (BLS, Census, NAHB, IBISWorld, manufacturer spec sheets, peer-reviewed journals, etc.). Our team adds new authorities by amending the registry, not by adding ad-hoc URLs to a published page.
  2. The cited URL is verified. Before we publish, we visit each URL referenced in the page and confirm it returns the cited content from the registered authority. A page referencing a 404 or a redirected URL does not ship.
  3. The citation uses the registry’s canonical citation format. The registry defines, per authority, how the citation should be rendered inline (publication name, year, optional report title). This keeps citation style consistent across pillars.

Inline references on a pillar page use two patterns:

  • Bel Covo data: a citation card showing the value, the sample size, the method, and the source identifier. The card pulls from the locked snapshot for that page so the value, n, and method on the card always match the table.
  • External authority: a direct link with anchor text matching the registry’s canonical citation format. The link target is the verified URL.

Cost pages, location pages, and blog posts elsewhere on the site that quote a research number use a third pattern: a citation card that links up to the pillar where the methodology is documented in full. This pattern is what closes the citation loop for downstream readers.

Freshness policy

Default cadence: annual refresh

Every pillar page is refreshed at minimum once per year. The refresh produces a new dated version (slug pattern: <topic-name>-<year>), bumps the snapshot freeze date, and the old slug 301-redirects to the new. The previous version is linked from the new version in a “previous version” footer so an external reader can audit what changed.

Triggered refresh: > 15% movement

The refresh cadence is interrupted whenever any metric on a published page moves more than 15% from its prior published value. The 15% threshold is calibrated so that ordinary noise (sample drift, seasonal mix) does not trigger a refresh, but a real shift in the underlying market does.

When a triggered refresh fires:

  1. We freeze a fresh snapshot with a new freeze date.
  2. We create a new dated page (<topic-name>-<year> or, for an intra-year refresh, <topic-name>-<year>-<month>).
  3. The old slug 301-redirects to the new slug.
  4. The new version’s footer links the old version with a brief note describing what moved.
  5. The methodology version is not bumped unless the underlying methodology rules also changed.

Triggered refresh is the mechanism that prevents “fresh stats that silently change.” Citations from journalists and AI engines point at a dated slug. That slug never silently updates its numbers. If the data moves materially, the old slug stays frozen and redirects forward; the new slug carries the new numbers under a new URL.

Methodology versioning

The methodology version (v1, v2, …) is a separate axis from the dated freeze date. We bump it only when the rules in this document change in a way that would invalidate prior numbers.

Bumps the methodology version

  • Changing the n-floor (e.g., raising the default from 30 to 50, or formalizing a per-cut override pattern that did not exist before).
  • Changing the definition of what counts as a Bel Covo quote (e.g., narrowing the consent flag, redefining “abandoned”).
  • Changing the anonymization floor (e.g., moving the geographic floor from ZIP3 to county, or relaxing it to ZIP5; the latter is forbidden by the current rules but the version axis exists in case a future rule change is justified).
  • Changing the citation format requirements in a way that retroactively invalidates prior citations.
  • Adding a new population (e.g., admitting historical pre-consent records as a permitted source under documented constraints).

Does not bump the methodology version

  • Re-pulling fresh data on the existing rules.
  • Adding a new authority to the citation registry.
  • Adding a new metric that was not previously published.
  • Adding a new region cut that was previously below the n-floor and is now above.
  • Re-running ingestion against the same historical records under the same v2 rules (ingest is data refresh, not rule change).

Each pillar page declares its methodology version in plain sight (the trust strip at the top of the page surfaces the version). A reader comparing two pillars at different methodology versions can cross-reference this document’s git history to see which rules differ.

Snapshot reference

Behind each pillar page is a structured reference document holding the aggregates and verified citations that the page renders against. We freeze the reference on the page’s freeze date and keep it available on request for any auditor who wants to verify a number against its underlying record. The reference shape is stable across pages so a reviewer auditing two pillars can apply the same checks to both.

What this is not

  • Not the press kit. A press kit is a marketing surface listing what’s available for journalists. Our pillar pages do that job by being citable on their own. If we ever ship a press kit, it links to pillars, not the other way around.
  • Not a guarantee of national applicability. Our numbers come from our quote records. Aggregates are most defensible inside the regions where we hold a critical mass of quotes (the floor enforces that automatically). National numbers are computed only when the region rolls all the way up; for any narrower regional question, we publish the cut at the level where the floor holds.
  • Not a substitute for consulting a local contractor. Our published medians describe the market we observe. They are not a quote on any individual project, and no number on a research page commits us or any other contractor to a specific price.

The rules above are the only thing standing between Bel Covo’s research pages and the failure modes that kill data publishers: noisy stats, re-identifiable cuts, broken citations, and silent post-citation drift. Every rule maps to a failure mode. Every rule is enforced. That is the whole document.