Method case study · Published 13 August 2026

What a relaunch actually moved.

George-Stone Gardens Limited rebuilt gsgardens.co.uk. We already held reads of the previous site. So rather than ask whether the new one is “better”, we re-ran the same checks and diffed the two.

This is a case study of the method, not a verdict on a firm. The previous site was the baseline, not a failure — it passed eleven of fourteen deterministic checks. What follows is what changed, what did not, what got worse, and the five measures we refused to diff because we could not compare them honestly.

Before
25 Jun / 26 Jul 2026
After
13 Aug 2026
Relaunch
between 3 and 13 August 2026
The baseline

What we actually held, including the thin parts.

A before/after is only as good as its before. Ours is real but uneven: plenty of runs, fewer complete ones, and no prior Continuous Assessment history at all.

20Machine-read scans (operator instrument)· 11 – 25 June 2026

Stored in `reports` joined to `intakes` on website. All 20 are of the previous site. The most recent (25 June, 19:54 UTC) is the one used as the before; the other 19 are earlier runs in the same fortnight and are not separately diffed.

19Consumer reads (the /start instrument)· 11 July – 3 August 2026

Stored in `friday_reports` under company_key gsgardens.co.uk. Only 4 are complete multi-section reads; the other 15 are partial runs of 1–4 sections. The 26 July run is the only one carrying a full deterministic technical audit, so it is the sole technical before.

59Cached page scrapes· 11 June – 3 August 2026

21 distinct URLs in the `scrapes` table. These are the raw page reads behind the reports above, and they are what establishes that the previous site was still live on 3 August.

0Prior Continuous Assessment deltas·

The `assessment_watches` and `assessment_deltas` tables are empty in production. No firm has ever been placed under a watch. This study is the delta engine's first run against real reports rather than fixtures.

Measured

Deterministic checks on the site's own code.

These involve no model. The same page produces the same result every time, so a difference between two reads is a difference in the site. This is the axis a before/after can be trusted on.

Four checks travel inside the machine read and span the whole window — 25 June to 13 August 2026. 2 moved, 2 did not.

CheckBeforeAfterMovement
Structured data (JSON-LD)Before: recognised @types found — Organization; no FAQPage or Review/AggregateRating. After: @types found — LocalBusiness, FAQPage.partialpassimproved
Answerability surfaceBefore: 1 question heading. After: 2 question headings plus an FAQ section, with pricing, process and credentials sections present.partialpassimproved
Contact integrityPhone number, email address and address signal found on both reads. The check tests for presence, not for stability — see the contact note below.passpassno change
Entity clarityBoth reads look for the declared legal name "George-Stone Gardens Limited" in the title or first 500 characters and do not find it. See the measurement artefact note — the underlying position did change even though the check did not.failfailno change

The fourteen-check technical audit

A wider check set over a shorter window (2026-07-26 to 2026-08-13). The audit code is unchanged between the two reads. Score 910 out of 10: 2 improved, 1 regressed, 11 unchanged.

CheckBeforeAfterMovement
Business identity for AI (schema.org)Before: Organization only, no FAQPage or Review/AggregateRating. After: LocalBusiness and FAQPage.warnpassimproved
Answerable content (headings & FAQs)Before: partial signals, 1 question heading. After: 2 question headings plus an FAQ section; pricing, process and credentials sections present.warnpassimproved
Page titleWarn on both reads, but the measured length moved the wrong way: 85 characters before, 96 after, against guidance of 15–65. The status did not change; the underlying number got worse.warnwarnregressed
Site loads cleanlyHTTP 200 on both reads.passpassno change
Secure connection (HTTPS)Served over HTTPS on both reads.passpassno change
Your homepage can be indexedNo noindex tag on either read.passpassno change
Search engines can crawl yourobots.txt permits crawling on both reads.passpassno change
AI assistants can read youAI crawlers not blocked on either read.passpassno change
Sitemap for search enginesA sitemap is published on both reads.passpassno change
Works on mobileMobile viewport set on both reads.passpassno change
Search result descriptionMeta description present on both reads.passpassno change
Canonical URL setCanonical tag present on both reads.passpassno change
Clear business name & locationFirm name found in title or opening content on both reads.passpassno change
Contact details machines can readPhone, email and address signal found on both reads.passpassno change

The published email address changed

The previous site published enquiries@gsgardens.co.uk. The relaunched site publishes info@gsgardens.co.uk, in the page body and in the LocalBusiness JSON-LD. Contact integrity passes on both reads because the check tests that an email is present, not that it is the same one. Phone (01242 234 929) and postcode (GL50 1SU) are unchanged. This is reported as an observation, not a fault — but anything that still cites the old address is now pointing at a different mailbox.

Entity clarity did not move, and the check is why

The check looks for the declared legal name — "George-Stone Gardens Limited" — in the page title. It failed before and it fails now, so the measure reads unchanged. But the titles are not equivalent. Before: "Landscape designers and landscape gardeners in Cheltenham | Home Landscaping Services", which contains no form of the firm's name. After: "Garden Designer Gloucestershire | Contemporary Design & Build | George-Stone Gardens, Cheltenham", which carries the trading name. The site got materially clearer about who it is; the check, which matches on the legal name including "Limited", cannot see it. We report the check as unchanged and the artefact alongside it.

Inferred

What the assistants now conclude, versus before.

Model-derived, so it carries variance the measured axis does not. The prompt and schema behind these figures are byte-identical across the two reads, which is what makes the comparison legitimate; the reading is still not deterministic, which is why the movement is a range.

AI visibility · 0–100

52 76 / 72

+20 to +24· two runs the same day

Reported as a range because two runs the same day returned 76 and 72 on the same page text. The second run also carried location and ICP arguments the first did not, so 4 points is an upper bound on the spread rather than a clean measurement of model noise — but it is enough to say that a single-run before/after on this measure should not be read to the point. The movement here is large enough to survive it.

Answerability across the generated questions

Each run generates five questions from live search demand, then judges whether the site answers them. The two runs share no question wording, so there is no per-question diff — but the distribution over each set is comparable, and it is the thing that moved.

RunFullyPartiallyNot foundOf
Before · 25 Jun 20260415
After · 13 Aug 2026, run 14105
After · 13 Aug 2026, run 23205

Would the assistant name the firm first?

Before: 4 of five, with 1 declined. After: 5of five, none declined — in both runs. On the previous site the assistant declined to recommend the firm on the reviews question — "I cannot confidently recommend based on the scraped pages alone." On the relaunched site it names the firm on all five generated questions, in both runs.

Excluded

The measures we refused to diff.

This register is the method. A measure that only one side produced is not movement, and a naive diff will happily report it as one. Each row below names what a careless before/after would have claimed.

Perception gap (6 dimensions)

No after-measure. The perception-gap leg requires both a stated ICP argument and surviving multi-engine perceptions. The first after-run had the engines but was not given the ICP; the second was given the ICP but our own budget cap blocked the engine calls. Neither produced the register.

Would have claimed: All 6 baseline dimensions resolved — a fabricated win. The delta engine was changed in this branch so it cannot make this claim.

Per-question money-question diff

The measure exists on both sides, but each run generates its five questions from live search demand and the two sets share no wording. Zero questions matched, so there is nothing to diff per question.

Would have claimed: Nothing changed — when in fact the answerability distribution moved from 0 fully answered to 3–4.

Multi-engine per-engine visibility scores

The instrument changed. Commit 57d2c43 (24 July 2026) coerces a phantom competitor slot to none_confident before scoring. This intake carries zero competitors, so every competitor slot is a phantom and the scoring basis is not the same on both sides.

Would have claimed: An average visibility jump from the low 70s to the mid 90s, part of which is the scoring change rather than the site.

The consumer report's narrated sections

No after-measure. Clarity, credibility, the-buyer, who-visits and ai-visibility were not re-run against the relaunched site, so there is no second read to diff.

Search and traffic outcomes

Out of scope and too early. This study covers what is publicly observable about the website plus our own checks against it. Ranking and traffic effects of a relaunch are not visible within ten days, and client analytics are confidential.

Instrument drift

Our own pipeline changed between the two reads.

Seven weeks passed between the baseline and the re-read, and we ship continuously. Some of the instrument moved. Where it did, the affected measure is excluded above; where it did not, we say so and name the check that establishes it.

Machine-read prompt and output schema (src/lib/prompt/)identical

`git diff 9dd1de4 HEAD -- src/lib/prompt/` is empty. The prompt, the system message and the response schema behind visibilityScore and the money questions are byte-identical across the two reads.

Effect: AI visibility and answerability are like-for-like. This is what makes the inferred axis usable at all.

Deterministic technical audit (src/lib/technical/)identical

Last modified 24 July 2026 (58c56dd), two days before the technical baseline was taken, and unmodified since.

Effect: The 14-check audit is like-for-like between 26 July and 13 August.

Page read layer (src/lib/firecrawl.ts)changed

Substantially changed since the baseline — metering, a reworked fallback chain, and a cached-excerpts path. The baseline and the first after-run both recorded 3 pages scraped, so the read scope matched, but the code producing that read did not.

Effect: A real caveat on every model-derived figure: the text the model saw was assembled by different code. It is not a reason to discard the comparison, and it is a reason not to read ±4 points as signal.

Multi-engine read (src/lib/engines/multi-engine.ts)changed

Commit 57d2c43, 24 July 2026 — phantom competitor slots coerced to none_confident before scoring and consensus.

Effect: Per-engine visibility scores excluded from the diff entirely.

Perception gap (src/lib/perception-gap.ts)identical

Unmodified since 11 June 2026.

Effect: Would have been comparable, but no after-measure exists — see the excluded register.

Cost

What the fresh reads cost to take.

$1.42

USD · 2 full machine-read runs

Measured as the increase in the OpenRouter account's lifetime usage figure across the session, which covers both full machine-read runs. It is not split per run: the account figure settles asynchronously, so a reading taken immediately after a run understates it. Treat 1.42 as the total for two runs, accurate to the cent only once settled. The deterministic technical audit costs nothing in model spend — it makes no model calls at all.

The run was budget-capped and the cap engaged. Our metering charges a conservative worst-case ceiling per call before making it; during the second run the global day ceiling reached 1,987p against a 2,000p limit and the multi-engine calls were blocked with no provider call made. That is the cap working, and it is why the second run has no perception gap.

Method & sources

Before: machine read of 25 June 2026 and deterministic technical audit of 26 July 2026, both retrieved from the production `audienceintel` database. After: two machine-read runs and one technical audit taken against the live site on 13 August 2026 through the same code paths. Instrument drift established by `git diff` between the commit that was HEAD at the baseline read (9dd1de4) and the commit the after-runs were taken on. Cost measured from the OpenRouter account's lifetime usage. No client-confidential data — no traffic, no revenue, no private commentary — appears in this study.

gsgardens.co.uk is George-Stone Gardens Limited of Cheltenham, Gloucestershire. It is not Graduate Gardeners (graduategardeners.co.uk), a separate Gloucestershire firm; the two are unrelated and are not conflated anywhere in this study.

Audience Intel · part of Bridge Intelligence · by Cambray Design