Measurement is the first thing a redesign should specify and the last thing it usually gets.
The catch is that success must be defined before the structure, layout or copy exists. That feels premature, but postponing measurement leaves stakeholders debating aesthetics against numbers that no longer mean the same thing.
That omission is why so many redesigns can only be argued about rather than assessed. A website redesign analytics plan is the commercial specification for the work, not an analytics task for launch week.
Broader commercial choices are covered in the redesign decisions B2B teams need to get right. The concern here is narrower: preserving a credible comparison from the old site to the new one.
The website redesign analytics plan starts before design
Success cannot be reverse-engineered from a polished GA4 dashboard after launch. The team must agree what a useful visit, enquiry and sales opportunity mean while those definitions can still shape the build.
Use this four-stage order:
- Measurement contract: define business outcomes, event definitions, data sources and accountable owners.
- Structure: assign each priority audience a route towards one of those outcomes.
- Layout and interactions: make progress through each route observable without mistaking curiosity for completion.
- Copy and visual treatment: form hypotheses about the behaviour each message or design choice should influence.
Tracking is designed before structure, layout and copy because every later decision is judged against it. That discipline belongs inside properly scoped website design and redesign work, not in a ticket added when development is nearly complete.
Architecture determines what deserves measurement. In our work for Lanteria, a broad Microsoft 365 HR platform had to route several stakeholder audiences through different capabilities. AfriCap Hub needed a catalogue, filtering and a registration journey. A single site-wide “conversion rate” would flatten the routes that the architecture was deliberately separating.
The same principle applies to interface work. Event definitions should inform the UX and UI decisions governing each interaction, including what counts as opening, starting, completing or abandoning a task.
Measurement earns priority because every design judgement depends on it.
Baseline capture for redesign analytics
A Friday export of last month’s dashboard is not a baseline.
Record the eight complete Monday-to-Sunday weeks immediately before any build starts, excluding the current partial week. Preserve the preceding eight complete weeks as a comparison window. The recent 8 weeks establish the baseline; the earlier 8 reveal obvious changes in demand, campaign mix or sales activity.
Store weekly values, not just one 8-week total. Every export needs four metadata fields: date range, timezone, applied filters and measurement-definition version.
Capture these nine baseline records:
- Acquisition volume: GA4 sessions and engaged sessions by source, medium and landing page.
- Audience mix: channel, device, geography and new-versus-returning shares.
- Paid-media inputs: Google Ads spend, clicks, campaign and active primary conversion actions.
- Search demand: Search Console clicks, impressions and query groups for priority landing pages.
- Primary outcomes: GA4 key-event counts and rates, with the denominator stated.
- Journey progression: form starts, validation failures, successful submissions, scheduler loads and confirmed bookings.
- Sales-system truth: HubSpot or equivalent CRM records matched through a unique submission or booking identifier.
- Commercial quality: valid enquiries, sales-qualified opportunities and the agreed rejection reasons.
- Data quality: consent acceptance, Consent Mode configuration, untagged pages and known internal or spam traffic.
A partial baseline is worse than none because it creates false symmetry. No baseline forces an honest admission of uncertainty; an incomplete one makes two differently defined numbers look comparable.
Apply three stop rules before design approval:
- If any primary outcome lacks an exact trigger definition or accountable CRM owner, pause approval and assign both.
- If GA4 successful submissions differ from CRM-created records by more than 15% in any 2 of the 8 weeks, reconcile duplicates, spam and integration failures first.
- If the baseline contains fewer than 30 CRM-valid enquiries, do not set a percentage-improvement target; extend collection to 16 weeks or use a count-based guardrail.
The measurement plan evaluates a chosen commercial journey. Deciding the offer ladder and site-wide acquisition model belongs in building a B2B website as a lead-generation system.
A complete baseline exposes uncertainty before design disguises it.
Event parity is a launch gate
A familiar event name proves nothing when its trigger has changed.
“Form submission” might previously mean a button click but now mean a confirmed server response. Treating those as one metric creates an apparent before-and-after change even when user behaviour is identical.
Adapt this six-row parity template to the site:
| Old event | New event | Matching definition | Verified by whom |
|---|---|---|---|
| form_start | lead_form_start | First interaction with a required field, once per form instance | Analyst and UX designer |
| form_submit | lead_form_success | Confirmed success response, once per unique submission ID; never the button click | Analyst, developer and CRM owner |
| scheduler_open | booking_widget_loaded | Cal.com or equivalent booking interface visibly loads | Analyst and developer |
| meeting_booked | meeting_confirmed | Confirmed booking with a unique booking ID written to the CRM | Marketing operations and sales operations |
| phone_click | contact_phone_click | Activation of the published tel: link on a priority page | QA lead and analyst |
| download_submit | asset_access_confirmed | Gated form succeeds and the asset is displayed or sent | Marketing lead and analyst |
Event parity requires three launch-gate checks:
- Trigger equivalence: the old and new actions represent the same completed behaviour.
- Payload parity: identifiers, page context, campaign data and consent state remain available.
- Destination reconciliation: one controlled action appears once in Google Tag Manager, once in GA4 and once in the relevant CRM or booking system.
A deliberately retired event needs a signed retirement decision. It cannot simply disappear.
One unexplained missing or duplicated priority event blocks launch. For each primary form, run 20 controlled submissions across supported device and consent states; any unexplained failure among those 20 returns the event to development.
No event parity means no credible before-and-after comparison.
Post-launch rules for the website redesign analytics plan
Launch-week dashboards mix staff testing, campaign changes and genuine visitor behaviour.
Do not judge the redesign inside the first two weeks. A 14-day snapshot is an instrumentation check; the full initial monitoring window is 42 days.
Use three monitoring phases:
- Days 1–14 — validation: reconcile events, CRM records, consent states and internal-traffic exclusions.
- Days 15–28 — stabilisation: inspect journeys by channel and device, but make no redesign verdict.
- Days 29–42 — assessment: compare matched audiences and traffic sources against the 8-week baseline.
Apply these four trigger-action rules:
- If a priority event differs from its source system by more than 10% for 2 consecutive days, declare an instrumentation incident and suspend performance conclusions.
- If a channel’s session share changes by at least 10 percentage points from baseline, segment the comparison by channel rather than crediting or blaming the redesign.
- If device-specific form completion falls at least 15% after 200 form starts on that device, inspect field behaviour, validation and interaction design.
- If the CRM-valid enquiry rate is at least 20% below baseline after 30 enquiries, investigate message, qualification fields and routing. With fewer than 30 by day 42, extend observation to day 70.
These are operating thresholds, not claims of statistical significance. They force a named action while preventing every small movement from becoming a design opinion.
When a decline is isolated to one commercial page with stable acquisition and sound tracking, use the service-page optimisation checks for proof, specificity and action rather than quietly changing the event definition.
Early data tests instrumentation; later data tests the commercial system.
If this analysis exposes wider gaps in your site, we can turn the evidence into a focused redesign brief — book a call
What does not prove a redesign worked
Green arrows remain persuasive even when their denominators have changed.
Four popular signals fail to establish a commercial outcome:
- More pageviews: additional views can come from increased media spend, irrelevant traffic or visitors struggling to find information.
- More CTA clicks: a click measures intent at one moment. It does not establish a successful submission, valid enquiry or sales opportunity.
- A cleaner dashboard: better colours and tidier charts cannot repair duplicate events, missing consent states or inconsistent CRM definitions.
- Seven days before versus seven days after: a 7-day window versus another 7-day window is too exposed to campaign timing, weekday mix and low event volume.
Heatmaps and session recordings can diagnose friction, but they cannot replace commercial outcome data. A dramatic scroll pattern has no decision value until it is connected to a defined audience, page purpose and downstream event.
Vanity metrics cannot settle a commercial redesign decision.
What redesign analytics can and cannot prove
A redesign changes several variables in one release: navigation, copy, visual hierarchy, forms, technical behaviour and tracking implementation.
Five common confounds are seasonality, campaign mix, sales response speed, consent acceptance and concurrent pricing changes. The honest limit here is that post-launch numbers can show how the new commercial system performed under observed conditions, but they cannot isolate which design variable caused the movement.
Take an illustrative UK consultancy with these five labelled monthly inputs:
- Google Ads spend: £8,000
- GA4 successful-form events: 32
- CRM-valid enquiries: 20
- Sales-qualified opportunities: 8
- Average opportunity value: £25,000
The arithmetic is:
Reported web CPL = £8,000 spend ÷ 32 GA4 events = £250
CRM-valid enquiry CPL = £8,000 spend ÷ 20 valid enquiries = £400
Opportunity acquisition cost = £8,000 spend ÷ 8 opportunities = £1,000
Illustrative pipeline = 8 opportunities × £25,000 average value = £200,000
Continue the illustrative example after launch. Assume spend remains £8,000, GA4 events rise from 32 before to 40 after, and CRM-valid enquiries remain 20 before versus 20 after.
Post-launch reported web CPL = £8,000 ÷ 40 GA4 events = £200
Post-launch CRM-valid enquiry CPL = £8,000 ÷ 20 valid enquiries = £400
The dashboard comparison reports £250 before versus £200 after. The commercial comparison remains £400 before versus £400 after. Calling that a redesign win would reward measurement volume rather than lead quality.
Our falsifiable claim is that a 32-to-40 GA4 increase alongside a 20-to-20 CRM result indicates measurement drift rather than stronger demand; 40 unique, accepted CRM records matched to those 40 events would prove it wrong.
Where stakeholders believe stronger brand expression influenced behaviour, record that as a hypothesis using a defined brand-personality framework, not as a causal conclusion.
Redesign analytics supports decisions, not clean causal attribution.
FAQ
Which system should own conversion truth?
Assign three distinct decisions. GA4 owns on-site behaviour, the CRM owns accepted enquiries and opportunities, and Google Ads owns bidding feedback. If a 7-day unique-ID reconciliation differs by more than 10%, remove the suspect action from Google Ads’ primary conversion set until corrected.
How should consent changes be handled?
Keep observed and modelled measurements identifiable when using Consent Mode. If consent acceptance differs by more than 5 percentage points between comparison windows, annotate the break and compare equivalent consent cohorts before issuing a verdict.
When is a segment too small to assess separately?
Use 20 CRM-valid outcomes as a practical operating floor. Below 20, extend the window or combine segments only when they represent the same commercial audience; do not hide a weak sample inside an unqualified site-wide rate.
Should Google Ads campaigns remain unchanged after launch?
Keep budgets, locations, match types and primary conversion actions stable for 14 days when commercially safe. If invalid leads exceed 50% across at least 10 enquiries, protect spend immediately, correct targeting or qualification, and restart the observation window.
Clear ownership keeps analytics useful after launch.
Summary
- Approve the measurement contract before structure, layout or copy receives sign-off.
- Capture 8 complete baseline weeks and preserve the preceding 8 as a comparator.
- Block launch when even 1 priority event has an unexplained parity failure.
- Use days 1–14 for validation and days 29–42 for the first assessment.
- When GA4 and CRM differ by over 15% in 2 weeks, fix measurement before judging design.
Measurement discipline turns redesign opinion into an accountable commercial decision.
Actualyse designs and rebuilds B2B websites that turn research visits into qualified pipeline. Book a call to talk through where yours stands.

