By Alex Host · Founder of Top Care Cleaning · Updated 2026-05-04

Hosted Reviews A/B test configuration screen — two SMS template variants side by side

Test one variable at a time — start with message length (short vs long), then personalization (tech name vs no tech name), then timing. As of 2026-05-04, Top Care Cleaning is actively running copy tests — interim results are noted below with the caveat that sample sizes are small and not yet statistically conclusive.


Why A/B testing review request copy matters

The baseline: 21% review-to-send at Top Care

Top Care's current template converts at 21% review-to-send (n=70 sends through 2026-05-04). That number is the result of the template we are using right now — one specific message structure, one set of merge tags, one ask format.

The question A/B testing addresses is: is this the best we can do, or is there a variation that performs meaningfully better?

A 5-percentage-point improvement from 21% to 26% does not sound large in isolation. In practice, for a business running 20 jobs per week, that difference is roughly 1 additional review per week — approximately 50 additional Google reviews per year. Compounded over three years, a single copy optimization can produce a materially different review count, which translates to local search visibility and inbound calls.

What a 5-percentage-point improvement means over 12 months

At 20 jobs per week (1,040 per year):

For a business that generates 3–5 new customers per 100 Google reviews (a reasonable local-market estimate), that difference translates to 1–3 additional customers per month from organic map pack traffic. The math favors systematic testing.

The caveat: sample sizes at Top Care

n=70 sends is not a large enough sample for statistically significant A/B testing. A valid A/B test for a binary conversion outcome (review or no review) typically requires several hundred sends per variant to detect a 5-percentage-point difference with 80% statistical power. At 20 jobs per week, Top Care needs roughly 6 months of sends per variant to reach meaningful sample sizes.

This is why this article is in the expansion cohort — we needed time for A/B test data to accumulate before this content could be based on something real rather than hypothetical.


What variables to test (in priority order)

Variable 1 — Message length (one sentence vs two sentences)

The Top Care default template is one sentence plus a link. The hypothesis for a longer version: adding a specific reference to the service or the tech's name in a second sentence increases personalization and conversion.

Variant A (current): "Hi {name}, thanks for having us out today. Quick Google review if you're happy: {link} — {tech}"
Variant B (longer): "Hi {name}, great having {tech} at your place today. If everything looked great, we'd love a quick Google review: {link}. Thanks so much, {tech}"

The directional expectation: short templates typically outperform longer ones in SMS because the ask is clearer and the link is more prominent. But this is a hypothesis worth testing, not a certainty.

Variable 2 — Personalization level (tech name vs no tech name)

The Top Care template includes the tech name twice (in the body and the signature). A stripped-down version removes the tech name entirely and signs from the business.

Variant A (with tech name): "Hi {name}, thanks for having us out. Quick Google review if you're happy: {link} — {tech}, Top Care"
Variant B (without tech name): "Hi {name}, thanks for using Top Care. Quick Google review if you're happy: {link} — Top Care"

The hypothesis: tech name matters most in businesses where the customer has a strong single-tech relationship (regular recurring clients). For one-time or infrequent jobs, the business name may be more relevant than the specific tech's name.

Variable 3 — The ask phrasing

Different ways to frame the ask:

Variant A: "We'd love a quick Google review"
Variant B: "Would you mind leaving us a quick Google review?"
Variant C: "Mind leaving a 2-minute Google review?"

The phrasing variants test whether a softer, permission-asking frame ("would you mind") outperforms a direct request ("we'd love"). Industry email data suggests permission framing slightly increases response rates, but whether this holds in SMS is an open question.

Variable 4 — Send timing (Day 1 vs Day 2)

Within the 24–48-hour window, does sending on Day 1 (same day as the job) vs Day 2 (next morning) produce meaningfully different results?

Top Care's timing data shows morning sends (6am–12pm) at 33% review rate, which suggests the next-morning send may outperform a same-day afternoon send. But we have not isolated "same-day" vs "next-morning" in a controlled test.

Reference: The 24–48 Hour Window: Why Timing Your Review Request Wins.


How to structure a valid A/B test for review requests

Minimum sample size per variant

For a binary outcome (review or no review) with a baseline conversion rate of 21%, to detect a 5-percentage-point lift (21% to 26%) with 80% statistical power and 95% confidence, you need approximately 400–500 sends per variant.

At 20 jobs per week, that is roughly 5–6 months of data per variant. This is why most small operators should think of A/B testing as a medium-term project, not a quick experiment.

A rough guide:

For help calculating exact sample sizes, use a publicly available A/B test significance calculator (e.g., from a statistics reference source) before starting — it helps set realistic expectations for when results will be conclusive.

How to isolate one variable at a time

Each A/B test should change exactly one thing. If you change both message length and phrasing simultaneously, you cannot know which change drove any difference in conversion.

Testing order recommendation:

  1. Message length (most impactful, easiest to isolate)
  2. Personalization level (tech name)
  3. Ask phrasing
  4. Send timing

Run tests sequentially, not simultaneously, unless you have very high job volume (100+ per week) that allows multiple concurrent tests across non-overlapping customer segments.

How to track results in Hosted Reviews

In Hosted Reviews, the A/B test feature splits sends between two template variants at a configurable ratio (50/50 by default). The dashboard tracks review-to-send rate per variant, click-through rate per variant, and time-to-review. Results accumulate in real time.

The key metric to watch is the review-to-send rate, not just the click-through rate — a variant may generate higher clicks but lower completed reviews if the review form experience creates friction.


Interim results from Top Care's tests

Test in Progress — update before publishing this article.

As of 2026-05-04, Top Care Cleaning has not yet run a completed A/B copy test cycle. The baseline conversion rate (21% review-to-send, n=70) is established. Planned first test: short vs long message (Variable 1), estimated start [date TBD], expected sample completion at n=200 per variant after approximately [X weeks at current job volume].

When A/B data is available, this section will be updated with:

Do not publish this article without real A/B test data in this section. Publishing with this section blank or labeled "coming soon" produces an article that is weaker than any competitor writing about copy testing — the entire moat of this article is "we actually ran the test."


What to do with your test results

How to read a conversion rate difference as meaningful vs noise

A difference of 22% vs 24% in review-to-send rate over 50 sends per variant is noise. A difference of 19% vs 27% over 400 sends per variant is signal. The practical test: would you change your default template based on this result if it held up over the next 6 months? If yes, the difference is large enough to act on. If you are not sure, run more sends before changing the default.

When to update your default template

When a test variant achieves statistical significance (or a sample size large enough to be directionally confident), update the default template in Hosted Reviews and retire the test. Do not run indefinite tests — at some point you are just varying the experience for customers who could all be receiving the better-performing version.

For the full current template library and how to update it, see Review Request Templates: SMS, Email, and In-Person Scripts That Work.

When to run the next test

After you have updated the default template with the winner, wait at least 90 days before starting the next test. This gives you a new baseline conversion rate to compare against. If you test too quickly in succession, you lose track of whether the most recent change improved things or whether an earlier change was responsible.


Testing reminder cadence as an A/B variable

The reminder timing (Day 3 vs Day 5) is a testable variable separate from copy testing. You can run a reminder cadence test while keeping the message copy constant.

The question: does a Day 3 reminder recover more customers than Day 5, or does it feel too soon?

For the full reminder cadence framework, see Reminder Cadence for Review Requests: How Many Follow-Ups, How Far Apart.


A/B testing and tech attribution

Per-tech performance data is a natural A/B lens: different technicians at Top Care produce different review conversion rates. If Tech A consistently converts at 30% and Tech B converts at 12%, one question is whether the difference is service quality or message quality. Running the same template for both techs and comparing conversion rates is an informal A/B test that requires no setup beyond what you already have.

For the full framework on per-tech review attribution and how to use leaderboard data to coach performance, see Tech-Attributed Reviews and Multi-Tech Jobs: How to Credit the Right Technician.


Frequently asked questions

How many review requests do I need for a valid A/B test?

400–500 sends per variant to detect a 5-percentage-point lift with statistical confidence. At 20 jobs per week, that is approximately 5–6 months per variant. Plan accordingly — this is a medium-term project, not a quick experiment.

Can I A/B test the screening message vs direct link send?

Yes — this is a meaningful test: two-step screening flow vs direct link delivery. The trade-off is that the direct link approach exposes unhappy customers to the public review form before you have a chance to route them to service recovery. The two-step approach protects against that at the cost of one extra tap for every customer.

How do I know if a difference in my results is meaningful?

Compare the difference against the sample size. A 22% vs 24% difference with n=50 per variant is within random variation. The same difference with n=400 per variant starts to become meaningful. Use an A/B test significance calculator to assess statistical confidence.

Should I test copy or timing first?

Timing first, then copy. Timing (day of week, hour of day) has a larger documented impact at Top Care (Tuesday 35% vs Wednesday 7%) than any copy variable we have tested. Get your timing right before optimizing copy.


The system that enables this

I built Hosted Reviews to automate this for Top Care Cleaning — and now for other local service businesses. The A/B test feature is part of the platform. 14-day trial, no card required.


About the author

Alex Host runs Top Care Cleaning, a Grand Rapids home services company with 400+ Google reviews, and built Hosted Reviews to automate what he was doing manually. Reviews-facet bio.

I run Top Care Cleaning, a Grand Rapids home services company with 400+ Google reviews, and built Hosted Reviews after manually asking for reviews for years. The data in these articles comes from our own system. — hostedbrands.com/about