Real Estate Follow-Up Tests: 5 Checks Before You Believe It
Before switching your team's follow-up script, check five things: that you counted a real outcome instead of a send, changed one variable, counted each prospect once, waited out the full response window, and put a confidence interval around each rate. If the two intervals overlap, the test decided nothing.
Eleven leads got the short opener and four wrote back. Nine got the long one and two wrote back. By Friday somebody has quietly retired the long version.
Run those twenty contacts through Fisher's exact test and the two-sided p-value comes out at 0.64. That is the kind of gap twenty coin flips hand you for free. But 36% and 22% look like different numbers on a whiteboard, so the script changes — and every new lead the team touches now gets the opener that won a coin flip.
Testing your follow-up is worth doing. The results just get read wrong in five specific ways, and the five compound. Here they are, in the order they bite.
Start from the assumption that your change did nothing
Ron Berman and Christophe Van den Bulte examined 4,964 effects from 2,766 experiments run on a commercial A/B testing platform. Using three methods, they estimate the false discovery rate — the share of statistically significant results that are actually null — at between 18% and 25% for tests conducted at 5% significance. Their own statement of the implication: decision makers should expect one in five interventions achieving significance at 5% confidence to be ineffective when deployed in the field. They attribute those rates mostly to the high fraction of true null effects, about 70%, rather than to low power (Berman and Van den Bulte, Management Science, 2022).
That is a platform where people run deliberate experiments at real sample sizes, and roughly one winner in five is a mirage. A team eyeballing twenty contacts on a Friday afternoon is not beating that rate.
Before anyone over-applies it: that study is about website experiments, not real estate follow-up. It does not predict your results. What it does say is that "no difference" is the normal outcome of a messaging change, which means your reading of the numbers has to be strong enough to survive that being true.
Check 1: A send is not an outcome, and a reply barely is
The number your tooling volunteers is sends, because sends are the one thing it knows for certain. Delivery is a fact about your software. It is not a fact about the lead.
Replies are better and still weak. "Take me off this list" is a reply. So is "wrong number," so is an out-of-office, and so is an angry one. If your denominator counts sends and your numerator counts inbound messages, a script that annoys people will beat a script that works.
Pick an outcome that only happens when something went right: the contact asked to speak with someone, or asked for a time. It is rarer, which makes every other check on this list harder to satisfy — that is the point. And before you count anything, confirm the activity you are counting is actually landing in your CRM, because not every channel logs the same way inside Follow Up Boss.
Check 2: Change one thing
Most team "tests" change the length, the tone and the offer in the same rewrite, then compare the new script to the old one. Whatever the result, it cannot be assigned to any of the three.
Hold everything else still. Same workflow, same channel, same lead source. A text and an email are not two versions of a message — they are two different experiments, and picking the channel is its own decision. Comparing a warm tone against a direct one means keeping both at the same length. Comparing a short message against a long one means keeping the tone fixed.
If you want a concrete place to start, take one of your existing follow-up email templates and cut it to half its word count, changing nothing else. That is a test. Rewriting it from scratch is not.
Check 3: One prospect, one vote, first touch only
A lead who got six messages is not six data points. A lead who re-enrolled in the same campaign twice is not two. And if a contact replied on touch five, the credit does not belong to touch one just because touch one is the version you are testing.
Count each prospect once, on their first outreach in that workflow and channel, and ignore everything after. It shrinks your numbers dramatically, which is the honest thing for it to do — the inflated version was counting the same person repeatedly and calling it evidence.
Check 4: Nobody counts until their week is up
This one silently ruins more comparisons than any other, because it looks like diligence. You add yesterday's leads to the tally. Yesterday's leads have had one day to respond; last month's leads had thirty. The group holding more recent contacts looks worse, and you conclude its script is worse.
Fix the response window before you count — seven days is a reasonable default for first outreach — and exclude anyone whose window has not closed. They are not a zero. They are not yet anything. The same discipline applies when you are watching for a lead going quiet: silence on day two and silence on day twenty mean different things.
Check 5: Do the interval, not the percentage
A percentage from a small group is a point estimate pretending to be a fact. What you need is the range the true rate plausibly sits in, and then whether the two ranges touch.
Use the Wilson score interval rather than the textbook proportion ± z × standard error. Agresti and Coull (1998) showed that the familiar Wald interval's coverage runs too low while Wilson's score interval holds coverage close to the nominal level even for very small samples; Brown, Cai and DasGupta, reviewing the problem again, recommend the Wilson interval for small n and note that common textbook prescriptions about the Wald interval's safety are misleading and cannot be trusted (Brown, Cai and DasGupta, Statistical Science, 2001). Every spreadsheet program can do the arithmetic; so can any of the free calculators.
Here is what it looks like on numbers of the size real teams actually have. These are worked examples, not results from any account — the counts are invented, the arithmetic is not.
| What you saw | Rates | 95% Wilson intervals | Verdict |
|---|---|---|---|
| 4 of 11 asked to talk, vs 2 of 9 | 36% vs 22% | 15–65% vs 6–55% | Not a test. Too few people to compare. |
| 5 of 20, vs 1 of 22 | 25% vs 5% | 11–47% vs 1–22% | Overlap. No decision — even at five times the rate. |
| 12 of 60, vs 5 of 58 | 20% vs 9% | 12–32% vs 4–19% | Overlap. Still no decision. |
| 25 of 60, vs 3 of 60 | 42% vs 5% | 30–54% vs 2–14% | Clear gap. Now you can act. |
Read row two again, because it is the uncomfortable one. Five times the rate, around twenty people in each group, and the honest answer is still "we don't know." That is not the tool being strict for the sake of it. That is what twenty people buys you.
The five-line version
- Outcome: did someone ask to talk, or did software send something?
- One variable: same workflow, same channel, one thing different.
- One vote each: per prospect, first touch only.
- Matured: nobody counts until their response window has closed.
- Intervals: if the two ranges overlap, you have not decided anything.
Pin that next to whatever KPIs your team already tracks. It is worth running by hand once a quarter even if nothing automates it for you.
What it looks like when the tool applies the same rules
We built those five checks into Autopilot's messaging observations, because a tool that quietly relaxes any one of them is worse than no tool — it launders a coin flip into a recommendation.
Inside a single account, over as much of its own last 90 days as it can read completely, Autopilot groups first outreach by workflow, channel, configured style and message-length band. A contact counts only once per workflow and channel, on their first message, and only after seven days have passed. A reply counts only when a real inbound email or text is matched back to that contact and verified against the CRM record. The outcome it actually compares on is narrower than a reply: the contact asked for an appointment, or asked to be put in touch with a person. Messages a human reviewed before they went out sit outside the comparison entirely, so a good agent's edits cannot be scored as the software's win.
Then the thresholds. At least 20 prospects in each group. At least 5 of those requests in the stronger group. Non-overlapping 95% Wilson intervals. Tone comparisons hold the length band fixed; length comparisons hold the tone fixed. Miss any of it and that comparison produces no observation at all — no hedged nudge, no "trending toward."
When it does say something, it hands the writing step a plain-language note along with the caveats: that these are associations rather than causal evidence, that recipient mix and timing can differ between the groups, that a request to talk is not a booked appointment. If your team has set an explicit style, that setting wins — the observation is not allowed to overrule it. It never reads message bodies to do any of this.
What none of this can tell you
Be clear about the ceiling, because the checks buy you less than they look like they do.
This is observational, not a randomized test. Your two groups came from whoever happened to be enrolled, not from a coin flip, so the difference you measure carries every other difference between those people along with it. A style setting is a record of what you configured, not proof of what any individual message actually said. And a comparison that clears every threshold is still a conservative heuristic on your own recent history — not a controlled experiment, and not adjusted for the fact that you are looking at several comparisons at once.
The bigger limit is volume. Row two of that table is the common case, not the edge case: most teams, most quarters, will not generate enough first-touch outreach in a single workflow and channel to separate two intervals. Autopilot's own-account comparison will stay silent for those accounts, which is the correct behavior and also a genuinely unsatisfying product experience. Anyone promising you a messaging insight from every account every month is not applying a threshold.
None of which is an argument for going back to gut feel. Gut feel is what produced the eleven-versus-nine decision at the top of this page. It is an argument for knowing which of your beliefs about your own follow-up are actually supported — and being willing to leave the rest labeled "unknown" until the numbers arrive. If you want the companion piece on auditing what an automation did rather than what it learned, that is here. And if you are still deciding which follow-ups should be automated at all, start there instead.
One question worth asking your team this week: which of your standing follow-up rules — the ones everyone repeats as though they were settled — has anyone actually measured?
Try Follow Up Ace in your Follow Up Boss
Free to start, no sales call. Connect Follow Up Boss in one click and Ace works inside your CRM.
Get Started Free