Lead Form Spam: Filter the Junk Without Losing Real Buyers

By the Follow Up Ace team· Last updated
Quick answer

A honeypot field and a submit-speed timer stop scripts, not sentences. Gibberish, test data and link spam arrive as well-formed leads, get assigned, and set off a speed-to-lead alert. Filter on what the submission says, hold rather than block, never judge a name or an email domain, and give every held lead a queue and a one-click release.

A long conveyor belt of pale envelopes running toward a lit archway; on a raised coral ledge alongside it three envelopes stand apart from the flow, one of them crumpled, with a small brass gate beside them standing open onto the belt
Three set aside on the ledge, and the gate back onto the belt standing open.

Two sign-ins come off the same open-house iPad ninety seconds apart. One is a name, a mobile number and “is the Oak Street place still available?” The other is asdkjh typed into every field. The form accepted both. The CRM assigned both to the same agent. That agent’s phone buzzed twice.

Both rows are illustrations rather than screenshots, but the shapes are the ones public capture forms really receive. What is not made up is the path between that iPad and that phone: on most setups, nothing standing in it can read either submission.

Your bot gate is a metal detector, not a reader

Nearly every hosted capture form ships with the same two defenses. A honeypot is a hidden input a person never sees and never fills; a script that fills every field on the page fills that one too and convicts itself. A render-to-submit timer throws out anything returned faster than a human could plausibly type it.

Both are good, and both are mechanical. They measure how a submission was made, never what it says, which is why all of this walks straight through:

What arrivesHoneypotTimerField validation
asdkjh in every boxclearsclearspasses
“test test” / [email protected] / “testing 123”clearsclearspasses
A backlink pitch pasted into the message boxclearsclearspasses
Abuse aimed at the agentclearsclearspasses
Answers that contradict each otherclearsclearspasses

On a typical setup each of those becomes a person in your CRM, routed to whoever is up next, and pushed to a phone at the urgency you reserve for a live buyer.

Automated traffic accounted for more than 53% of all web traffic in 2025, up from 51% the year before, per the 2026 Imperva Bad Bot Report. That describes the whole web, not your sign-in page — read it as why a public form gets found at all, not as an estimate of your own junk rate. Yours is something to go measure, and there is a way to do that below.

Which surface gets hit is predictable. A QR-code sign-in used at one open house is a small target; an indexed public profile page with an open message box is what link-spam runs actually find.

“Block” is the wrong verb for an inbound lead

A brass balance scale on a dark cloth-covered table by a window. The raised left pan holds a single blank paper tag; the lowered right pan holds an iron house key with a coral ribbon tied through it
A held junk row costs somebody a minute. A dropped buyer costs the whole client, and you never find out.

Spam filtering borrowed its vocabulary from email, where both mistakes cost about the same and either one is fixable in a folder. Lead capture is not like that. A junk row costs an agent a few seconds. A real buyer refused at the form is gone silently, and the only evidence is a listing somebody else sells.

So the rule worth adopting is about reversibility rather than accuracy. Let a rule drop a submission only where a false positive is close to impossible — a hidden field a human cannot see qualifies. Anything that has to interpret a sentence belongs in a reversible lane: store the lead, hold back the parts that spend someone’s attention, and let a person wave it through.

That cuts both ways in our own stack, and it is worth saying plainly. A honeypot conviction on an Ace capture form does discard the submission, and the visitor still sees an ordinary success screen. That is the right trade for a rule almost never wrong about a human, and the wrong trade for anything that has to read meaning.

The line you do not cross: judge the text, never the person

On 2 May 2024, HUD issued Fair Housing Act guidance on the use of AI in tenant screening and housing advertising. It states that the “use of third-party screening companies, including those that use artificial intelligence or other advanced technologies, must comply with the Fair Housing Act,” and that screening should be “transparent, accurate, and fair” (HUD No. 24-098).

That guidance covers screening and ad delivery, not lead capture, and nobody should quote it as though it did. The principle underneath it travels anyway: an automated step deciding which housing enquiries reach a human is not neutral plumbing, and a filter that quietly holds back people with unfamiliar names is a Fair Housing problem wearing a spam filter’s clothes.

Four rules keep a junk filter on the right side of that line:

Homegrown rules usually fail right here, because the easy heuristics are all identity heuristics: a blocked country-code domain, a regex on name shape, a list of “suspicious” first names somebody added after a bad week. Each keys on who the person appears to be instead of what they wrote. Your outbound copy has the mirror-image problem, which we worked through in the ten-line Fair Housing filter test.

The real cost is the notification nobody trusts anymore

A dirty database is the obvious cost and the smaller one; cleaning up is a known job with a known cure, from merging duplicates to archiving what is dead. The expensive cost is what repeated worthless alerts do to the people receiving them.

The best evidence on that comes from hospitals. Studying 112 ambulatory primary care clinicians, Ancker and colleagues found the likelihood of a clinical practice reminder being accepted dropped by about 30% for each additional reminder received in the same encounter, and that repeated alerts rather than sheer workload predicted the decline (Effects of workload, work complexity, and repeated alerts on alert fatigue, BMC Medical Informatics and Decision Making, 2017). A review of 23 studies put average override rates for computerized physician-order-entry alerts between 46.2% and 96.2% (Poly et al., JMIR Medical Informatics, 2020).

Those are clinicians and drug alerts, not agents and leads, and nobody has run the equivalent study on speed-to-lead notifications. Read it as a mechanism, not a number you can apply to your team. The mechanism carries: people stop answering a channel in proportion to how often it has wasted them. Your new-lead push buys minutes only while the agent still believes it. That is what junk is really spending, and why which events earn an alert and which channel they arrive on deserve more thought than they get.

A twenty-minute audit of your own forms

Do this from a phone on cellular data, not the desk where you built the form. Send six submissions to your own live capture surfaces and, for each, write down three things: did it reach the CRM, was it assigned, did a notification fire.

#Send thisWhere it should land
1Gibberish in every field, typed slowlyNot assigned, no agent notification
2“Test” as the name, a throwaway address, “just testing” as the messageNot assigned, no agent notification
3A link-building pitch in the free-text boxNot assigned, no agent notification
4A real name, a real question about a real addressAssigned and notified, fast
5A real name written in a non-Latin script, plus a plain questionAssigned and notified, fast
6A four-word question with two typos in itAssigned and notified, fast

Rows 1 to 3 tell you whether anything reads content. Rows 4 to 6 matter more, because a filter that fails those is worse than no filter at all. If 5 or 6 lands anywhere other than straight with the agent, fix that before tuning anything else.

Then count. Pull your last 30 days of form captures, mark every row no agent could have worked, and set that against the new-lead notifications your team received in the same period. That ratio is what you are managing. Two questions finish the audit: where do junk rows go today, and can you get one back when the answer was wrong? “Deleted” is a decision nobody can revisit.

What Ace does with a junk sign-in

Ace reads each submission for coherence on three of its public capture forms — Lead Magnets, Ace Forms and Ace Pages — and on the open-house question box, before anything about it reaches Follow Up Boss or an agent’s phone. Two other public paths, Client Packet referrals and Home Almanac sign-ups, run the honeypot and the timer but no coherence read, so that step is not part of those today. A submission that reads as junk is held rather than refused, and both halves of that word are load-bearing.

The lead is still written to your account, still counted, still listed. What waits is everything that spends attention or writes into your CRM: the push to Follow Up Boss, and with it the direct assignment, the follow-up task and the action-plan enrollment; the owner notification; the agent’s alert; the referral credit. A held lead was never pushed, so it also never enters your source ROI reporting — a spam run cannot quietly make one of your lead sources look worse than it is.

The visitor sees no difference: same confirmation, same report or Home File if the page promised one. A spammer who cannot tell whether a run worked has nothing to iterate against, and a real person caught by mistake is never made to feel like a suspect.

The Ace captures leads inbox: a table of captured contacts with columns for Who, Contact, Came in through (Ace Form, Lead Magnet, Ace Page or Client Packet, with the specific page name), Assigned agent, In FUB status showing Delivered or Pending, and When
Every capture in one list, whichever tool it came through, with whether it has reached Follow Up Boss. Demo account, not real contacts.

Held rows sit in that same inbox behind a “Held as junk” filter, each showing what it was held for. Releasing one is a single click that runs the exact work the capture deferred — the same code path, not a re-implementation — and records who released it and when. A second click is a no-op, not a duplicate contact.

The judgement is fenced by the rules above. The instruction not to treat an unfamiliar name, script, domain or imperfect English as junk is written into every question the classifier is asked, and a test fails the build if that clause goes missing. A confident read of “this is an ordinary enquiry” overrides a suspicious one, because two signals disagreeing is what an unsure case looks like, and unsure resolves toward keeping the lead.

What a junk filter does not prove

An accuracy percentage means little unless somebody swept the filter against a labeled set of real submissions from surfaces like yours. Ours is tuned from the asymmetry rather than from such a sweep — a deliberate choice, and a real limit on what we can claim. The better question for any vendor, ours included, is not how accurate the filter is. It is what happens to a lead it gets wrong.

A hold is invisible to the visitor by design, so a mistaken hold is invisible too until somebody opens the queue and looks. That is a standing habit, not a solved problem. Put it beside whatever weekly review you already run.

Coherence is not quality. A polite enquiry from somebody who will never transact is a real submission and goes straight through. Separating serious from curious is a different job with different tools — scoring, tagging, deciding which contacts your automations can reach — and conflating the two is how teams end up suppressing real buyers. If you sort with tags today, the traps in tag-based exclusion come first.

Common questions

Is a hold the same as marking a lead as spam?

No. Marking as spam is a verdict, and usually a deletion. A hold is a deferral: the record exists, nothing was thrown away, and no human was interrupted yet. That distinction decides what a mistake costs you.

Will a junk filter stop bots from submitting my form?

It will not, and neither will a honeypot. Both stop what happens next. Reducing submission volume means rate limits, edge bot protection and CAPTCHAs, each of which also costs you real people. Filtering after the fact is the layer that does not cost you conversions.

Can I just block disposable email domains?

Use it as a supporting signal, never the whole reason. Plenty of real buyers put a secondary address on a public form precisely because they do not want a hundred calls, and a domain you have never heard of usually means nothing more than that.

What should I do with held leads that turn out to be real?

Release them, then look at what they had in common. Two or three false holds pointing the same direction is worth more than any accuracy number, because it tells you which shape of real submission your filter reads badly.

Does this apply to leads from Zillow or Realtor.com?

Not in the same way. Portal leads arrive through an authenticated integration rather than an open public form, so the exposure is different. This is about surfaces you host yourself: sign-in pages, home-value pages, public profile pages, anything with a URL somebody can find.

Try Follow Up Ace in your Follow Up Boss

Free to start, no sales call. Connect Follow Up Boss in one click and Ace works inside your CRM.

Get Started Free