You know exactly what kind of customer you want more of. What you do not have is a list of them with working email addresses, and every time you try to make one you end up with companies the wrong size, contacts who left two years ago, and addresses that bounce.
AI lead research is the part of outbound sales that decides whether everything after it works. The email, the follow-up, the call, all of it lands on whoever is on the list. A great list with a mediocre email beats a mediocre list with a great email, every time.
This article explains how the list gets built properly in 2026: defining the fit, where the names come from, what enrichment tools do, how a language model can read a company’s website and tell you whether to write to them, verification, and the do-not-contact list.
What this actually is
Lead research answers three questions about each potential customer before you contact them. Who are they, the company and the specific person? Do they fit, meaning do they have the problem you solve? And why now, meaning is there something specific that makes your message relevant this month?
Enrichment is the industry word for filling in the blanks. You start with a company name and a tool adds the rest: employee count, industry, location, the people who work there with titles, their email addresses, recent hiring. Tools such as Apollo, Clay and ZoomInfo do this by combining public data, purchased databases and pattern guessing.
AI adds a step the enrichment tools cannot do well: reading. A language model (software that reads and writes text, the engine behind Claude, ChatGPT and Gemini) can read a company’s homepage, about page, careers page and news and answer in plain words: “Is this company a fit, and if so, what is the one specific thing we should mention?” The ordinary-business analogy is a sharp junior researcher with a clear brief. They do not decide who to sell to; you do. They find the candidates, read up on each, and hand you a short note per company. Their work is only as good as your brief.
1. Write the fit definition before anything else
Everything downstream depends on a written description of who you want, and “small businesses that need our services” is not one. Write down the industry, the size range, the geography, the role you want to reach, and two or three visible signals that a company has the problem you solve. Then write down who is not a fit: too small to afford you, too large to buy from you, industries you cannot serve.
Be honest about the signals. If you sell dispatch software for plumbers, a signal might be “more than five trucks” or “recently posted a job for a dispatcher”. Those are visible from outside. “Struggling with scheduling” is not, and a tool that claims to detect it is guessing. This document becomes the instructions for every tool and person in the process. When a list comes back full of companies that do not match, the fault is usually in the definition.
2. Know where the raw names come from
A prospect list starts as a list of companies, and companies come from a limited number of places: the databases inside tools like Apollo and Clay, public directories (state registries, licensing boards, association member lists), review sites and maps for local businesses, job boards, trade press, and your own records of past quotes.
Each source has a character. Database tools are broad but shallow, with details often years out of date. Licensing boards are accurate for regulated trades but give you no contact person. Job boards are the freshest signal of what a company is doing now. Combine them: a database pull filtered by your definition, cross-checked against an authoritative directory, with a freshness signal from job postings. Every company should carry a note saying where it came from, because when bad leads show up, the source column is how you find the leak.
3. Use enrichment tools for facts, not judgment
Apollo, Clay, ZoomInfo and similar tools are excellent at the mechanical part: given a company, return its size, location, website, the people who work there with titles, and a best-guess email for each. Clay in particular works like a spreadsheet where each column can call a different data provider, so you can chain lookups.
What these tools are bad at is deciding fit. Their industry categories are coarse, their employee counts are estimates, and their “intent signals” are mostly inference from web traffic. Watch how each tool charges: most use credits, one per email found, more per phone number, and the bill climbs fast if you enrich everything. Pull the company list cheaply, filter hard on fit, and spend credits only on the survivors.
4. Let AI read the company and decide fit
For each company that passes the basic filters, have a language model read its website (homepage, about, services, careers) and recent news, and answer a short structured set of questions: does this company match the fit definition, yes or no, with a one-sentence reason; what is the single most specific fact about it relevant to what we sell; and what page did it come from.
The rules matter more than the model. It may only use information on pages it actually read, and it must point to the source for every fact. If it finds nothing relevant, it says “nothing found” rather than inventing a hook. A generic observation (“a growing company focused on customer service”) is a failure. Expect a meaningful share of companies to come back as no fit or nothing found, and treat that as the step working. Running this across hundreds of companies and checking it on a sample is in AI Personalization at Scale for Outreach That Reads as Human; how the fact becomes an email is in AI Cold Email Outreach Best Practices for Business Owners.
5. Find the right person, not just a person
A company on the list is not a lead. A specific person who could plausibly say yes, or route you to the person who can, is a lead. For a small business that is often the owner; for a mid-sized one it is a department head. Your fit definition says which titles you want, and the model reading the website can often tell who runs the relevant area even when titles are vague.
When the tool gives you three people at a company, pick one. Emailing the owner, the operations manager and the office administrator with the same pitch gets you noticed for the wrong reason, because they compare notes. Record the reasoning: “chose the practice manager because the about page says she oversees vendors” is a line the salesperson can use.
6. Verify every address before it is used
An email address from an enrichment tool is a guess, sometimes a very good one, built from a name and the company’s usual address pattern. Some are wrong, and sending to wrong addresses produces bounces, the fastest way to damage a sending domain (Email Deliverability for Cold Outreach, Explained Plainly).
Run every address through a verification service before it enters a sequence. It checks whether the mailbox exists and returns a status: valid, invalid, risky, or catch-all (the domain accepts mail to any address, so nothing can be confirmed). Send only to valid. Verify close to the time of sending, because people change jobs.
7. Build and enforce the do-not-contact list
Before the first email goes out, assemble a list of everyone who must never receive one: current and former customers, everyone who ever asked to be removed, vendors and partners, anyone with an open quote. Include every person at those companies, because cold-pitching the owner’s colleague the week after the owner said no is a story that gets told.
The list must be enforced by software at the moment of sending, not by a person remembering to check. The sequencer checks it as the last step before each message, and it refreshes from the CRM (the software where you track contacts and deals) automatically so a new customer is excluded the day the deal closes (AI CRM Automation That Makes the CRM Do the Work for You). Keep opt-outs forever, and match on company domain as well as email address, so a person who opted out is protected even if a fresh pull finds them at a different address.
8. Keep the data clean and know what you may hold
A prospect list is personal data. In the United States there is no general restriction on holding business contact data for outreach, but state privacy laws are expanding and some give people the right to ask what you hold and have it deleted. In the EU and UK, holding it at all requires a lawful basis, usually legitimate interest, and a way to honor objections. The broader picture is in AI Data Privacy for Business, What Happens to Customer Data.
The practical habits are the same everywhere. Keep one master list rather than copies in five tools. Record the source and date of every record. Delete what you are not using. And check what any AI provider does with data you send it; the major providers offer business terms that exclude your data from training, but consumer defaults often do not. Clean data is also just good sales.
Picture a business like this one
The business below is a composite of the kind of company that writes to us, not a client. The numbers describe the shape of the problem, not a case study.
Picture a business like this one: a regional distributor of restaurant equipment with sixty employees that wants to add independent restaurant groups (three to fifteen locations) in two neighboring states. The sales team had been working from a purchased list of “restaurants” that was mostly single locations and chains, with a bounce rate high enough to get one mailbox blocked.
What a distributor like this would build:
- A one-page fit definition: independent groups with three to fifteen locations in the two states, the owner or director of operations as the contact, signals such as a new location or a job posting for a kitchen manager.
- A company pull from a database tool, cross-checked against state health department licensing lists to confirm multi-location groups, joined to job board data for freshness.
- A model-driven reading step returning fit yes or no with a reason, the single most relevant fact with its source page, and the likely right contact. About forty percent of the pull is dropped here.
- Enrichment credits spent only on the survivors, one contact per company, every address verified the week before sending.
- A do-not-contact list seeded from the CRM, matched on email and domain, checked by the sequencer at send time and refreshed nightly.
- Every record carrying its source, fit reason and hook, written into the CRM.
What changes: the sending list is a few hundred names rather than several thousand. Bounces drop to a level that no longer threatens the domain. Replies reference the fact in the first line (“yes, we just opened in Tacoma, who told you?”), and the salesperson has the answer in the CRM.
What it costs to run
Enrichment is the most variable cost. Apollo, Clay and their competitors have free tiers covering a few dozen lookups a month and paid plans from roughly fifty to several hundred dollars a month, with credits consumed per email or phone number. Check the current pricing pages; the bill tracks how disciplined you are about enriching only companies that passed the fit filter.
Verification is fractions of a cent per address. The AI reading step is a usage cost: with current models, reading a company’s site and returning a structured answer costs a fraction of a cent to a few cents per company, so a thousand companies is a few dollars to a few tens of dollars. A workflow tool or small server adds ten to twenty dollars a month. The real cost is the person who writes the fit definition, reviews the model’s fit calls on a sample every week, and maintains the exclusion list.
The mistakes we see most
Buying a list and calling it research. Without a fit definition, a reading step and verification, a purchased list is a bounce generator.
Enriching everything before filtering. Credits get spent on companies that were never going to fit.
Letting the model invent a hook. Without a “nothing found” option and a source requirement, a language model produces a plausible fact for every company, and some will be wrong.
Contacting three people at the same company. They compare notes. Pick one.
A do-not-contact list that lives in someone’s head. It has to be enforced by the sending tool and fed from the CRM.
When to bring in help
An owner or salesperson can do the first version by hand with off-the-shelf tools: a written fit definition, a database tool on a low tier, a verification service, and a spreadsheet with source, fit reason and hook columns. Reading twenty company websites yourself and writing the hook for each is the best training there is for what the automated step should produce.
A developer becomes worth it when the list needs to refresh continuously, when the reading step has to run over hundreds of companies a week with rules you can audit, when every tool needs to share one do-not-contact list without a human copying between them, and when you want a log of every decision so you can find out why a bad lead got through.
Levelbrook builds lead research pipelines like this for businesses, at a fixed price from a written scope. All of it runs in accounts you own. The form below is how a conversation starts.