The AI PioneerPlain-language field notes on putting AI to work in a real business. From Levelbrook.

The AI Pioneer / Outreach and salesNo. 27

AI personalization at scale: mail merge with adjectives versus a note a person wrote

What real personalization is made of, what to research per company, how to write templates that do not read as templates, how to run it across hundreds of prospects, and the review sample that keeps it honest.

11 minute read. Updated 2026-09-17. Ask about your business

“Hi Sarah, I saw that Northgate Dental is a leading provider of dental care in Portland and I was really impressed by your commitment to patient experience.” You have received this email. It knew your name, company and city, and it was still obviously written by nobody, to nobody. That is what most vendors mean by AI personalization at scale.

Real personalization is different in kind, not degree. It is a message that could only have been written to this one person, because it says something true and specific about their business that took effort to find. AI makes the effort cheaper, not unnecessary, and a tool that skips it produces the email above, faster.

This article explains what personalization consists of, what is worth researching, how to build a template that does not read as one, how to give a language model rules it will follow, and the review habit that catches failures before your prospects do.

What this actually is

Personalization at scale means producing many one-to-one messages, each shaped by facts about its recipient, without a person writing each from scratch. The “scale” part is a template and an automated research step. The “personalization” part is the facts. If the facts are shallow (name, company, city), the output is a mail merge with better grammar. If they are specific (what the company did last month, what it is hiring for), the output is something a good salesperson would have written.

A language model (software that reads and writes text, the engine behind Claude, ChatGPT and Gemini) does two jobs here. It reads about each company and extracts the facts, and it writes the message that uses them. The reading is the valuable job. The writing is the easy one, and it is where the model most needs constraining, because left alone it produces the confident, empty prose it learned from millions of marketing emails.

The ordinary-business analogy is the difference between a form letter with the name filled in and a note from someone who met you at a conference last week. The second mentions the thing you talked about, is short, and asks one thing. You answer it because ignoring it would be rude to a person, and a form letter has no person behind it.

1. Decide what a real observation looks like for your business

Before any research, write down the two or three kinds of facts that would make a message from you relevant to a prospect. For a staffing agency, a fact is “posted four listings for the same role in two months”. For a commercial cleaner, “opened a second location in March”.

Then write down what does not count. Adjectives about the company are not facts; “growing”, “innovative”, “customer-focused” all fail. Restating the company’s tagline fails. The test: could this sentence have been written without visiting their website? If yes, it is not an observation. This list is the brief for the research step and the standard for the review sample.

2. Research from sources that actually contain facts

The sources that reliably contain such facts are the company’s own website (especially the about, services, news and careers pages), job boards, public reviews, local news, trade press, and the prospect’s own public posts. Where the list comes from is covered in AI Lead Research and Enrichment for Business Outreach.

Give the model a fixed reading list per company and a fixed set of questions. “Read these pages. Return up to two facts matching the definition below. For each, quote the exact sentence it came from and name the page. If nothing matches, return NOTHING FOUND.” The quote requirement keeps the model honest; a fact it cannot quote is a fact it made up. In our experience, a research step that never returns empty is one that has been allowed to fabricate.

3. Build the template around slots, not paragraphs

A template that survives personalization has a fixed skeleton and a small number of slots. The skeleton is short: a first line that is entirely the observation, one or two sentences of offer, one question. Everything outside the slots is fixed and written by a person.

The mistake is a paragraph with blanks: “I noticed [FACT] and I think [PRODUCT] could help [COMPANY] [BENEFIT].” That reads as a template whatever goes in the blanks. Instead, write three or four complete example emails by hand, to real prospects, in the voice you want. Those become the model’s instruction; it sees the shape without a fill-in form, and its output varies the way a person’s would. Keep the message under about one hundred and twenty words (AI Cold Email Outreach Best Practices for Business Owners).

4. Give the model rules it cannot talk its way around

The instructions you give a language model are a specification, and a vague specification produces vague output. The rules for the writing step should be explicit, short, and checkable. Use only the facts returned by the research step. Write in plain words; here is a list of banned words. No sentence about the sender’s company history. No exclamation marks. One question only. Sign off with this exact signature. Output only the email.

Include the hand-written examples and tell the model to match their tone and length, not their content. Add one or two examples of bad output with a note on why each is bad. Keep the prompt (the written instructions, examples and rules) in a versioned file, not a text box someone can edit and forget, so when output quality changes you know exactly what the instructions were that day. Treating prompts as specifications is the subject of Prompt Engineering for Business Applications That Work.

5. Personalize the first line, not the whole email

The temptation with a capable model is to personalize everything: a custom offer, a custom benefit, a custom sign-off. This multiplies the ways the output can go wrong and makes the email longer. A recipient decides in the first line whether to keep reading, and that line is where personalization earns its keep.

After the first line, consistency is a feature. The same offer in the same words to everyone on a list means you can tell what is working. If the offer is rewritten per recipient, you cannot compare replies, and the model drifts toward salesy phrasing over hundreds of variations. The one exception is the question, which sometimes benefits from being tied to the observation (“who handles vendor selection for the new location?”).

6. Review a sample every week, by hand, against the standard

An automated step produces hundreds of emails a week, and nobody reads them all. Somebody must read some. Pick twenty at random and read each against the standard from rule one: is the observation a fact, is it quoted from a real page, is the email under length, does it sound like a person. Score each pass or fail.

Do this every week, not just at launch. Models change, websites change, and quality drifts in ways invisible from a dashboard. A reply rate that slides over a month is usually a research step returning shallower facts, and the weekly sample catches that in week one. When the sample fails, fix the instructions, not the individual emails. Change one thing, run next week’s sample, and see if the score moves. Keep the samples and scores; they are the log that proves the system is working.

7. Handle the ones where nothing was found

A share of every list will come back from research with nothing worth saying. There are three honest options: send a plainer email that does not pretend to an observation (“I work with independent pharmacies on X; is that something you handle in house or with a vendor?”), hold the company for a later pass, or drop it.

What you must not do is let the writing step fill the gap. A model told to write a personalized email with no facts will produce a sentence that sounds like one, and it will be wrong or generic. Make “nothing found” a hard branch in the workflow that routes to the plain template, never to the personalized one. In our experience the plain, honest email performs respectably, precisely because it does not pretend.

8. Keep the human in the loop where it counts

Full automation from research to send is reasonable for a well-tuned system with a weekly sample. But two places need a person. The first is the initial batch: for the first fifty to one hundred emails from any new list or template, someone reads every message before it goes.

The second is replies. Personalization ends the moment someone answers. The reply is a sales conversation between two people, and it needs to be written by one. A model can classify it and draft a response for approval, but the send button belongs to a salesperson. Where replies land is covered in AI CRM Automation That Makes the CRM Do the Work for You. Everything in between, research, drafting, and sending at the safe volumes in Email Deliverability for Cold Outreach, Explained Plainly, can run without a person touching each message, provided the sample review is happening.

Picture a business like this one

The business below is a composite of the kind of company that writes to us, not a client. The numbers describe the shape of the problem, not a case study.

Picture a business like this one: a payroll and HR services firm with twenty-five employees that sells to companies with twenty to two hundred staff in its state. It had an outreach tool configured by a contractor that used “AI personalization”. The first lines were things like “I see that Acme Manufacturing is a leader in precision manufacturing solutions.” The owner stopped the campaign after a customer forwarded one with a joke attached.

What a firm like this would build:

  1. A one-page observation standard: facts that count (job postings for HR or payroll roles, a new location, a named payroll vendor on the careers page), and facts that do not.
  2. A research step that reads five fixed pages per company plus recent job listings and returns up to two quoted facts with source, or NOTHING FOUND. About a third of the list returns empty.
  3. Four hand-written example emails from the owner, each under one hundred words: one observation, one fixed offer sentence, one question about who handles payroll internally.
  4. A writing prompt with the examples, a banned-word list, a length cap, and the rule that only research-step facts may be used. Empty-research companies route to a plain three-sentence template.
  5. A weekly review of twenty random outputs scored against the standard, with the scores kept in a sheet.
  6. Replies classified by the model, drafted for approval, and answered by the owner or sales lead within the hour.

What changes: the first line of every email is a quoted fact or a plain question, never a compliment. Replies open with the fact (“yes, we lost our payroll person in June, how did you know?”). The sample review catches a drift in month two when the model starts quoting careers-page boilerplate as a fact, and the standard is tightened.

What it costs to run

The AI usage is small. With current models, reading a handful of pages per company and producing two quoted facts costs a fraction of a cent to a few cents per company; writing the email costs less. A thousand companies through both steps is a few dollars to a few tens of dollars. Check the current pricing pages of whichever provider you use.

The surrounding tools are the same as any outreach setup: a sending tool with warm-up and sequencing (typically thirty to one hundred dollars a month at small volumes), enrichment credits, a verification service at fractions of a cent per address, and possibly a workflow tool or small server at ten to twenty dollars a month. The real cost is the weekly hour on the review sample and the time spent writing the examples and the standard at the start.

The mistakes we see most

Confusing merge fields with personalization. Name, company and city in the first line is a form letter, and the recipient knows it.

Letting the model personalize everything. More variable parts means more ways to fail, longer emails, and no way to compare what works.

No “nothing found” branch. A model asked for a hook will always produce one. Without a way to say there is none, invented hooks go out with the real ones.

A prompt that lives in a text box. When quality changes, nobody knows what the instructions were that day.

No review sample after launch. Drift is silent and the weekly sample is the only alarm.

When to bring in help

An owner can do the honest version by hand for a small list: read twenty websites, write twenty first lines, send twenty emails, and learn what a good observation for their business looks like. That is the most valuable step in the whole process, because the automated system is only as good as the standard you wrote from that experience. Off-the-shelf tools can then run a version of it, with limited control over the research step.

A developer is worth it when the research step must read specific sources with your rules and refuse to invent, when the writing prompt needs to be versioned and testable, when “nothing found” has to route to a different template automatically, when outputs and review scores need logging, and when replies must flow into your CRM with the research attached.

Levelbrook builds personalization pipelines like this for businesses, at a fixed price from a written scope. The prompts, logs and data all live in accounts you own. The form below is how a conversation starts.

Questions owners ask

What is the difference between personalization and mail merge?

Mail merge fills a template with fields you already have: name, company, city. Personalization uses a fact about the recipient that took effort to find, such as a recent hire or a new location, and builds the first sentence around it. The first is a form letter; the second reads as a note from a person.

How much personalization is enough for a cold email?

One specific, verifiable fact in the first line is enough, and often better than more. Keep the offer and the question fixed so you can tell what is working. Personalizing every sentence makes the email longer and more generic.

Can AI personalize emails without making things up?

Yes, if it is required to quote the source sentence and page for every fact and allowed to answer "nothing found". Without those two rules a language model will produce a plausible fact for every company, and some will be wrong.

How do I know if my outreach tool's personalization is any good?

Read twenty random outputs. If the first lines are facts you could only know by visiting the company's website, it is working. If they are compliments or restatements of the company's tagline, it is mail merge with extra steps.

Should a person review every AI-written email before it sends?

Every one for the first fifty to a hundred from a new list or template, then a random sample of about twenty a week. Every reply should be written or approved by a person.

Want this done properly for your business?

Tell us what the task is and what it costs you today. You get a reply from an engineer with a couple of questions, an honest view of whether it is worth doing, and a fixed price if it is.

One reply within a business day, from the engineer who would do the work. No newsletter, no sales sequence.
Sent. We read every one of these and will reply within a business day with a couple of questions and, if it makes sense, a time to talk.