All articles
Audience Intelligence6 min read

Customer language when you have no reviews

Every message mining guide assumes reviews exist. How to find customer language when there are none, and the one thing that corpus cannot tell you.

A client sells compressed-air leak audits to factories. Six-figure contracts, a sales cycle measured in quarters. I went looking for the buyers' own words first, the way I always do.

Amazon: nothing. App Store: nothing. Reddit: one thread from 2019 about a compressor brand, seven comments, all engineers arguing about PSI. Capterra had four reviews of a competitor and three read like staff wrote them. Two hours in, my quote doc held four sentences.

Every message mining guide starts at step two, after the reviews exist.

Short answer: When a category has no public reviews, you swap the public corpus for the private one your client already owns: sales-call notes, the objections sitting in the inbox, chat transcripts, free-text form fills. Then you rank those quotes by intensity rather than by count, because thirty quotes cannot support a frequency read.

The takeaways

  • The private corpus is the substitute. Sales calls, support email and form free-text hold longer, more specific language than most public reviews, and they exist for every business that has ever answered a phone.
  • Thirty quotes cannot be counted. At that sample size, two mentions is two people. Rank by the consequence a quote names instead: money lost, a deadline missed, someone they had to apologise to.
  • Your inbox has no cold end. Everyone in it contacted you first, so it holds Solution-Aware and Product-Aware language and none of the Unaware stage Eugene Schwartz put at the top of his ladder.

Where does customer language live when nobody reviews you?

In four first-party places, all of them already sitting on your client's servers. Sales-call notes, especially the objections a rep hears in the first ten minutes. The support inbox, where a frustrated customer writes four paragraphs nobody would ever type into a review box. Chat transcripts. And the free-text field on the inbound lead form, which people fill in a hurry and therefore fill honestly.

That last one is the most underrated source in this job. "What can we help you with?" gets answers like "boiler keeps cutting out at night and my tenant is calling me". Nobody edits a form field for style. You get the problem in the words a person uses when they are annoyed, which is the register a scroll-stopping hook needs. Ask the sales lead to forward you fifty emails and you have a corpus by Thursday.

Why does a thirty-quote corpus break frequency counting?

Because at that size, a percentage is a costume. Two mentions out of thirty reads as 7% in a slide, and it is two people who happened to email in the same week. Run the same collection a month later and the ranking reshuffles. Frequency needs volume you do not have here, so borrowing it from review-mining guides produces confident nonsense.

Rank by intensity instead. Intensity means the quote names a consequence: a number, a deadline, a person the buyer had to answer to. "The compressor ran all weekend and I found out on Monday from the invoice" outranks five people saying the service was helpful, because it holds a cost, a timeline and an emotion, and each of those is a hook. A small sample just makes the shortcut obvious; a pain map should be ordered this way at any size.

What is your own inbox structurally missing?

The entire cold end of the funnel. Everyone in a private corpus crossed a threshold before they turned up in it: they noticed a problem, decided it was worth solving, went looking, found you. By the time their words reach your inbox they sit at Solution-Aware or Product-Aware on the five awareness stages Eugene Schwartz set out in Breakthrough Advertising.

The Unaware stage is absent by construction. A plant manager who has never thought about compressed air leaks does not email a leak-detection company about them. So if you build every hook from your own files, you write ads that address people who do not yet know they have the problem, in the vocabulary of people who already solved it.

That mismatch is invisible in the doc and expensive in the account. It shows up as cold creative with a decent CTR and no conversions, because the clicks came from the small warm slice your copy was written for. There is a longer treatment of the ladder in matching ad copy to awareness stages.

Where do you get the cold end of the funnel?

From the category one tier up, which almost always has the public corpus yours lacks. Compressed air leak audits have no reviews. "Energy audit", "facility maintenance", "utility bill too high" have thousands, written by exactly the people you need to reach before they know your category exists.

Two moves work here. First, read the one-star reviews of the adjacent product, because complaints describe the before-state in plain language and the before-state is what a cold hook has to name. Second, widen the forum: skip the niche subreddit that does not exist and read the general trade forum where your buyer posts about something else entirely.

If a direct competitor with real reviews does exist, start there instead. That case has its own method, covered in audience research for a new product.

How do you know when you have enough?

When ten new quotes stop producing new phrasings. That is saturation, and it is the only honest stopping rule available at this sample size. You will not reach statistical comfort with a private corpus, so waiting for it just delays the launch by a fortnight.

For a narrow B2B niche that usually lands around forty to sixty quotes. Sort them into three or four pain themes, write three hooks per theme in the buyer's own wording, and put them live. The CTR spread across those hooks beats another week of collecting: it is the first time your reading of the corpus meets people who were not already talking to you. Then the winners go back into the doc, and the ads become a source of their own.

The part a tool cannot fix for you

This is where a research tool's limits show, so it is worth saying plainly. Adscalr's audience research searches five named sources: Reddit, Amazon reviews, the App Store, Google Play, and custom forums you seed with your own and your competitors' URLs. It pulls the exact phrasing out of each quote (two to four phrase markers, plus emotion and register), maps quotes onto the five awareness stages, and ranks the pain map by intensity.

In a thin niche, four of those five come back close to empty and the seeded forums carry the whole result. That is a coverage limit worth naming, and the private corpus in your client's inbox is what closes it. Same method either way: collect the exact words, rank them by what they cost the person who said them, stay honest about which part of the funnel you heard from. The audience intelligence pillar covers how the sourcing and the ranking work.

This is the thinking behind Adscalr.

See the product