The Convergence Nobody Ordered

HMG Thinking · Copywriting

The Convergence Nobody Ordered

What happens to distinctiveness when competitors solve the same communication problem with the same models, trained on the same category language.

Marketing leaders have started noticing something difficult to articulate precisely: AI-assisted copy, across a growing number of categories, is beginning to sound like it came from the same source — even when it didn’t. The headlines differ. The exact phrasing varies. But the underlying shape of the argument, the rhythm of the claims, the structure of the pitch, feels increasingly familiar across competitors with no direct relationship to one another.

The common explanation is that people are prompting badly. Ask a generic question, get a generic answer. Edit harder, brief better, train the model on brand voice, and the problem resolves. This explanation is not wrong. It is incomplete, and the incompleteness matters, because it locates the entire problem at the production layer — as something technique can fix — when the evidence suggests part of the cause sits somewhere else: in what these systems were built to do, and in what organizations feed them before a single word gets generated.

This is not an argument that AI makes brands sound the same. That claim is already common and too shallow to guide an investment decision. The more useful question is this: if convergence is partly a documented behavior of how these systems generate language, and partly a consequence of what goes into them, where does the actual leverage against sameness sit — and is that where organizations are currently investing?

What the Research Actually Shows

Start with what a large language model is doing when it generates text. It is not retrieving a stored answer. For a given prompt, the model estimates a probability distribution over what could plausibly come next, based on patterns learned during training, then generates by selecting or sampling from that distribution — a process shaped by the training data, the model’s architecture, the settings applied at generation time, and the specific prompt provided. This is a meaningfully different claim than “the model always picks the single most likely word.” It can and does introduce variation. The question this raises is not whether variation is possible, but how much genuine diversity survives once a model has been trained and aligned to produce responses people rate as good.

That is precisely what a 2025 study examined directly. “Artificial Hivemind: The Open-Ended Homogeneity of Language Models,” by researchers at the University of Washington, Carnegie Mellon University, and the Allen Institute for AI, received a Best Paper award at NeurIPS 2025 and is available as a preprint. The authors built a dataset of roughly 26,000 open-ended queries — questions with no single correct answer, such as requests for a metaphor or a piece of advice — and sampled 50 responses per query across several major model families, including GPT-4o, Claude 3.5 Sonnet, Llama 3.1, and Qwen 2.5.

The findings were specific and measurable. Individual models tended to repeat themselves closely across independent samples of the same prompt — what the authors call intra-model repetition. More strikingly, different models built by different companies frequently converged on similar or even identical answers to the same open-ended question — inter-model homogeneity. In one illustrative example the researchers describe, asking for a metaphor about time produced outputs that clustered almost entirely around two ideas, one dominant, rather than scattering across the wide range of metaphors a group of people might generate. The paper attributes this largely to how models are aligned after initial training — techniques like reinforcement learning from human feedback appear to reward a narrow, consensus notion of a “good” answer, which can prune away responses that are equally valid but less conventional.

It is worth being precise about what this establishes and what it does not. This is a technical study of open-ended text generation under controlled conditions, using a specific evaluation methodology the authors themselves note has limitations — including reliance on embedding-based similarity measures, and a focus on English-language queries. The study did not measure brand perception, marketing performance, or any commercial outcome, and its authors explicitly stop short of claiming a fully worked-out causal mechanism. What it does establish, credibly and directly, is that convergence in model outputs is a documented behavior across major model families, not simply a symptom of any one person’s careless prompting. Whatever role prompting plays, it is not the whole explanation.

The Category-Level Question

Extend this from model behavior to a competitive market, and a more specific risk comes into focus — stated as a risk, not a certainty the research proves.

Most marketing language, at the category level, is not private. Common positioning phrases, standard claims about speed or ease of use, and the general shape of a persuasive pitch in a given industry are heavily represented in whatever text a model was trained on. When a company in that category asks an AI system to help write a homepage or a sales email, it draws on a model whose sense of that category reflects an average of what has already been written, weighted toward what appeared most often.

Where many competitors in the same category rely on a small number of widely used foundation models, prompted with broadly similar requests, informed by similar public category knowledge and increasingly standardized prompting conventions, the conditions favor convergence. None of this guarantees identical output. But it describes circumstances under which sameness becomes more likely — a shared input field, processed by systems shown to compress toward a narrow set of favored responses, queried by many parties solving a similar problem at roughly the same time.

This is the deeper version of the “AI sameness” complaint. The issue is not that any individual output is poorly made. It is that when competitors share the same models, the same category knowledge, and similar production habits, there is less material available to distinguish one company’s output from another’s — unless something distinctive entered the process before the prompt was written.

The Homepage as Evidence of an Older Condition

If this dynamic operates at the market level, there should be some visible sign of it beyond a controlled study. There is, though it requires the same care applied above.

Research from Wynter and Contentifai, cited in industry commentary on B2B messaging, found that a large majority of B2B SaaS homepages converge on a small set of recurring phrases and a limited number of hero-section structures. This is a useful illustration of category-level sameness in language and layout, though it reaches us through secondary reporting rather than direct access to the original studies, and should be weighted accordingly.

It is also, importantly, not proof that AI caused this. B2B homepage convergence predates the current generation of AI writing tools. Best-practice templates, agency conventions, competitive imitation, and swipe-file culture have pulled marketing language toward familiar patterns since long before a model could draft a headline. What the homepage evidence demonstrates is that the underlying condition — a category converging on shared language — already existed. AI did not invent this problem. What the NeurIPS findings suggest is that AI-assisted production may be well-suited to scale it: a documented tendency toward narrow, consensus output, deployed at far greater speed and volume than manual drafting allowed, applied to categories already prone to convergence before the tool existed.

This distinction changes what kind of problem this is. Not a new failure introduced by new technology, but a familiar tendency in marketing language, now operating at a different speed and scale.

Why Editing Alone Doesn’t Resolve It

The most natural response to all of this is some version of: have a skilled human rewrite it. This deserves serious engagement, because it is partly right.

A skilled editor can improve clarity, sharpen rhythm, correct tone, and introduce language that sounds distinctly like a particular brand rather than a generic voice. None of that should be minimized.

But stylistic distinctiveness and strategic distinctiveness are not the same accomplishment. An editor can take a generic underlying proposition — we help you do this faster, with less effort, backed by great support — and make the sentence describing it more elegant and more on-brand. The proposition itself, the actual claim about what makes this company different from the alternative beside it, remains exactly as generic as before the rewrite. The sentence improved. The distinction did not.

An honest version of this argument also has to acknowledge that human writers converge for reasons that have nothing to do with AI. A writer trained on the same category conventions, familiar with the same competitor sites, working from the same assumptions about what a SaaS homepage is supposed to say, will often produce language close to what a model would generate — not because a machine was involved, but because the category exerts its own pull toward convention. Removing AI from the process does not automatically solve this. The lever is not who holds the pen. It is what specific, non-generic material that person or system had to work with before writing began.

The Investment Question

This is the question that should actually reach a CMO’s or CEO’s desk, and it is about resource allocation, not philosophy.

Marketing organizations are investing, often significantly, in AI licenses, prompt engineering capability, workflow automation, and the raw capacity to generate more content, faster, at lower cost. This is not a mistake. Speed and production efficiency are legitimate capabilities, and nothing here argues against building them.

The question this evidence raises is narrower: is that investment matched by a comparable investment in the inputs a shared model, drawing on shared category knowledge, cannot generate on its own? Proprietary research findings. Direct customer insight that hasn’t been published anywhere the model could have learned it. A genuinely different point of view about the category, arrived at through the organization’s own experience rather than assembled from what competitors have already said. A specific mechanism or piece of evidence that is true of this company and not generically true of the category.

This is not a choice between the two. Production capability and distinctive input are not competing for the same budget line in any strict sense, and an organization can reasonably build both. The risk is this: where competitors in a category have access to comparable AI tools, comparable prompting conventions, and comparable public category knowledge, widely shared production capability is unlikely, on its own, to be a durable source of distinction — because a capability everyone has access to is a weak candidate for what sets any one of them apart. Whatever is genuinely distinctive about a company’s position has to come from material the organization possesses and the shared training data does not.

A model can generate substantial variation in phrasing, structure, and tone. It cannot manufacture a position the organization has not yet worked out. Feed it a category-average understanding of the business, and the output — however fluent, however fast, however well-edited afterward — is likely to read as a highly competent version of what everyone else in the category is already saying.

What This Does Not Mean

A few boundaries are worth stating plainly, since this argument is easy to overstate in either direction.

This is not an argument to stop using AI in content and copy work, or that doing so is inherently a strategic error. It does not mean every AI-assisted output is generic — plenty of AI-assisted copy is grounded in genuinely distinctive input and reads accordingly. It does not mean prompting quality is irrelevant; a poorly specified prompt can waste even a strong underlying position. It does not mean human-written copy is automatically more distinctive than AI-assisted copy — the evidence suggests the opposite risk exists in human work too, whenever writers lean on shared category convention. And it does not mean proprietary data or a distinctive point of view guarantees differentiated output; weak execution can still flatten a strong position.

Nor does the NeurIPS research prove that any specific brand’s market performance has suffered because of model-level convergence. It establishes a documented behavior worth taking seriously, not a business outcome. The honest position is that this technical finding, the homogeneity visible in B2B homepages, and the long-running decline in brand differentiation that Kantar’s research has tracked for years are compatible with one another and worth examining together — not proof that one causes another. These are separate evidence fields, from separate methodologies, answering separate questions, and the value of placing them side by side lies in the pattern they suggest, not in a chain of proof they do not provide.

What the evidence does support, carefully stated, is this: independent research has documented that major language models converge toward a narrow band of favored outputs on open-ended tasks, and marketing categories were already prone to linguistic convergence before this technology existed. Put those together, and the practical question for leadership is not whether to use AI, but what is being put into it — and whether the organization has invested as seriously in answering that question as it has in the tools generating the sentences.

If production capacity keeps becoming more widely available and more evenly distributed across an entire competitive category, it becomes a weaker and weaker basis for distinction, since a capability everyone holds cannot easily be what sets any one of them apart. The harder, more durable question — the one a shared model cannot answer on a company’s behalf — is what that organization knows, believes, or has evidence for that its competitors, working from the same public information and the same tools, cannot simply generate for themselves.

Public Source References

Sources used in the examination.

These references support separate parts of the examination. Technical model behavior, B2B homepage convergence, and brand differentiation remain distinct evidence fields and are not presented as a causal chain.

Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)

Liwei Jiang, Yuanjun Chai, Margaret Li, et al. — University of Washington, Carnegie Mellon University, and Allen Institute for AI · 2025
Primary technical research; NeurIPS 2025 Datasets & Benchmarks Track.
Primary preprint →
NeurIPS conference record →

How to Simplify B2B Messaging Without Dumbing Down Your Product

Wynter / Contentifai research as reported by Pitch Kitchen · 2025–2026
Secondary reporting used only as an illustration of B2B homepage convergence; the underlying studies were not directly verified in the final research pass.
View source →

Brand difference is in crisis

Kantar BrandZ research via Contagious, commentary by Dom Boyd · 2025
Used as a separate longitudinal signal concerning meaningful brand difference; not presented as caused by model convergence.
View source →

HMG Thinking

← Return to HMG Thinking

Scroll to Top