Duplicate keyword combinations in a keyword combiner come from two separate sources: repeated lines in your input lists and collisions where different pairs concatenate to the same final phrase. A complete Cartesian product of two one-per-line lists exposes both problems at once, which is why a clean output rarely matches the raw input count. The tool removes exact case-sensitive duplicates within each list in first-seen order, then keeps the first phrase it generates whenever distinct pairs collide to the same result, and reports the totals in the summary panel. Knowing where duplicates appear prevents broken generations: the truncated output vanishes because the entire request fails before output is generated when the 10,000 input-pair limit is exceeded, not at the request layer where partial samples would still slip through. That distinction keeps a string-combination utility from pretending to be keyword research, because the deduplicated list is still only a candidate set that requires human review for relevance, intent, and demand before it is used in any campaign or publishing workflow.

Where Duplicates in a Keyword Combiner Actually Come From
Every duplicate in the final output traces back to one of two roles played by the inputs and the concatenation step. The first role is the source editor itself. Spreadsheets, ad-group exports, and scraped suggestion lists routinely contain the same term typed twice, pasted with a stray trailing space, or carried through a copy/paste that lowercased only some rows. The second role is the concatenation step. Once two lists are crossed, pairs whose joined output matches a pair that was already generated earlier are considered a collision, not an input error. A trivial case is `ab` plus `c` against `a` plus `bc`: with an empty separator, both pairs produce the literal string `abc`, so the second pair becomes a collision that the tool has to remove.
Empty separators make collisions far more common than separated ones. With a hyphen, `ab-c` and `a-bc` are visibly different; with no separator at all, `abc` is produced twice and one occurrence is dropped. Long modifiers also raise the collision rate because more letters are shared across pairs. Repeating the same modifier across an entire sheet of unrelated inner rows can quietly double the apparent output without doubling the useful phrases.
How the Tool Handles Repeated Input and Colliding Output
Normalization happens in a fixed order before any pair is generated. Each non-empty line is trimmed at its outer edges, empty lines are ignored, and any line that contains a control character or a Unicode line or paragraph separator is dropped. Each term may contain up to 100 Unicode code points. Each source editor accepts up to 100,000 UTF-16 code units and up to 500 unique terms. Exact duplicates within a list are removed after trimming while the first occurrence is preserved. Deduplication is case-sensitive, so `Shoe` and `shoe` remain distinct because they may be intentional variants for exact match and phrase match campaigns.
Once the lists are normalized, the pair count is computed before generation. If that number is above 10,000, the entire request fails and no output is produced. If the number is within the limit, the selected loop order runs deterministically, the chosen separator is inserted between the two strings, and the final collisions are deduplicated in first-seen order. Internal spaces, punctuation, accents, emoji, and non-Latin scripts are preserved unchanged throughout. Nothing is normalized beyond outer whitespace and the case-sensitive duplicate check.
Clean Up Duplicate Combinations in Keyword Combiner
The fastest way to reach a clean list is to clean the inputs, then read the panel that reports what the work actually did. Use the Keyword Combiner as a deterministic Cartesian product, and treat the summary line at the bottom as an audit trail.
- Open the tool and paste one base term per line in list A, one modifier or second term per line in list B.
- Audit each list on the page: remove any row that has stray leading or trailing whitespace, any row that duplicates another row after trim, and any row that contains a control or paragraph separator that the editor will reject.
- Decide whether you need the empty string or a visible separator. Pick a single space, a hyphen, or a plus sign unless you have a specific reason to leave the separator empty.
- Pick A-then-B or B-then-A as the outer-to-inner loop order. The choice changes phrase order, not pair count, so pick whichever order makes your final list easier to scan.
- Run the combine action and read the five summary numbers at the bottom of the result block.
- Compare the theoretical input-pair count against the unique output count. If the gap is wider than expected, collisions are the cause.
- Compare the unique output count against the number of input duplicates the summary reports. If the input duplicate count is large, the source lists need trimming before the next run.
- Sort the output or paste it into a spreadsheet and remove any phrase that has no clear landing page, intent, or demand signal.
- Press the copy button to capture the full block of one phrase per line, with no ellipsis or sampled tail.
Read the Summary Panel to Catch What You Missed
The summary block at the bottom of the result is not a marketing footer; it is the only audit trail the tool provides. Five numbers matter. The unique output count is the number of distinct phrases after input collisions removed. The normalized list sizes are the counts of unique case-sensitive terms in list A and list B after trim. The theoretical input-pair count is the raw A multiplied by the raw B before either set of collisions is removed. The removed input duplicates number is the count of source rows that were discarded as exact repeats. The removed output collisions number is the count of generated phrases that were discarded because a previous pair produced the same string.
If the unique output count is unexpectedly large, narrow the source lists before the next run. If the unique output count is unexpectedly small, two things are usually true at once: the source lists contain exact repeats, and the chosen separator is collapsing distinct pairs into the same string. The two problems feed each other, and the summary numbers expose both at once.
| Hard limit in the tool | What it bounds | What happens at the boundary |
|---|---|---|
| 100 Unicode code points per term | Each line in list A or list B after trim | Longer lines are rejected as a single line, not truncated |
| 100,000 UTF-16 code units per editor | Total pasted text per list | Excess text is rejected at the editor boundary |
| 500 unique case-sensitive terms per list | Distinct rows after source deduplication | Additional unique rows are not accepted into that list |
| 20 Unicode code points in the separator | The literal string placed between A and B | Longer separators fail validation; line breaks and controls are never accepted |
| 10,000 input pairs | Theoretical A multiplied by B before any output | The entire request fails before output is generated; no partial result |
Separator, Order, and Case Choices That Cut Noise
The separator is the single biggest control over collision frequency, because the tool inserts exactly that literal string with no hidden padding. A one-character separator such as a space, hyphen, or plus sign almost never produces accidental collisions unless the inputs themselves contain the same character. An empty separator collapses adjacent letters and turns far more pairs into identical strings. If your inputs are normalized, your separators can be short. A short separator has another benefit: it keeps the output readable in a spreadsheet column and keeps the phrase length predictable for ad-group policies that cap keyword length.
Case matters more than most readers expect. Because deduplication is case-sensitive, `Shoe` and `shoe` are two distinct terms, two distinct input rows, and two distinct output phrases for every pair they participate in. That behavior is deliberate, because exact match and phrase match in Google Ads are case-insensitive while broad match modifiers and many editorial taxonomies care about capitalization. Use that property deliberately when you want both variants in the candidate list, and use the source editor to lowercase first if the property was not intentional.
| What the tool does | What you still have to do |
|---|---|
| Builds the complete A × B Cartesian product in first-seen order | Pick the pair budget that fits the 10,000 input-pair limit |
| Removes exact case-sensitive duplicates inside each list | Lowercase or normalize terms only if the case variant was not intentional |
| Deduplicates colliding output in first-seen order | Read the removed output collisions number to know how many were dropped |
| Trims outer whitespace and ignores empty lines | Remove control characters and stray paragraph marks before pasting |
| Returns one phrase per line, copyable as a single block | Remove irrelevant or unnatural phrases before using the list |
| Runs entirely in the current browser tab, lists never sent to a server | Validate demand, intent, and landing-page fit through a separate source |
Review the Output Before You Use It in a Campaign
A deduplicated output is a candidate list, not a finished keyword strategy. The tool does not check search volume, competition, cost per click, intent, language, relevance, or grammatical quality, and it does not query an advertising account, search engine, keyword provider, or AI service. Some combinations will sound unnatural, overlap in meaning, or have no measurable demand. Treat each line as a candidate that still needs human review and, when appropriate, first-party or paid keyword evidence before it is committed to a budget.
The deterministic output helps that review step. Because order is stable, rerunning the same inputs yields the same text, which means a fixed candidate list can be diffed against a later run after a single edit. That property makes the tool well suited for version comparison across drafts, internal taxonomy reviews, and ad-group exploration before any phrase is shipped to a live campaign or publishing workflow.
Treat the Tool as Documentation
The separation between string combination and keyword research is what makes the output trustworthy for what it actually does. Combining two lists is a mechanical operation; deciding which phrases are worth bidding on or writing for is a judgment call. Keep that boundary clear in your workflow. Use the summary numbers as a record of what was removed, treat the output as a candidate set, and apply the rules of the destination system yourself before any phrase reaches a campaign, a brief, or a published page.
For a deeper look, see Keyword Density Checker on iPhone: A Safari Workflow.