Word documents store every external link as a relationship inside the .docx package, so removing them safely starts with reading that relationship file rather than scanning the visible text. A modern .docx is actually a ZIP archive that follows the Office Open XML standard; its external hyperlinks live in a small XML file named word/_rels/document.xml.rels, alongside the body text, styles, footnotes, embedded fonts, images, and revision history. When you copy visible text, you only see the human-readable anchor — a phrase, a button label, or the alt text on a shape — and you miss any link attached to an image, a citation, or a short fragment that does not read like an address. That is exactly why manual extraction is slow and unreliable for documents with dozens of references. The Word Hyperlink Extractor opens that relationship file locally in your browser, decodes the http, https, and mailto targets, deduplicates them, and gives you one plain-text line per URL that you can paste into a spreadsheet or a URL checker. The list is deliberately conservative: it ignores internal bookmarks, file-system paths, relative references, malformed values, and any protocol that is not http, https, or mailto, so every line is a safe external destination recorded by Word itself.

Where Word Actually Stores Its Links
A .docx document is more than the text you see on screen. Microsoft Word writes the file as a ZIP archive whose layout is defined by the Office Open XML standard, the same specification that governs modern Excel and PowerPoint files. Inside that archive, every piece of content — body text, styles, footnotes, embedded fonts, images, settings, even revision history — lives in its own XML part. External URLs are no exception.
The hyperlink destinations themselves are not stored in the visible paragraph XML. Instead, Word records them as relationships in a separate file: word/_rels/document.xml.rels. Each relationship has a type, an identifier, and, when it points to something outside the package, a TargetMode of External and a Target string. Hyperlinks specifically use the relationship type whose name ends in /relationships/hyperlink. The visible text inside the document references those relationship identifiers through an r:id attribute, which is why copying the visible paragraph never copies the underlying destination.
This structure explains two everyday problems. First, a link can be attached to an image, a shape, or a citation field whose displayed text is a short label, a number, or even empty, so manual copying misses it. Second, the relationship file can hold targets that are not safe to treat as web links, such as internal bookmarks pointing to headings, relative paths to local files, or malformed values left over from older documents. The relationship mechanism is documented in the ECMA-376 standard for the Office Open XML file formats, which describes how Relationship entries with hyperlink type and external TargetMode are the authoritative source for clickable external destinations.
What "Removing" Really Means Before You Touch the Document
The phrase "remove word links" usually covers two different actions. The first is the click-and-erase approach inside Microsoft Word: right-click a hyperlink, pick Remove Hyperlink, and the formatting disappears but the text remains. That method works for one or two visible links, but it cannot tell you what you actually have. The second meaning is the audit approach: before deleting anything, build a list of every external destination the document contains, review each URL for safety or relevance, and only then decide what to remove and what to keep.
The audit approach is what the Word Hyperlink Extractor was built for. Reading the relationship file gives you the same list Word itself uses when it resolves clicks, without requiring you to scroll through every page, hover over every image, or open the document in protected view. Once you have the list, you can paste it into a spreadsheet, run it through a URL checker, compare it against your publishing requirements, and only then open Word and remove the links you do not want.
Treating extraction as the first step has another practical benefit: it leaves a paper trail. If a colleague asks why a reference was deleted, you can show them the extracted URL and the visible paragraph side by side, then explain the decision. For compliance reviews, citation audits, or template handovers, that paper trail is often more valuable than the actual removal, because it lets you defend each deletion against the original source.
Extract Every External Link From Word
The Word Hyperlink Extractor runs entirely in your browser. The .docx file never leaves your device, and no part of the document text, comments, or macros is sent to a service. Follow these steps to produce a clean URL list.
- Save the Word document you want to audit as a modern .docx file. Older .doc files, encrypted packages, and password-protected archives are out of scope — only modern .docx packages are processed — so convert or unlock the file first if needed.
- Open the Word Hyperlink Extractor in your browser.
- Choose the .docx file from your device using the file picker. The browser validates the ZIP package before it unpacks anything, so a malformed archive is rejected without producing a list.
- Wait while the browser reads the external hyperlink relationships from the document package. There is no upload and no network call — the relationship XML is parsed locally in your tab.
- Review the safe destinations shown on the page. Each line is one external URL, with duplicates removed.
- Download the TXT report using the on-page button. The file contains one URL per line, ready to paste into a spreadsheet, URL checker, or any other local text workflow.
The whole process is deterministic: the same .docx always produces the same URL list. There is no scoring, no reputation ranking, and no crawl. The tool only reads the relationship XML that Word itself wrote, then returns the http, https, and mailto targets as plain text.
What the Tool Includes and What It Skips
The extraction is intentionally conservative. The table below summarises the relationship targets the Word Hyperlink Extractor returns and the targets it deliberately drops, along with the reason for each decision.
| Relationship target | Included? | Reason |
|---|---|---|
| http://… destination | Yes | Standard web link recorded as an external relationship |
| https://… destination | Yes | Standard secure web link recorded as an external relationship |
| mailto:… destination | Yes | Mail link recorded as an external relationship |
| Internal bookmark (for example _Toc12345) | No | Points inside the document, not to a web address |
| File-system path or relative URL | No | Local target, not safe to treat as a public web address |
| JavaScript-style or custom protocol URL | No | Not http, https, or mailto; could behave unexpectedly |
| Malformed target value | No | Cannot be parsed cleanly; including it would mislead the audit |
| Duplicate destination after normalisation | No | Each URL appears once so the count stays meaningful |
The "Yes" rows are the only categories the tool is designed to return. Anything else is treated as out of scope on purpose, so the resulting list always represents safe external destinations recorded by Word, not every clickable object Word might display on a page.
Review the URLs Before You Remove Them in Word
Once the TXT report is on your machine, the next stage is review, not deletion. Open the original .docx alongside the URL list and compare each address with the visible paragraph, citation, or image it labels. That comparison matters because Word can also display text that looks like a URL but is not actually a hyperlink, and the extractor will not list it — that is correct behaviour, not a missed link.
For a sensible workflow, work through the list in this order:
- Sort the URLs and skim them for any address you do not recognise. Unknown domains are the most common reason an audit catches a problem before publication.
- Compare the count against your expectations. A 10-page report with five references should not return 200 URLs; a large template with hundreds of citations should not return only three.
- For each URL you want to remove, return to the .docx, find the matching paragraph or shape, and use Word's built-in Remove Hyperlink command, or press Ctrl+Shift+F9 with the link selected, to strip the formatting while keeping the visible text.
- For each URL you want to keep, leave it in place. The extractor never modifies the source document.
For a deeper walkthrough of the same workflow with screenshots and edge conditions, the related guide How to Get All Links from a Word Document covers additional formats and recovery scenarios. If you want a second opinion before publishing, paste the URL list into an independent URL checker, a domain reputation tool, or a spreadsheet pivot table to spot duplicates and unusual top-level domains. Because the TXT file contains one URL per line, it drops into almost any text-processing workflow you are already using.
Privacy and Local Processing
The extractor runs as a single-page browser tool. The .docx file you select is read locally by JavaScript inside your tab; it is not uploaded to a server, and the relationship XML is parsed in memory only after you choose the file. The browser also rejects malformed archives, oversized relationship XML, unsupported ZIP64 directory records, and any package that exceeds its local size limits before it shows a result, which keeps a focused extractor from turning into a general document parser.
Because the tool never visits the URLs in the list, it cannot tell you whether a destination is currently live, whether a domain is trustworthy, or whether a link is safe to click. Treat each listed address as a document-supplied string that still needs human review. For sensitive audits, keep the original .docx, the TXT report, and your notes together so the trail is reproducible if anyone questions a removal decision later.
If you're weighing options, How to See Invisible Characters in Word (Beyond Show/Hide) covers this in detail.