Remove URLs from Text
Paste a block of text and every URL — whether it starts with https://, http://, or is buried inside Markdown syntax — is stripped out. The surrounding sentences stay intact.
Hidden character report
Paste text to scan
When bulk URL removal actually matters
The situation comes up more often than people expect. You copy a section from a web article to paste into an email, and the pasted block arrives with a dozen raw links breaking across lines — https://example.com/some/very/long/path?ref=homepage&utm_source=newsletter — right in the middle of a sentence. The recipient's email client turns each one into an underlined blue string that pulls the reader's eye away from the point you were making. Deleting them by hand, one at a time, is the kind of task that takes ten minutes when it should take ten seconds.
The same problem appears when you pull paragraphs from a PDF into a Word document. PDF-to-text converters are inconsistent about whether they preserve hyperlinks as live objects or as bare URL strings. If they land as bare strings, you end up with visible citations like https://doi.org/10.1000/xyz123 scattered through prose that was supposed to read cleanly. Academic writers doing lit-review notes hit this every week. So do editors who aggregate content from multiple sources before a human rewrites it.
A third common trigger is social-media archiving. Exported tweets or LinkedIn posts come with URLs expanded to their full tracking form — often 80 or 90 characters of utm_campaign and fbclid parameters. Before feeding that text to any analysis tool, or before sharing a cleaned version with a colleague, the URLs need to go. This tool handles all three scenarios in one paste.
Link stripping versus link extraction — two different jobs
Link extraction means pulling URLs out of text so you can do something with them: build a list, run them through a crawler, check for broken links. The URLs are the product. Link stripping — what this tool does — means the opposite: you want the readable prose and the URLs are noise. The output you care about is the text with the gaps closed, not a spreadsheet of addresses.
Confusing the two leads to using the wrong tool. A link extractor will give you a list of URLs and discard the surrounding text. If you then paste that list back into your document, you have accomplished nothing. The correct tool for stripping is one that operates on the text as a whole, removes the URL tokens, and hands back a coherent block of prose — collapsing any whitespace the removal leaves behind so sentences do not start with a stray space.
There is a middle case worth naming: sometimes you want to replace each URL with a placeholder like [link] rather than deleting it entirely. That approach preserves sentence structure when the URL was carrying meaning — for example, a sentence like 'Read the full report at https://example.com/report' becomes 'Read the full report at [link]' rather than the grammatically incomplete 'Read the full report at'. This tool removes URLs completely, so if your text relies on URLs for grammatical completion, a quick manual pass after the fact will catch those cases.
Why the simple regex https?://\S+ is not enough
The pattern https?://\S+ is the first thing any programmer reaches for, and it works on the easy cases. But \S+ matches any non-whitespace character, which includes punctuation that is not actually part of the URL. A sentence ending with a URL followed by a period — 'See the docs at https://docs.example.com.' — will have the period consumed into the match. The cleaned text then reads 'See the docs at' with no sentence-ending period, and you have introduced a punctuation error into prose you were trying to clean.
The same problem occurs with parenthetical citations. Wikipedia-style text often wraps sources in parentheses: '(https://example.com)'. A naive \S+ match grabs the closing parenthesis as part of the URL. After removal you get '(' floating alone. Right parentheses, right square brackets, and commas are the most common offenders. A robust stripper unwinds trailing punctuation — checking whether the character after a potential URL end is a sentence-final mark rather than a path component — before finalizing the match boundary.
There is also the question of bare domains. Some text contains addresses without a protocol prefix: 'Visit example.com for details.' A pure https?:// pattern misses these entirely. Whether you want to remove bare domains depends on context — a block of technical documentation might legitimately mention domain names as concepts, not as clickable links — so this tool targets protocol-prefixed URLs by default, which covers the vast majority of real-world cases. If you also need to strip bare domains, running the result through a second pass with a looser pattern is the cleanest approach.
Short links and UTM-heavy URLs — handling the two extremes
Short links like bit.ly/AbCd3 or t.co/xyz99 are structurally identical to any other URL as far as pattern matching is concerned — they have a protocol, a domain, and a path, just a very short one. The regex handles them the same way. The practical difference is interpretive: a short link carries no readable information about its destination, so there is no risk of stripping something that a human reader needs in order to understand the sentence. With a full URL like https://company.com/blog/q3-report, you might want to preserve the path for context. With t.co/xyz99, there is nothing to preserve.
UTM-tagged URLs are the opposite extreme. A link shared through an email campaign might look like https://example.com/landing?utm_source=newsletter&utm_medium=email&utm_campaign=october2026&utm_content=cta-button. That is 120 characters of tracking parameters after a question mark, none of which is meaningful to the reader. The tool removes the entire URL including its query string, so the tracking noise disappears along with the base address. If your goal is to keep the base URL and strip only the tracking parameters, that is a different operation — parameter cleaning rather than URL removal — and would need a different tool.
A related scenario is URLs that appear inside angle brackets in plain-text email format: <https://example.com>. Some mail clients and API responses serialize links that way. The angle brackets are not part of the URL itself, but after the URL is stripped they remain as an empty pair of angle brackets. This tool removes the URL characters, and the leftover < and > are then just stray punctuation. For clean output, a final pass to remove isolated angle brackets is worth adding if your source text uses that format consistently.
Markdown links and the dangling bracket problem
Markdown uses the syntax [link text](https://example.com) for hyperlinks. A stripper that only targets the URL portion removes https://example.com and leaves [link text]() — a broken Markdown link with an empty parenthetical. If the output will be rendered as Markdown, that broken syntax produces no visible link but the square-bracketed text remains, which is fine. If the output is plain text destined for an email or a Word document, the square brackets look like formatting debris.
The cleaner approach for Markdown text is to replace the entire [text](url) construct with just the link text — stripping the URL and the surrounding syntax together, leaving the readable words behind. That way 'Check the [release notes](https://github.com/org/repo/releases) for details' becomes 'Check the release notes for details', which reads naturally. The distinction matters most for README files, documentation drafts, or content exported from tools like Notion or Obsidian, which use Markdown internally.
Images in Markdown follow the same pattern with an exclamation mark: . If you are stripping URLs from Markdown, image references are worth addressing separately — removing the URL here also removes the image from the rendered output, which may or may not be what you want. This tool removes protocol-prefixed URLs; if your Markdown image references need to survive, a targeted substitution that replaces only [text](url) patterns while leaving  patterns intact is the more precise approach. For most prose-cleaning tasks, where the goal is readable text rather than valid Markdown, that level of distinction is unnecessary.
What URL removal does not clean up
Removing the URL from a sentence does not remove every piece of tracking information that may have traveled with it. If someone copied a block of text from a web page that injects invisible Unicode characters around links — zero-width spaces, word joiners, or directional markers — those invisible characters survive URL removal and remain embedded in the surrounding text. They will not cause visible problems in most contexts, but they can break string matching in databases or produce unexpected results when the text is processed programmatically.
Path components sometimes encode personally identifiable information. A URL like https://app.example.com/users/jane.smith/settings contains a username that is part of the URL string itself. After stripping, the username is gone too — which is usually the right outcome if your goal is to anonymize the text. But if the URL appeared in a context like 'User jane.smith updated her profile at https://app.example.com/users/jane.smith/settings', the name appears twice: once in the prose and once in the URL. Stripping the URL removes one instance, not both.
For text that needs thorough cleaning before sharing externally — say, a support ticket transcript going to a vendor — URL stripping is a useful first pass but not a complete privacy review. After removing URLs, a scan for email addresses, phone numbers, and names in the surrounding prose is still worthwhile. The invisible character remover on this site handles the Unicode residue problem, and combining the two tools covers the most common cases. For anything requiring compliance-level review, human eyes on the output are still the final check.
Characters this tool looks for
25 invisible characters and formatting controls, with the action taken on each one.
| Character | Code point | Action | What it does |
|---|---|---|---|
| Zero-width characters | |||
| Zero Width Space | U+200B | Removed | Invisible separator with no width. Commonly left behind when text is copied out of a chat window. |
| Zero Width Non-Joiner | U+200C | Removed | Invisible character that blocks two letters from joining in scripts such as Arabic or Devanagari. |
| Zero Width Joiner | U+200D | Removed | Invisible character used to glue emoji sequences together. Stray copies can break word wrapping. |
| Zero Width No-Break Space (BOM) | U+FEFF | Removed | Byte order mark. Harmless at the start of a file, invisible noise anywhere else. |
| Word Joiner | U+2060 | Removed | Tells a text engine not to break a line here. Invisible, and easy to carry along by accident. |
| Mongolian Vowel Separator | U+180E | Removed | A legacy invisible separator from Mongolian script that now renders as nothing at all. |
| Bidirectional controls | |||
| Left-to-Right Mark | U+200E | Removed | Directional hint for mixed-script lines. Invisible on screen but present in the character stream. |
| Right-to-Left Mark | U+200F | Removed | Mirror of the left-to-right mark, used to steer the order of mixed-direction text. |
| Arabic Letter Mark | U+061C | Removed | Invisible directional marker for Arabic text. Rarely intentional outside its original document. |
| Left-to-Right Embedding | U+202A | Removed | Opens an embedded left-to-right run. Unpaired copies can scramble the order of a sentence. |
| Right-to-Left Embedding | U+202B | Removed | Opens an embedded right-to-left run. Left behind after copy-paste, it can flip nearby punctuation. |
| Pop Directional Formatting | U+202C | Removed | Closes an embedded directional run. Without its partner it is pure invisible noise. |
| Left-to-Right Override | U+202D | Removed | Forces following characters to display left-to-right regardless of their own direction. |
| Right-to-Left Override | U+202E | Removed | Forces following characters to display right-to-left. A common cause of text that looks reversed. |
| Left-to-Right Isolate | U+2066 | Removed | Modern replacement for the embedding controls. Still invisible, still travels with copied text. |
| Right-to-Left Isolate | U+2067 | Removed | Isolates a right-to-left run from the surrounding paragraph direction. |
| First Strong Isolate | U+2068 | Removed | Isolates a run whose direction is decided by its first strong character. |
| Pop Directional Isolate | U+2069 | Removed | Closes an isolate opened by one of the isolate controls above. |
| Formatting controls | |||
| Soft Hyphen | U+00AD | Removed | A hyphen the reader only sees if the word happens to break at the end of a line. |
| Carriage Return (CR) | U+000D | Removed | Half of a Windows line ending. Kept on its own it produces doubled spacing in most editors. |
| Markdown Asterisk Marker | U+002A | Removed | Bold or italic marker such as ** left behind when text is copied in its raw form. Turn this group off to keep Markdown styling marks. |
| Markdown Backtick Marker | U+0060 | Removed | Inline-code backtick left behind after copying raw assistant output. Turn this group off to keep inline code marks. |
| Invisible spaces | |||
| No-Break Space | U+00A0 | → space | Looks exactly like a space but stops a line from breaking there. A frequent copy-paste leftover. |
| Figure Space | U+2007 | → space | A space as wide as a digit, used to align numbers in tables. Replaced with a normal space. |
| Narrow No-Break Space | U+202F | → space | A thin, non-breaking space from French typography. Replaced with a normal space. |
Frequently Asked Questions
▸Does the tool remove URLs that start with http:// as well as https://?
Yes. Both http:// and https:// prefixed URLs are matched and removed. The underlying pattern targets the protocol prefix, so any URL beginning with either variant — including less common forms like http://localhost:3000 used in development environments — will be stripped.
▸Will it remove URLs that appear in parentheses without leaving stray brackets?
The tool removes the URL string itself. If a URL appears inside parentheses — such as (https://example.com) — the parentheses are not part of the URL and remain after the URL is stripped. You would see empty parentheses () in the output. A quick find-and-replace for () afterward handles that case if it comes up in your text.
▸What happens to Markdown image syntax like ?
The URL portion inside the parentheses is removed, leaving ![alt](). For most plain-text cleaning tasks this is acceptable. If you need to preserve Markdown image references while stripping other URLs, the remove Markdown tool can help flatten the entire Markdown layer first, after which a URL strip leaves clean readable text.
▸Does it handle URLs that wrap across two lines?
No. The tool processes URLs as continuous non-whitespace token sequences. A URL that has been broken across a line boundary by a line break character appears as two separate fragments — the first ending with a hyphen or slash, the second starting mid-path — neither of which looks like a complete URL to the pattern. You would need to remove the line break first using the remove line breaks tool, then run the URL removal on the rejoined text.
▸Can I use this to clean text before feeding it to an AI model?
Yes, that is a common use case. Raw URLs in a prompt or document add token count without adding meaning. Stripping them before submitting the text to an AI reduces noise and keeps the input focused on the readable content. It is particularly useful when processing exported articles, email threads, or scraped pages where the link density is high.
▸Does the tool also remove email addresses?
No. Email addresses do not start with http:// or https://, so they are not matched by the URL pattern. If you need to remove email addresses from text, that requires a separate pattern targeting the user@domain.tld structure. This tool is focused specifically on web URLs and does not affect email addresses, phone numbers, or other contact information.