Python parses an absolute URL with the urllib.parse module, whose urlparse function returns a six-field named tuple containing scheme, netloc, path, params, query, and fragment. For the query portion, parse_qs and parse_qsl add a second pass that splits the query string into a Python dict or an ordered list of pairs. Most Python tutorials stop at that two-call pattern and never look again, which is fine for round-tripping RFC 3986 URLs but routinely produces surprising results when the input contains a non-ASCII hostname, a default port written explicitly, or a query parameter whose value uses plus signs to encode spaces. The browser-side counterpart, the WHATWG URL parser, normalizes those cases differently, and most cross-language bugs trace back to the gap between the two grammars. URL Parser runs the WHATWG algorithm directly in your browser tab and exposes every field the JavaScript URL and URLSearchParams objects expose, so you can paste a suspect URL, see exactly what the browser sees, and copy a normalized JSON result to compare against urlparse(...)._asdict().

how to parse url in python
Parse a URL in Python and Verify It Against the Browser

Python's urllib.parse Returns Six Top-Level Fields

This is the Python standard-library path and the one most code reviews will expect to see. Importing urllib.parse gives you urlparse, which accepts the URL as a string and returns a SplitResult or ParseResult named tuple. The six attributes are stable across Python versions and behave identically for ASCII inputs, so they remain a reasonable default for tasks like routing, logging, or pulling the host out of a referrer. The named tuple is read-only, but you can convert it to a dict with the private _asdict helper or by feeding it to dict(). The split variant adds a separator keyword that lets you parse URL-like strings whose query delimiter is not a question mark, an option urllib.parse documents but most tutorials skip.

There are exactly two situations where urlparse silently gives you a wrong answer: when the host contains internationalized characters that the WHATWG parser would Punycode-encode, and when the input carries a username or password inside the netloc. urlparse preserves the userinfo portion of the netloc as-is and offers no attribute for stripping it, so any https://user:[email protected]/ string slips through with the credentials embedded in netloc and lands in your logs without warning.

Where urllib.parse and the Browser Parser Disagree

The disagreement is not a bug in either library; it is two grammars maintained by different groups. The table below maps the urllib.parse attribute on the left to its closest counterpart in the URL Parser output on the right, with the practical difference noted in the third column.

urllib.parse (Python) URL Parser (WHATWG) Practical difference
scheme protocol urllib returns the lowercase scheme name; WHATWG appends the trailing colon, returning "https:".
netloc host, hostname, port urllib returns one combined string; WHATWG splits into host (with port), hostname alone, and the separate port value.
path pathname urllib returns the raw path part; WHATWG normalizes the leading slash and preserves percent escapes byte-for-byte.
params (no equivalent) urllib's deprecated semicolon-params field has no WHATWG counterpart and is rarely useful.
query search urllib strips the leading "?"; WHATWG preserves the leading "?" inside the search field.
fragment hash urllib strips the leading "#"; WHATWG preserves the leading "#" inside the hash field.
(no equivalent) origin WHATWG adds origin as scheme + hostname + effective non-default port; urllib requires manual assembly.
(no equivalent) filename WHATWG adds filename as the final pathname segment; urllib requires a manual split on "/".

The two parsers also differ on three behaviors that affect any code doing conditional routing, IDN handling, or query decoding. First, default-port stripping: the input host example.com:443 is 15 characters, and the normalized URL Parser host becomes example.com, which is 11 characters, a 4-character reduction that comes entirely from the stripped colon and digits. Second, internationalized hostnames: a URL containing 例え.jp produces an example.xn--r8jz45g.xn--zckzah ASCII form in the URL Parser hostname field, while urllib.parse preserves the original Unicode characters inside netloc. Third, plus-sign decoding: parse_qs leaves plus signs as literal characters in decoded values, but URL Parser decodes them as spaces because URLSearchParams follows the WHATWG form-urlencoded rule. That difference shows up most often in search and OAuth callbacks where plus was used as a space substitute.

Parse an Absolute URL

Use URL Parser whenever you want to inspect what the browser would actually do with a URL string that misbehaves in your Python code. The tool runs entirely inside your current browser tab and never sends the URL anywhere, so it is safe for production URLs that contain internal hostnames or signed query tokens you do not want to leak.

  1. Paste one absolute HTTP or HTTPS URL into the input area, without a username or password, and review the input character count before parsing. The input cap is exactly 8,192 UTF-16 code units, and any surrounding whitespace, raw backslash, or ASCII control character causes the URL to be rejected rather than silently trimmed.
  2. Select Parse URL and inspect the normalized components, including protocol, origin, host, hostname, port, pathname, search, hash, and filename. The serialized filename is the final segment after the last slash; a pathname that ends in a slash produces an empty filename rather than a directory marker.
  3. Read the ordered, decoded query parameter list, review any potentially sensitive query or hash data, and copy the complete parsed JSON to your clipboard if appropriate. Editing the input clears the old result, so a failed URL never leaves an earlier successful parse visible alongside a new error.

If you also need to focus only on the query portion without re-pasting the whole URL, the same URLSearchParams decoding rules apply to the dedicated workflow described in Parse Query Strings in JavaScript the Browser-Native Way.

Reading the Parsed JSON Output

The output JSON is built from the same fields the URL and URLSearchParams browser objects expose. Every entry is two-space formatted and the complete document is capped at 50,000 UTF-16 code units; if your URL would produce more than that, the whole result is rejected rather than truncated. A typical parse of https://example.com:443/path/to/page?tag=python&tag;=web&q;=url+parsing#section produces protocol "https:", origin "https://example.com", host "example.com", hostname "example.com", port "", pathname "/path/to/page", search "?tag=python&tag;=web&q;=url+parsing", hash "#section", filename "page", and a query array of three ordered name-and-value objects: tag equals python, tag equals web, and q equals "url parsing" with the plus sign decoded to a space. The query array preserves order and repeats rather than collapsing duplicate keys into a dict, which is the same anti-duplicate-collapse rule that URLSearchParams follows internally. Percent escapes inside pathname and filename stay percent-encoded so the output remains byte-faithful to the serialized URL; only query names and values are decoded.

Rules URL Parser Enforces Before Producing Output

The tool is deliberately an inspector and refuses to act like anything else. Inputs that begin with anything other than a case-insensitive http:// or https:// prefix are rejected before any URL object is constructed, which means a path like /docs/page, a scheme-relative reference like //example.com/path, and a bare domain like example.com all produce a clear error rather than being resolved against the current page. After the prefix check, only http: and https: are accepted; javascript:, data:, file:, ftp:, blob:, mailto:, and other schemes the browser URL API would otherwise parse are still rejected, because presenting them as if they shared an HTTP origin, host, port, and query model would mislead anyone copying the output.

URLs containing a nonempty username or password are rejected outright: the tool does not redact, mask, or partially display userinfo, so a URL like https://user:[email protected] returns a credentials error and produces no parsed fields at all. The query parameter array is capped at exactly 200 entries; the 201st name-value pair is rejected, and no earlier entries are silently sampled or skipped. ASCII control characters are rejected before URL construction so the browser cannot silently strip a newline or tab from pasted text. The tool never navigates to the supplied URL, never performs a reachability check, never tests DNS, and never claims to verify that a site is safe or that a link is malicious.

If you're weighing options, Convert XML to JSON in Python and in the Browser covers this in detail.