URL Encoder / Decoder
Enter some text above to see the encoded results.
Percent-Encoding Mechanics: How RFC 3986 Translates Unsafe Characters
Uniform Resource Identifiers (URIs) depend on a restricted set of US-ASCII characters to guarantee dependable transmission across web servers, proxies, and client browsers. When a web address contains characters outside this allowed set, or includes characters with reserved syntactic roles, standard percent-encoding transforms those bytes into a safe triplet format.
The mechanism operates at the byte level. Each character outside the unreserved character set converts into its underlying UTF-8 byte sequence. The encoder replaces each byte with a percent symbol (%) followed by two hexadecimal digits representing the byte value. For instance, a standard space character has a decimal ASCII code of 32, which translates to hexadecimal 20, producing %20 in the final URL.
Multi-byte international characters and symbols follow the exact same rule across their full byte representation. The character ü occupies two bytes in UTF-8 (hexadecimal C3 and BC), producing %C3%BC. Modern search engine crawlers and HTTP parsers require this strict octet representation to process international domain paths and UTF-8 query parameters without losing data.
Character Classification in Uniform Resource Identifiers
RFC 3986 divides all characters into unreserved characters and reserved characters. Unreserved characters never need encoding because they carry no structural meaning in URI grammar. Reserved characters serve as syntax boundaries, such as separating the host from the path or delimiting individual query parameters. When a reserved character is part of the actual data payload rather than a structural boundary, you must percent-encode it.
The unreserved character set consists of uppercase letters (A-Z), lowercase letters (a-z), decimal digits (0-9), hyphens (-), underscores (_), periods (.), and tildes (~). All other characters, including whitespace, control characters, and reserved delimiters, require percent-encoding when used inside query values or path segments.
RFC 3986 Reserved and Unsafe Character Reference
Common reserved characters, their percent-encoded equivalents, and how unencoded data disrupts HTTP parsers.
| Character | Name | RFC 3986 Hex | URI Classification | Parser Impact If Unencoded |
|---|---|---|---|---|
| Space | Space | %20 | Unsafe whitespace | Breaks request line parsing and triggers HTTP 400 Bad Request errors |
| & | Ampersand | %26 | Sub-delimiter | Prematurely splits query strings into multiple unintended parameters |
| = | Equal sign | %3D | Sub-delimiter | Breaks key-value parameter pairing syntax |
| ? | Question mark | %3F | General delimiter | Starts a new query string prematurely and disrupts URL routing |
| # | Number sign (hash) | %23 | General delimiter | Marks start of fragment; servers drop all trailing characters |
| / | Forward slash | %2F | General delimiter | Creates false path hierarchies and breaks endpoint matching |
| : | Colon | %3A | General delimiter | Confuses port number, scheme, or auth credential boundaries |
| + | Plus sign | %2B | Sub-delimiter | Deserialized into an empty space by form-urlencoded parsers |
| % | Percent sign | %25 | Escape prefix | Causes invalid percent-escape errors or double-encoding defects |
| @ | At sign | %40 | General delimiter | Interpreted as userinfo credential boundary in URI authorities |
The Space Character Dilemma: %20 Versus the Plus (+) Sign
One of the most frequent sources of data corruption in web APIs is the confusion between %20 and the plus sign (+). RFC 3986 specifies %20 as the universal percent-encoded representation of a space character across the entire URI, including path segments and query components.
In contrast, the plus sign convention comes from the W3C HTML form specification under the MIME media type application/x-www-form-urlencoded. In standard web forms, browser engines replace spaces with plus signs before submitting form fields. Many backend web frameworks, including Express, Spring, Django, and PHP, automatically decode a plus sign in a query string into an empty space.
This dual behavior causes bugs when handling literal plus signs. If an API endpoint receives an email alias like dev+updates@domain.com or a phone number like +14155552671 without encoding the plus sign as %2B, the receiving server decodes the plus sign as a space, producing dev updates@domain.com and invalidating the request data.
Component Encoding vs. Full URI Encoding in Production Code
Developers building web applications frequently encounter defects by using the wrong encoding method for their data context. JavaScript provides two native functions, encodeURI() and encodeURIComponent(), each designed for a different architectural task.
Use encodeURI() when preparing an entire URL string with existing path structures and query delimiters that must remain intact. Use encodeURIComponent() when handling individual query parameter values or route variables. In modern client and Node.js code, constructing query strings with URLSearchParams handles these distinctions automatically.
Troubleshooting Double Encoding and Gateway Deserialization Errors
A common bug in microservice architectures and reverse proxy setups is double encoding. This happens when an already percent-encoded value passes through a second encoding pass. Because the percent symbol (%) is itself an active escape prefix with hex value 25, the secondary encoder converts %20 into %2520, and %26 into %2526.
Intermediate proxies, API gateways (such as Nginx, AWS CloudFront, or Cloudflare), and load balancers often normalize incoming request paths before forwarding them upstream. If a gateway decodes a path segment once, but your downstream application expects raw encoded data, routing mismatches and signature verification failures in HMAC or OAuth2 headers occur.
Our client-side URL Encoder and Decoder helps you isolate these issues instantly. By inspecting and toggling your string between encoded and decoded formats right in your browser, you verify whether your API payloads match RFC 3986 requirements before dispatching network requests.