Understanding URL Encoding (Percent-Encoding) & Standards
Uniform Resource Identifiers (URIs) and web URLs are restricted to a specific subset of ASCII characters defined by internet specification RFC 3986. Any character outside this allowed set—such as non-English alphabets, spaces, symbols, or reserved control characters—must be translated into percent-encoding (also known as URL encoding) to be safely transmitted across the World Wide Web.
encodeURI vs. encodeURIComponent: When to Use Which?
In web development, choosing the correct encoding strategy is essential to prevent broken links or corrupted query strings:
- encodeURIComponent: Encodes all special characters including structural URL delimiters like
/,?,&,=, and#. Use this method when encoding individual query string parameter values (e.g.,?query=val). - encodeURI: Preserves functional URL structural symbols (such as
http://, slashes, and question marks) while encoding invalid characters like spaces or accent marks. Use this method when encoding a complete full URL.
Reserved vs. Unreserved URL Characters
RFC 3986 divides characters into two distinct buckets:
Unreserved Characters
Unreserved characters never require encoding. They include upper/lowercase letters (A-Z, a-z), numbers (0-9), and four special symbols: hyphen (-), period (.), underscore (_), and tilde (~).
Reserved Characters
Reserved characters hold structural significance in URLs (e.g. ? for query strings, # for fragments, & for key-value delimiters). When passed as plain data values, they must be percent-encoded (e.g. & becomes %26).
Common Percent-Encoding Reference Table
-
Space
%20 or + (Form Encoding) -
Ampersand (&)
%26 -
Equals (=)
%3D -
Question Mark (?)
%3F -
Forward Slash (/)
%2F -
Hash / Fragment (#)
%23