Definition:
A URI (Uniform Resource Identifier) is a sequence of characters that identifies a resource through a specified scheme. The resource may be a document, image, service, person, concept, or any other entity that the scheme can distinguish. A URI provides identification; it does not necessarily imply that the resource can be retrieved through a network.
Table of contents
Differences between URI, URL, URN, and IRI
URI is the general concept defined by Internet syntax. Within that framework, several names have been used for different functions or representations:
- URL: identifies a resource by indicating a mechanism or location through which it can be accessed. HTTPS addresses on the web are the most common example. Under the RFC 3986 classification, a URL is a URI.
- URN: is a URI that uses the
urnscheme and a namespace. It is assigned with the intention of providing a persistent, location-independent identifier, as with certain bibliographic identifiers. - IRI: extends the representation of URIs to support characters from the Unicode repertoire. To be exchanged in contexts that expect URI syntax, those characters require appropriate conversion or encoding.
- URI reference: may be a complete URI or a relative reference that is resolved using a base URI.
On the current web platform, browsers apply the WHATWG URL Standard to parse and serialize addresses. This standard generally uses the term URL and defines algorithms compatible with actual browser behavior. The URI concept remains useful for understanding identifiers that do not operate like conventional web addresses.
A URI should not be confused with the display URL used in certain advertisements. The latter is a representation shown to the user and may not match the final address character for character, subject to the advertising platform’s rules.
Components of a URI
The generic syntax in RFC 3986 organizes a URI around a scheme and several optional components. The general form can be expressed as scheme:, although each scheme defines its own requirements.
In https://, the components are:
- Scheme:
httpsindicates how the rest of the identifier should be interpreted. - Authority:
www.example.com:443may contain user information, a host, and a port. Credentials inside an address present risks and should not be exposed. - Path:
/folder/resourceidentifies a hierarchical location within the authority. - Query:
language=ensupplies data that the resource or application may use. - Fragment:
sectionidentifies a secondary part or view. The fragment is not sent to the server in an HTTP request.
Not every URI contains an authority, hierarchical path, or query. For example, mailto:user@example.com uses the mailto scheme and a structure defined specifically for it. A scheme is not always a network protocol: it also establishes how the identifier is interpreted and processed.
Characters, encoding, and relative references
Some characters have reserved functions as separators. Others can be represented through percent-encoding, using a percent sign followed by the hexadecimal value of the corresponding bytes. Encoding or decoding without knowing the component can change the meaning of the identifier, so it should not be applied indiscriminately to the entire string.
A relative reference omits the scheme and is interpreted against a base. From https://, the reference other may resolve to https://, while /other starts from the root of the same host. The . and .. segments take part in this resolution process.
A browser, API, or application should use an appropriate URL parser instead of separating components through simple text operations. Cases involving IPv6, internationalized characters, ports, repeated queries, or encoded paths can produce incorrect results when processed manually.
Equivalence, the web, and security
Two different URIs may identify the same resource, and two similar strings may identify different resources. The scheme and host of an HTTP URL are case-insensitive, but the path may be case-sensitive depending on the server. Default ports, percent-encoding, and certain path segments also affect comparison.
On a website, accessible variants of the same content can fragment crawling and measurement signals. Redirects, consistent internal links, and rel canonical help declare the preferred version, but they do not make all variants identical strings. A slug is only one part of the path and does not represent the complete URI by itself.
A URI also does not prove that a destination is safe or legitimate. The HTTPS scheme protects the connection when validation and configuration are correct, but it does not guarantee that the content is trustworthy. Before using an identifier, it is useful to inspect the scheme, actual host, encoding, possible lookalike characters, and any sensitive data present in the query.
URI provides an extensible syntax for identifying resources across different systems. Understanding its components and rules makes it possible to construct, resolve, and compare references without reducing the concept to a web address or assuming that every identifying string leads to an accessible resource.
