Definition:
An algorithm is an ordered set of instructions for performing a task or solving a problem. In computing, it describes the steps used to process data and obtain a result, such as sorting a list, finding information or calculating a route.
A search algorithm is an application of this concept: it retrieves information related to a query. A web search engine uses different algorithms and systems to interpret the request, select documents and rank results.
Search engines update these systems to improve the relevance and quality of results, as well as to combat spam and manipulation. These changes are part of the context analyzed in search engine optimization work.
Table of contents
How search algorithms work
When a query is submitted, a search engine must find potentially useful information among the content in its index. To do so, it may combine term matching with systems that interpret the meaning and intent of the search.
A common structure in information retrieval is the inverted index: a list of the documents in which each term appears is maintained. This makes it possible to find candidates without scanning every document from the beginning.
In a search that requires several words to be present, the intersection of their lists can be calculated. However, this procedure does not by itself describe how a modern web search engine works: it may also recognize variants, relationships between concepts and other indications of relevance.
After retrieving candidates, the systems rank them using signals related to the query and the documents. These may include content, links, location or the freshness of information when relevant. Results may be presented in different formats, not necessarily in fixed groups of ten links.
Google describes several ranking systems, not a single algorithm with a public, unchanging list of factors. Clear content and an appropriate user experience help make information accessible and useful, but do not guarantee a particular position.
Some well-known algorithms and applications on the internet
Algorithms are used in search engines, social networks and other digital services. A specific algorithm should be distinguished from a system that combines several algorithms or a field in which they are used:
- Google PageRank: analyzes the structure of links between pages. From its original formulation, it considers not only the number of incoming links, but also the importance of the linking pages and how they distribute their links. PageRank has evolved and remains part of Google’s systems, even though its former public indicator is no longer available. It does not determine the order of results on its own.
- Facebook feed ranking: EdgeRank is a historical reference. Current systems combine signals and predictions to rank posts and recommend content for each user. They cannot be described as a single fixed formula or by one change related to false news.
- High-frequency trading (HFT): is a field of financial trading that uses automated systems to process information and submit orders with very low latency. It is not the name of a single algorithm; different strategies use different procedures.
These examples illustrate different functions: evaluating relationships between documents, selecting content or executing operations. Not all algorithms use artificial intelligence or learn from data while they run.
Algorithm optimization
Optimizing an algorithm means improving aspects such as execution time or memory use while retaining the requirements it must meet. In a search engine, the data structures and infrastructure that execute queries can also be optimized.
Techniques used in information retrieval systems include the following:
- Work distribution: spreading documents and operations across multiple machines to process tasks in parallel. The outcome also depends on coordination and the cost of communication between them.
- Memory and caching: keeping frequently accessed data available or reusing previously calculated results when they remain valid.
- Processing order: for certain queries requiring several terms, starting with the shortest document lists may reduce the number of comparisons needed.
- Candidate pruning: using bounds and scoring criteria to avoid processing documents that cannot make it into the top results. Sorting solely by PageRank does not solve this problem in general.
- Term and phrase information: storing position data or word combinations may facilitate certain queries, at an additional storage cost.
- Compression: reducing the space occupied by indexes and the volume of data transmitted, while accounting for the work needed to compress and retrieve them.
These techniques concern the design of the search system. They are not instructions a website owner can apply to modify Google’s algorithm. Search engine optimization works on content, technical accessibility and other aspects of a website, without controlling the search engine’s systems.
