Definition:
Dark data is data that an organisation collects, processes or stores during its normal activity but does not use effectively for operations, analysis or decision-making. It may include known but inactive information, datasets whose owner or purpose has been lost and data that the organisation has not properly inventoried.
The concept belongs to the field of Big Data, although it does not depend on volume or a particular format. Dark data can be structured, semi-structured or unstructured. It is defined by the absence of use and governance, not by a lack of labels or technical accessibility.
Scope of Dark Data
An organisation may know that data exists but not use it because there is no purpose, owner or process for doing so. In other cases, information is distributed across applications, accounts, devices, backups and repositories that do not appear in the corporate inventory.
The iceberg analogy represents the difference between visible, used data and information that remains below the surface. There is no universal percentage that determines how much of an organisation’s data is dark, useless or potentially valuable. The proportion depends on its activity, architecture, policies and management capability.
A data lake can store information in different formats without automatically turning it into dark data. Likewise, a data warehouse can contain tables that are no longer used. The status depends on the knowledge, purpose and actual use of each dataset.
Causes of Dark Data
Dark data commonly accumulates through a combination of technical and organisational decisions:
- Collection without a defined purpose: Systems, forms and devices retain data because they can collect it, even when there is no use case or agreed retention period.
- Silos and lost context: Data remains isolated in tools, departments or accounts without an owner, documentation or links to other systems.
- Technology changes: Migrations, application replacements and legacy formats leave files and copies that remain stored but no longer belong to active processes.
- Lack of capability or priority: The organisation knows about the data but lacks sufficient quality, permissions, tools, time or expertise to evaluate and use it.
Data that was never collected is not dark data. The term applies to information that already exists within the organisation’s environment, even when it is difficult to locate or interpret.
Sources of Dark Data
Examples vary by activity but usually fall into four sources:
- Operations and systems: Log files, telemetry, events, errors, access histories and data generated by devices or applications.
- Communications and content: Emails, chats, documents, recordings, images, videos, transcripts and intermediate versions.
- Commercial relationships: Forms, customer service interactions, CRM histories, surveys, requests and account-related data.
- Administrative and legacy files: Backups, exports, old databases, financial documents, former employee records and duplicated or obsolete files.
A file type is not dark data by nature. One log may be essential for security or diagnosis, while another identical file may have no purpose and be retained longer than necessary.
Dark Data management
Managing this data does not mean analysing all of it. The process needs to establish what exists, which obligations apply to each dataset and which decision is justified:
- Inventory: Locate repositories, sources, copies, owners, formats, access permissions and information flows.
- Classification: Identify purpose, sensitivity, quality, age, duplication, restrictions and retention requirements.
- Decision: Retain necessary data, activate data with a legitimate use, archive information that must be preserved and delete data without a purpose or obligation.
- Control: Apply permissions, encryption, traceability, retention periods and periodic reviews throughout the lifecycle.
Data mining techniques can help find patterns in particular datasets, but they should be applied after checking origin, quality, permissions and purpose. The fact that data can be analysed does not mean that it is reliable, relevant or lawful to use for every objective.
Dark Data risks
Uncontrolled storage consumes capacity and can multiply copies, versions and operating costs. It also makes it harder to determine which information is current and increases the chance that later analyses combine incomplete, duplicated or out-of-context data.
Forgotten repositories can contain personal data, credentials, contracts or confidential information with inadequate permissions. Lack of use does not remove security, access, retention and deletion obligations. Keeping information indefinitely is not, by itself, a compliance measure.
Some dark data may become valuable when a new question arises or analytical capabilities improve. Other data will remain useless or need to be deleted. Its assessment should balance utility, quality, cost, risk and a legitimate basis for retention without assuming that all hidden data represents a missed opportunity.
