3 4 5 A B C D E F G H I J K L M N O P Q R S T U V W X Y Z

What is Big Data

Big Data Definition:

Big Data refers to data sets and flows whose scale, speed, diversity or complexity require methods and infrastructure different from those used in conventional processing.

The concept covers data capture, integration, storage, querying, analysis, governance, security and communication. It may combine structured, semi-structured and unstructured information from systems, sensors, transactions, documents, images or digital interactions.

Big Data is associated with techniques such as predictive analytics, but accumulating data does not automatically produce knowledge or correct decisions. Usefulness depends on quality, representativeness, purpose, the analytical method and the ability to interpret results.

The size of Big Data

Big Data is not defined through a fixed number of terabytes or petabytes. A volume may be manageable for one organisation and exceed the acceptable capacity, time or cost of another. A Big Data problem can also exist at a lower volume when data arrive rapidly, have highly diverse formats or require complex relationships.

Architectures may distribute storage and processing across several systems, use cloud services or combine batch and real-time processing. The choice depends on the use case, required latency, cost, security and available skills. A relational database may remain suitable for part of the system and coexist with other technologies.

The three Vs of Big Data

The classic model describes Big Data through three characteristics:

  • Volume: The amount of data that needs to be stored, processed or transferred within the project’s constraints.
  • Velocity: The rate at which data are generated, received and processed, and how quickly a response is required.
  • Variety: The diversity of sources, structures, formats and meanings that need to be integrated.

Additional Vs have been proposed over time to describe further challenges. There is no single universal extended list, but the following are common:

  • Veracity: Data quality, reliability, traceability and uncertainty.
  • Value: The ability to use data for a particular purpose and obtain a benefit that justifies the cost and risk.
  • Variability: Changes in the meaning, distribution or behaviour of data over time.
  • Visualisation: Ways of representing complex results to support exploration and communication without distorting them.
  • Volatility: The period during which data remain useful and the conditions for their retention or deletion.

Examples of Big Data use

Applications depend on the availability and legitimate use of the data, the problem being addressed and validation of the results. Examples include:

  • Digital: In digital marketing, Big Data can support journey analysis, audience segmentation and campaign measurement. It may also support recommendations and improvements to user experience, provided personalisation respects consent and correlation is not confused with intent.
  • Health and medicine: Analysis of clinical, epidemiological or genomic data may support research, surveillance and decision-making under strict privacy, quality and medical validation controls.
  • Marketing and sales: Customer, product and channel data may be used to study demand, attribution, satisfaction and responses to offers while avoiding discriminatory profiles or unsupported inferences.
  • Finance: It is used in fraud detection, risk assessment, compliance and transaction monitoring, with oversight of errors, bias and false positives.
  • Retail: It can support demand forecasts, inventory, assortment and pricing. Models need to adapt to seasonal changes, promotions and events absent from historical data.
  • Transportation and logistics: It supports the analysis of routes, capacity, delivery times and equipment maintenance using operational and sensor data.
  • Energy: It can be used to monitor grids, forecast demand, detect incidents and combine weather data with renewable generation.

Big Data evaluation

The value of a project depends on its governance as much as on its infrastructure. Major challenges include:

  • Privacy and security: Collection needs a purpose and legitimate basis, access needs to be limited and protection measures applied throughout the life cycle.
  • Data quality: Errors, duplicates, missing values, bias or changing definitions can produce misleading results even when the volume is large.
  • Data management: Architecture, cataloguing, permissions, traceability, retention and costs require clear responsibilities and processes.
  • Interpretation of results: A statistical pattern does not prove causation. Conclusions need to include uncertainty, context, limitations and validation against the objective.
  • Data integration: Combining sources requires identifiers, formats, timing and meaning to be resolved without losing provenance or creating incorrect associations.