3 4 5 A B C D E F G H I J K L M N O P Q R S T U V W X Y Z

What is Data Scientist

data scientist

Definition:

A data scientist is a professional who uses statistics, programming and contextual knowledge to formulate questions, analyse data, develop and evaluate models and communicate results.

Their work may use structured, semi-structured or unstructured data. The role does not necessarily require large data volumes or machine learning in every project: responsibilities depend on the problem and the organisation.

How a data scientist works

The process begins by turning a general requirement into a question that can be investigated using data. Before selecting a technique, the expected outcome, analysed population, constraints and evaluation criteria need to be defined.

The work may include:

  • Problem formulation: Define the question, related decisions, units of analysis and conditions determining whether the result is useful.
  • Data acquisition and understanding: Locate sources, review their provenance and document how they are generated, updated and related.
  • Preparation and quality control: Correct formats, handle missing values, identify duplicates and verify that transformations do not alter meaning.
  • Exploration and analysis: Describe distributions, relationships and changes, formulate hypotheses and identify information requiring further investigation.
  • Modelling and validation: Select statistical or computational methods, separate training and evaluation data and compare the result with appropriate baselines.
  • Communication and monitoring: Explain methods, uncertainty and limitations, deliver reproducible results and check their performance when they are used continuously.

The United States Bureau of Labor Statistics includes identifying useful data, analysing it, creating and validating models and presenting findings among these professionals’ duties. Its occupational description of data scientists illustrates the breadth of the role without establishing one mandatory workflow.

Differences from other data roles

The boundaries between job titles are not universal. Two organisations may use the same title for different responsibilities or distribute one project across several roles.

A data analyst commonly focuses on queries, reports, visualisations and descriptive or diagnostic analysis. A data scientist may also perform these tasks but typically incorporates experimentation, inference or predictive modelling when the problem requires it.

A data engineer designs and maintains systems for ingestion, transformation, storage and availability. A data scientist consumes these structures and may prepare specific datasets but does not necessarily replace engineering work.

A machine learning engineer commonly focuses on turning models into reliable, scalable and observable software components. A data scientist may create prototypes and evaluate models, while operating them in production also requires engineering, security and operational work.

Business intelligence organises data, indicators and reports to support monitoring and decisions. It may share tools and sources with data science but is not equivalent to every statistical analysis or predictive model.

Data scientist competencies

Statistics supports sample design, relationship estimation, uncertainty quantification and result evaluation. Programming supports querying, transforming and analysing information through languages and environments such as SQL, Python or R.

Data mining provides methods for discovering patterns in information. Machine learning supports systems that estimate outcomes or classify cases from data. Both disciplines may form part of the work but do not define the role by themselves.

Domain knowledge helps interpret variables, detect incorrect assumptions and assess whether a relationship is meaningful. Communication allows the result to be explained to people who did not participate in the analysis and prevents an estimate from being presented as a certainty.

The particular tools depend on the sources, volume, frequency, technical environment and intended use. A project also needs documentation, version control, transformation traceability and the ability to reproduce results.

Data work governance

A model needs to be evaluated with data and metrics appropriate to its purpose. A strong technical score does not automatically demonstrate usefulness, profitability or impact. Simpler alternatives and the procedure used before the model should also be compared.

Prediction and causation are different objectives. A model may anticipate an outcome without demonstrating what causes it. Causal decisions require additional designs, assumptions and evidence.

Data may contain errors, omissions and imbalances or represent only part of the population. These conditions may affect groups, periods and situations not observed during development in different ways.

When an analysis uses personal or confidential data, its purpose, access, retention, security and deletion options need to be defined. Reusing information for a different objective is not justified merely because it is technically possible.

Deployed models require monitoring because data, processes and behaviour may change. Automation does not remove the responsibility to review errors, document decisions and control the system’s consequences.