Jump to content

Data

From IdeaWazaWiki

Data are recorded observations, measurements, symbols, descriptions, values, or other representations that can be used for learning, analysis, communication, decision-making, and research. Data can be numerical, textual, visual, auditory, geographic, biological, behavioral, or otherwise structured or unstructured.

Data are fundamental to modern science, technology, business, government, education, and everyday life. Humans collect data in order to understand the world, test ideas, identify patterns, make predictions, solve problems in living, and evaluate whether particular actions or interventions are producing desired results.

Data by themselves are not necessarily knowledge. A large collection of numbers, measurements, or documents may be difficult to interpret without context. Data become more useful when they are organized, analyzed, compared, and connected with meaningful questions.

The study of data involves fields such as statistics, computer science, data science, mathematics, information science, research methodology, economics, psychology, sociology, medicine, and artificial intelligence. Data is like bulk information.

What counts as data?

Almost anything that can be recorded in a reasonably systematic way can potentially become data.

Examples include:

  • Temperature measurements.
  • Survey responses.
  • Census records.
  • Financial transactions.
  • Photographs.
  • Audio recordings.
  • Satellite images.
  • Scientific measurements.
  • Website traffic.
  • Medical test results.
  • Geographic coordinates.
  • Interviews.
  • Historical documents.
  • Social media posts.
  • Experimental observations.
  • Sensor readings.
  • Computer logs.
  • Voting records.
  • Prices of goods and services.
  • School attendance.
  • Biological sequences.

Whether something functions as useful data depends partly on the question being asked.

For example, a list of daily temperatures could be used to study local weather patterns. The same information could later be combined with decades of measurements to investigate long-term climate trends.

Quantitative and qualitative data

One common distinction is between quantitative data and qualitative data.

Quantitative data are represented numerically. Examples include age, income, temperature, distance, population size, blood pressure, test scores, and the number of people visiting a website.

Quantitative data can often be examined through statistics.

Qualitative data describe characteristics, experiences, meanings, perceptions, or events that are not necessarily best represented by numbers.

Examples include:

  • Interview transcripts.
  • Written observations.
  • Diaries.
  • Open-ended survey responses.
  • Historical documents.
  • Photographs.
  • Video.
  • Descriptions of social interactions.

Qualitative research can examine themes, meanings, patterns, experiences, and relationships that might be difficult to represent only through numerical measurement.

Researchers can also combine qualitative and quantitative methods through mixed methods research.

Structured and unstructured data

Structured data are organized according to a consistent format.

A spreadsheet containing a row for each person and columns for age, income, education, and location is an example of structured data.

Databases often contain highly structured data so that information can be searched, filtered, compared, and analyzed efficiently.

Unstructured data do not necessarily follow a rigid predefined format.

Examples include:

  • Books.
  • Emails.
  • Videos.
  • Photographs.
  • Audio recordings.
  • Web pages.
  • Research articles.

Modern artificial intelligence systems have increased the ability to analyze large amounts of unstructured data, especially text, images, audio, and video.

There are also forms of semi-structured data, such as JSON, XML, and some types of web content, which contain organizational structure without being arranged like a conventional table.

Primary and secondary data

Primary data are collected directly for a particular research project or purpose.

Examples include conducting a survey, running an experiment, interviewing participants, or collecting environmental measurements.

Secondary data already exist and are reused for another purpose.

Examples include:

  • Census data.
  • Government statistics.
  • Existing scientific databases.
  • Historical archives.
  • Previously published research datasets.
  • Public financial records.

Secondary data can make research less expensive and allow researchers to study questions that would otherwise require enormous amounts of time or money.

Researchers must still consider why the original data were collected, how they were measured, and what limitations may exist.

The data lifecycle

Data can move through several stages.

A simplified data lifecycle might include:

  1. Identifying a question.
  2. Determining what data are needed.
  3. Collecting the data.
  4. Recording and storing the data.
  5. Cleaning the data.
  6. Organizing the data.
  7. Analyzing the data.
  8. Interpreting the results.
  9. Communicating findings.
  10. Preserving, sharing, archiving, or deleting the data.

Each stage can affect the reliability of the final conclusions.

For example, sophisticated statistical analysis cannot fully repair data that were collected poorly in the first place.

Data quality

Not all data are equally useful.

Important characteristics of data quality can include:

  • Accuracy - whether the data reasonably represent what actually occurred.
  • Completeness - whether important observations are missing.
  • Consistency - whether information is recorded in compatible ways.
  • Timeliness - whether the data are sufficiently current for the intended purpose.
  • Validity - whether a measurement actually captures what it is intended to measure.
  • Reliability - whether similar measurements would produce reasonably consistent results.
  • Representativeness - whether the data adequately reflect the population being studied.

Poor-quality data can result in misleading conclusions even when the analysis itself is mathematically correct.

This is sometimes summarized by the computing expression garbage in, garbage out. If the information entering a system is inaccurate or inappropriate, the output may also be inaccurate.

Data collection and bias

Data are not necessarily neutral simply because they contain numbers.

Researchers decide what to measure, how to measure it, whom to include, how questions are worded, which records are available, and how missing information is handled.

These choices can introduce bias.

For example, an online survey may exclude people without reliable internet access. A study of customers who purchased a product may reveal little about people who considered buying it but decided not to.

Historical datasets may also reflect older institutional practices, discrimination, incomplete records, or changing definitions.

Understanding data therefore requires understanding how the data were produced.

Data analysis

Data analysis involves examining information in order to answer questions or identify patterns.

Basic analysis might include:

  • Calculating averages.
  • Comparing groups.
  • Measuring percentages.
  • Creating graphs.
  • Identifying trends.
  • Examining correlations.
  • Looking for unusual observations.

More advanced methods can include regression, machine learning, Bayesian analysis, causal inference, network analysis, time-series analysis, and natural language processing.

An important distinction exists between correlation and causation.

Two variables can change together without one causing the other. Researchers need additional evidence and carefully designed studies before drawing strong causal conclusions.

Data visualization

Data visualization uses visual representations to make information easier to understand.

Examples include:

  • Bar charts.
  • Line graphs.
  • Scatter plots.
  • Maps.
  • Histograms.
  • Network diagrams.
  • Interactive dashboards.

Good visualization can reveal patterns that are difficult to notice in tables of numbers.

Poor visualization can also mislead. Manipulated axes, inappropriate scales, selective time periods, and excessive visual complexity can create inaccurate impressions.

Learning how to interpret graphs critically is therefore an important part of data literacy.

Open data and research

Open data are data made available for people to access, reuse, analyze, and redistribute, subject to appropriate legal and ethical restrictions.

Open scientific data can help researchers:

  • Reproduce studies.
  • Detect errors.
  • Develop new research questions.
  • Combine information from multiple sources.
  • Avoid unnecessary duplication of data collection.
  • Build educational resources.

Government open-data projects can also allow citizens, journalists, businesses, researchers, and community organizations to study public services and social conditions.

Open data does not mean that every dataset should be public. Personal information, confidential records, security-sensitive information, and other restricted data may require protection.

Privacy and ethics

The collection of data can create significant ethical questions.

Personal data may include information about someone's identity, location, finances, communications, health, employment, relationships, activities, or preferences.

Important questions include:

  • Did people consent to the data collection?
  • Do people understand how their data will be used?
  • Who owns or controls the data?
  • Who can access it?
  • How long should it be stored?
  • Can supposedly anonymous data be connected back to individuals?
  • Can the data be used for purposes different from those originally stated?

Data security is also important. Large databases can create significant risks if unauthorized parties gain access to them.

The benefits of collecting information should therefore be balanced against privacy, autonomy, security, and possible misuse.

Data and artificial intelligence

Modern artificial intelligence depends heavily on data.

Machine-learning systems identify statistical patterns in training data and use those patterns to generate predictions, classifications, recommendations, text, images, or other outputs.

The quality of AI can therefore be affected by the quality and composition of its training data.

Problems can arise when training data are inaccurate, incomplete, outdated, biased, duplicated, or poorly matched to the task.

AI also creates new possibilities for analyzing very large datasets. Researchers can use AI to summarize documents, identify patterns in images, classify information, generate hypotheses, and assist with data analysis.

AI-generated conclusions should still be verified. A system can generate plausible explanations that are unsupported by the underlying data.

Data as a tool for solving problems

Data can help people identify and solve practical problems.

A basic approach can be:

  1. Define the problem.
  2. Determine what information would help explain the problem.
  3. Collect or locate relevant data.
  4. Evaluate the quality of the data.
  5. Analyze the information.
  6. Develop possible solutions.
  7. Implement an intervention.
  8. Collect additional data to determine whether the intervention worked.

This approach can be applied to social problems, business problems, scientific problems, health problems, technological problems, and many other forms of problem solving.

Data do not automatically determine what people should do. Values, priorities, costs, rights, uncertainty, and practical limitations can also affect decisions.

  • What is the difference between data, information, knowledge, and wisdom?
  • Can data ever be completely neutral?
  • What makes a dataset reliable?
  • What are examples of situations where more data do not necessarily produce better decisions?
  • What is the difference between quantitative and qualitative data?
  • How can researchers identify bias in a dataset?
  • Why does correlation not necessarily demonstrate causation?
  • What information should governments make available as open data?
  • When should personal privacy outweigh the potential value of data for research?
  • Ask an AI system to identify ten possible sources of bias in a dataset. Evaluate which ones actually apply to a real dataset.
  • Ask an AI system to propose a research question and identify what data would be needed to answer it.
  • Select an open government dataset and develop three research questions that could be investigated using it.
  • Compare two visualizations of the same information and determine which one communicates the data more effectively.
  • How might access to better data help humanity solve problems in living?
  • What types of data could be useful for measuring progress toward solving a major social problem?
  • How should AI-generated data or synthetic data be distinguished from observations collected from the real world?

Readings

Wikipedia

See also