Security scan is checking this page.

Chapter 1

What Are Data and Data Science?

Read chapter 1 in the book

Summary

Data science is a field of study that investigates how to collect, manage, and analyze data of all types in order to retrieve meaningful information.

Key terms

data science
a field of study that investigates how to collect, manage, and analyze data in order to retrieve meaningful information from some seemingly arbitrary data
data
anything that we can analyze to compile some high-level insights
information
some high-level insights that are compiled from data
dataset
a collection of related and organized information or data points grouped together for reference or analysis
data collection
the systematic process of gathering information on variables of interest
data science cycle
a process used when investigating data
qualitative data
non-numerical data that generally describe subjective attributes or characteristics and are analyzed using methods such as thematic analysis or content analysis
quantitative data
data that can be measured by specific quantities and amounts and are often analyzed using statistical methods

Chapter 2

Collecting and Preparing Data

Read chapter 2 in the book

Summary

The gathered data is crucial for making sound interpretations and gaining meaningful insights.

Key terms

gathered data
crucial for making sound interpretations and gaining meaningful insights
BCNF (Boyce-Codd Normal Form)
a normal form that is similar to 3NF but stricter in its requirements for data independence; it ensures that all attributes in a table are functionally dependent on the primary…
cloud storage
type of online storage that allows users and organizations to store, access, and manage data over the internet with data. This data is stored on remote servers managed by a cloud…
data validation
procedure to ensure that the data is reliable and trustworthy for use in the data science project, done by performing checks and tests to identify and correct any errors in the…
log transformation
data transformation technique that requires taking the logarithm of the data values and is often used when the data is highly skewed
normal form
a guideline or set of rules used in database design to ensure that a database is well-structured, organized, and free from certain types of data irregularities, such as…
object storage
method of storage where data are stored as objects that consist of both data and metadata and often used for storing large volumes of unstructured data, such as images and videos
observational data
data that is collected by observing and recording natural events or behaviors in their normal setting without any manipulation or experimentation; typically collected through…

Chapter 3

Descriptive Statistics: Statistical Measurements and Probability Distributions

Read chapter 3 in the book

Summary

These measures can help indicate where the bulk of the data is concentrated and are often called the data’s central tendency . The mean , or average, (sometimes referred to as the arithmetic mean ) is the most commonly used measure of the center of a dataset.

Key terms

probability
a numerical measure that assesses the likelihood of occurrence of an event
probability distribution
a mathematical function that assigns probabilities to various outcomes
descriptive statistics
the organization, summarization, and graphical display of data
trimmed mean
a calculation for the average or mean of a dataset where some percentage of data values are removed from the lower and upper end of the dataset; typically used to mitigate the…
outliers
data values that are significantly different from the other data values in a dataset
bulk of the data
concentrated and are often called the data’s central tendency
central tendency
the statistical measure that represents the center of a dataset
arithmetic mean )
the most commonly used measure of the center of a dataset

Chapter 4

Inferential Statistics and Regression Analysis

Read chapter 4 in the book

Summary

Data scientists interested in inferring the value of a population truth or parameter such as a population mean or a population proportion turn to inferential statistics. Inferential statistics provides methods for generating predictive forecasting models , and this allows data scientists to generate predictions and trends to assist in effective and accurate decision-making. A data scientist is often interested in making generalizations about a population based on the characteristics derived from a sample; inferential statistics allows a data scientist to draw conclusions about a… In addition, inferential statistics is used by data scientists to assess model performance and compare different algorithms in machine learning application.

Key terms

inferential statistics
statistical methods that allow researchers to infer or generalize observations from samples to the larger population from which they were selected
regression analysis
a statistical technique used to model the relationship between a dependent variable and one or more independent variables
bootstrapping
a method to construct a confidence interval that is based on repeated sampling and does not rely on any assumptions regarding the underlying distribution
confidence interval
an interval where sample data is used to provide an estimate for a population parameter
prediction
a forecast for the dependent variable based on a specific value of the independent variable generated using the linear model
proportion
a measure that expresses the relationship between a part and the whole; a proportion represents the fraction or percentage of a dataset that exhibits a particular characteristic…
data scientist
often interested in making generalizations about a population based on the characteristics derived from a sample; inferential statistics allows a data scientist to draw c
analysis of variance (ANOVA)
statistical method to compare three or more means and determine if the means are all statistically the same or if at least one mean is different from the others

Chapter 5

Time Series and Forecasting

Read chapter 5 in the book

Summary

Time series analysis allows for the examination of data points collected or recorded at specific time intervals, enabling the identification of trends, patterns, and seasonal variations crucial for making informed… Some examples of the areas in which time series analysis is used to solve real-world problems include the following Making predictions about future (unknown) data based on current and historical data may be the most important use of time series analysis.

Key terms

time series
data that has a time component, or an ordered sequence of data points
forecasting
making predictions about future, unknown values of a time series
time series analysis
the examination of data points collected at specific time intervals, enabling the identification of trends, patterns, and seasonal variations crucial for making informed…
trend
long-term direction of time series data in the absence of any other variation
exponential moving average (EMA)
type of weighted moving average in which the most recent values are given larger weights and the weights decay exponentially the further back in time the values are in the series
order (or degree)
the number of terms or lags used in a model to describe the time series; for a sequence that is modeled by a polynomial formula, the value of the exponent on the term of highest…
stationary
characterizing a time series in which the variance is relatively constant over time, an overall upward or downward trend cannot be found, and no seasonal patterns exist

Chapter 6

Decision-Making Using Machine Learning Basics

Read chapter 6 in the book

Summary

A machine learning (ML) model is a mathematical and computational model that attempts to find a relationship between input variables and output (response) variables of a dataset.

Key terms

unsupervised learning
machine learning methods that do not require data to be labeled in order to learn; often, unsupervised learning is a first step in discovering meaningful clusters that will be…
machine learning (ML) model
any algorithm that trains on data to determine or adjust parameters of a model for use in classification, clustering, decision making, prediction, or pattern recognition
overfitting
modeling using a method that yields high variance; the model captures too much of the noise and so may perform well on training data but very poorly on testing data
underfitting
modeling using a method that yields high bias; the model does not capture important features of the data
typical mathematical model
that the user does not have full control over all the parameters used by the model
supervised learning
machine learning methods that train on labeled data
weak learners
individual models that are trained on parts of the dataset and then combined in a bootstrap aggregating method such as random forest

Chapter 7

Deep Learning and AI Basics

Read chapter 7 in the book

Summary

The goal of neural networks is for a computer algorithm to be able to classify these digits as well as a human

Key terms

deep learning
training and implementation of neural networks with many layers to learn hierarchical (structured) representations of data
bias
value b that is added to the weighted signal, making the neuron more likely (or less likely, if b is negative) to activate on any given input
weight
value w that is multiplied to the incoming signal, essentially determining the strength of the connection
neural network
structure made up of neurons that takes in input and produces output that classifies the input information
fully connected layers
layers of a neural network in which every neuron in one layer is connected to every neuron in the next layer
step function
function that returns 0 when input is below a threshold and returns 1 when input is above the threshold
sparse categorical cross entropy
generalization of binary cross entropy, useful when the target labels are integers

Chapter 8

Ethics Throughout the Data Science Cycle

Read chapter 8 in the book

Summary

Data collection is an essential practice in many fields, and those performing the data collection are responsible for adhering to the highest ethical standards so that the best interest of every any party affected by… First and foremost, it is important to ensure that data is used appropriately within the boundaries of the collected data’s purpose and that individual autonomy is respected—in other words, that individuals maintain…

Key terms

autonomy
in data science, the ideal that individuals maintain control over the decisions regarding the collection and use of their data
informed consent
the process of obtaining permission from a research subject indicating that they understand the scope of collecting data
data sharing
processes of allowing access to or transferring data from one entity (individual, organization, or system) to another
data collection
responsible for adhering to the highest ethical standards so that the best interest of every any party affected by…
data security
steps taken to keep data secure from unauthorized access or manipulation
cookies
small data files from websites that are deposited on users’ hard disk to keep track of browsing and search history and to collect information about potential interests to tailor…
data governance protocols
set of rules, policies, and procedures that enable precise control over data access while ensuring that it is safeguarded
data privacy
the assurance that individual data is collected, processed, and stored securely with respect for individuals' rights and preferences

Chapter 9

Visualizing Data

Read chapter 9 in the book

Summary

Univariate data includes observations or measurements on a single characteristic or attribute, and the data need to be converted in a structured format that is suitable for analysis and visualization. Bivariate data refers to data collected on two (possibly related) variables such as “years of experience” and “salary.” The data visualization method chosen will depend on whether the data is univariate or bivariate…

Key terms

data
univariate or bivariate…
bivariate data
paired data in which each value of one variable is paired with a value of a second variable
Pareto chart
a type of bar chart where the bars are arranged in order of decreasing height
data visualization
the use of graphical displays, such as bar charts, histograms, and scatterplots, to help interpret patterns and trends in a dataset
histogram
a graphical display of continuous data showing class intervals on the horizontal axis and frequency or relative frequency on the vertical axis
data scientist
interested in creating various visualizations, graphs, and charts to facilitate communication and assist in detecting trends and patterns
univariate data
observations recorded for a single characteristic or attribute

Chapter 10

Reporting Results

Read chapter 10 in the book

Summary

Such comprehensive reporting not only builds trust with the project team, but also contributes to the advancement of knowledge by enabling peer review and constructive feedback. Report writers often use a structured format , in which the report is divided into various standard parts for ease of understanding and consistency

Key terms

report
divided into various standard parts for ease of understanding and consistency
audience
the person or group that will be reading the report
alt text
brief descriptions of images that accompany the image or may be embedded within the image data, meant to aid readers who are not able to view the images and graphics due to…
assumption
statement that is thought to be true without being verified or proven; foundational hypotheses or beliefs about the structure, relationships, or distribution of data that guide…
k-fold cross-validation
cross-validation strategy that works by dividing the dataset into k equally sized subsets or folds, where the model is trained on k - 1 of the folds and tested on the remaining…
layered approach
writing strategy in which a report is organized into sections or appendices that provide different levels of detail and complexity for different audiences
leave-one-out cross-validation (LOOCV)
special case of k-fold cross-validation where k is set to the number of data points in the dataset so that the model is trained on all data points except one, which is used as…
constraint
limitation or restriction that is imposed on a project or its solution

Summaries and key terms on this page are taken from that chapter’s material already kept for this desk. They follow the OpenStax book. Margins is not affiliated with OpenStax. Resources, policy, and site safety