A dataset is a collection of related data points organized for analysis. It might be a single table, a set of files, or a database extract. A dataset has a purpose. The Iris dataset contains measurements of flower petals and sepals, used to teach classification. The MNIST dataset contains handwritten digits, used to benchmark image recognition. A company's sales dataset contains transactions, used to forecast revenue. The dataset defines the boundaries of the analysis.
Datasets vary widely in size, quality, and structure. Some are small enough to open in a spreadsheet. Others span petabytes across distributed storage. Some are carefully curated with documented provenance. Others are scraped from the web with unknown reliability. The quality of a dataset determines the quality of the analysis built on it. A dataset with missing values, biased sampling, or inconsistent labeling produces unreliable results. Before using a dataset, examine it. What is the source? How was it collected? What is the time period? What is missing? What are the known limitations? These questions are not optional. They determine what the data can support. A dataset that works for one analysis may be unsuitable for another. The data does not know its own limitations. The analyst must.
Dataset characteristics
- Size — number of records and features
- Structure — tabular, hierarchical, or unstructured
- Quality — accuracy, completeness, consistency
- Provenance — where it came from and how
- License — terms of use and redistribution
A dataset is a snapshot of the world. The snapshot has a frame. Understanding the frame is part of understanding the data.
Comments
No comments yet. Be the first to share a thought.
Leave a comment