Data quality measures how fit data is for its intended use. Accurate, complete, consistent, timely, and valid. Bad data has many flavors. A missing value. A duplicate record. A date in the wrong format. An address that does not exist. Each defect has a cost. A marketing campaign sends mail to dead addresses. A financial report overstates revenue because of duplicates. A machine learning model makes biased predictions because the training data was incomplete.
Data quality is not a one-time fix. It is a practice. Define quality dimensions for each dataset. Measure them regularly. Set thresholds for acceptable quality. Investigate when the numbers drop. The root causes are usually process problems, not technology problems. Data entry without validation creates errors. Systems that do not share a common customer ID create duplicates. Processes that allow incomplete records create gaps. Fixing the process prevents the errors. Fixing the data without fixing the process means the errors come back. Data quality tools help with profiling, cleansing, and monitoring. They do not replace ownership. Someone must be accountable for each dataset. Without that accountability, quality drifts. It always drifts.
Data quality dimensions
- Accuracy — data reflects reality
- Completeness — no missing values
- Consistency — same data across systems
- Timeliness — data is current
- Validity — data conforms to rules
- Uniqueness — no duplicates
Data quality is everyone's job. When it is nobody's job, it fails.
Comments
No comments yet. Be the first to share a thought.
Leave a comment