Semi-structured data has some organization but does not fit neatly into tables. JSON, XML, YAML, and email headers all fall into this category. The data has tags or keys that describe its structure, but the structure can vary from record to record. One JSON object might have a field that another lacks. One XML document might nest elements differently than the next. The schema is flexible, sometimes called schema-on-read. You apply structure when you query the data, not when you store it.
Semi-structured data is common in modern applications. APIs return JSON. Log files use key-value pairs. Configuration files use YAML. NoSQL document databases store JSON natively. The flexibility is useful when the data model evolves or varies. It is also a challenge for analysis. Querying semi-structured data requires different tools and techniques than querying tables. You need to navigate nested structures, handle missing fields, and flatten arrays. SQL engines have adapted. PostgreSQL, MySQL, and Snowflake all support JSON functions. You can query inside JSON documents without extracting them first. That blurring of structured and semi-structured data is making the distinction less relevant. The important question is not whether data is semi-structured. It is whether you can query it effectively.
Semi-structured formats
- JSON — key-value pairs, arrays, nested objects
- XML — tagged elements, hierarchical
- YAML — human-readable configuration
- CSV — delimited but no type information
- Log files — timestamped events with variable fields
Semi-structured data is the middle ground. It has enough structure to be useful and enough flexibility to be messy.
Comments (2)
Leave a comment