The seven datasheet sections as seven boxes, producing a document and a JSON file you can check in beside the data.
Sections answered
3 of 7
Every unanswered section is printed into the document with its questions, so the gap is visible to whoever inherits the dataset instead of being invisible to everyone.
Sections answered
3
Sections left blank
4
Questions in the framework
19
Words
319
Sections still unanswered — they appear as holes in the document:
Covering
Licence
Maintainer contact
Where it lives
Preprocessing, cleaning and labelling
Uses
Distribution
Maintenance
What this assumes: the seven groups and their questions come from the datasheets-for-datasets framework; the answers are entirely yours, and nothing is inferred from one answer to fill another. The text this page opens with is an example about an invented dataset, there to show the shape of a real answer — replace it. A blank section is printed as a blank with its questions attached rather than quietly dropped, because a datasheet whose gaps are invisible is a datasheet that gets waved through. The row count and the period are text, not measurements — this page cannot see your data. Everything on this page runs in your browser. Nothing you paste is uploaded, logged or sent anywhere.
Every dataset has a person who knows what is wrong with it, and that person leaves. The questions above exist because the knowledge that disappears with them is specific and predictable: which period the scrape actually covers, why the January rows are duplicated, which label was applied by one contractor with a different instruction sheet, and the fact that the “random sample” was the first 50,000 rows by id.
The section that pays for itself
Distribution. A model trained on data you were not licensed to train on is not a data-quality problem, it is a problem you cannot fix by retraining, because the question arrives months later and the answer has to be reconstructed from memory. Writing down where each part came from and under what terms takes ten minutes now and is unbuyable later.
Answer the ones you cannot answer
“We do not know how these were labelled” is a complete and useful answer. It tells the next person not to trust the labels more than the data supports, which is the entire point of the document, and it is far better than a blank that reads as “nothing to see here”. The completeness figure above counts answers, not reassurances — write the awkward ones down and the number goes up for the right reason.