TSV Stats
A first look at an unfamiliar file. For every column you get an inferred type, how many values are present and missing, how many are distinct, and — where the values are numeric — minimum, median, maximum, mean and standard deviation. Underneath, a short list of the things worth a look: the column that is 60% missing, the ID that is not unique, the price whose mean sits nowhere near its median.
How to use
- Paste or drop the file. The summary appears as a table you can copy.
- Read Worth a look at the bottom first. It is the short list: near-empty columns, constant columns, stray whitespace, values that broke a numeric column, and means that sit far from their median.
- Read the missing counts — they tell you which columns can be trusted.
- Compare unique counts to the row count to find candidate keys, and to spot columns that are constant.
- Turn on Top values when you want the actual category list rather than just how many there are.
- Copy or download the summary to paste into a ticket or a data review.
What counts as missing
Not just the empty string. A column of 9,999 numbers and one N/A is a numeric
column with a hole in it, and treating it as text — which is what a strict reading does —
throws away the minimum, the maximum and the mean for most real exports. So
NA, N/A, NULL, nil, none,
-, --, ?, nan, undefined and
Postgres's \N all count as missing, case-insensitively and after trimming. Add
your own in Also missing — -999 and TBD are the
usual local dialect.
The type column
Inferred from the values, not declared anywhere: integer, number,
date, boolean, category (few distinct values relative
to the rows), text, text (unique), or empty.
number* is the one to read carefully. It means at least 90% of the filled values
are numeric and the rest are not — a real numeric column with something wrong in it. The
numeric statistics are computed over the values that parse, and the offenders are quoted by
name in Worth a look. Below 90% the column is not called numeric at all.
What each number is good for
Missing is the most useful column in practice: a field that's mostly empty can't support the analysis someone is about to do with it. Unique equal to the filled count means the column is a candidate key — worth confirming before you join on it, since duplicate keys multiply rows. Unique equal to 1 means the column is constant and can probably be dropped with delete columns.
Min and max catch sentinel values that pretend to be data: a min of
-999 or a max of 9999-12-31 is almost always a placeholder rather
than a measurement.
Median next to mean is there because a mean quoted on its own is the classic way to be misled by one bad row. When the two are far apart the column is skewed or has an outlier, and the summary says so and names the extreme value. Stdev is the sample standard deviation (n−1), so a single-value column reads 0 rather than an error.
For a non-numeric column, min and max are lexicographic —
first and last alphabetically. That is genuinely useful for ISO dates and IDs, and misleading
for numbers stored as text, where "10" sorts before "9". The
summary spells this out under the table so the distinction is never silent.
Where to go next
For the distribution of a single categorical column — the real category list, plus the typos hiding in the tail — use value counts. For types and nullability expressed as DDL, the schema generator. For aggregates grouped by a column rather than over the whole file, group and aggregate.
FAQ
Are non-numeric values counted in the mean?
No — they're skipped, not treated as zero. That means the mean is over the numeric subset, which is why the count column matters when you read it.
What counts as null?
An empty cell, or one containing only whitespace. A literal N/A, -, or NULL string counts as a value — value counts will show you those, and find and replace can convert them to genuinely empty.
Is there a median or standard deviation?
Not on this page. Group and aggregate offers median (group by a constant column to get it for the whole file); standard deviation isn't offered anywhere here.
Does it handle numbers with thousands separators?
Yes — 1,234.5 is read as a number. Currency symbols are not stripped, so $1,234 counts as non-numeric.
Privacy
100% client-side. No upload. See the privacy policy.