Tabular Data Embeddings for Row Similarity Search
10 min read · updated August 11, 2026
“Find me customers like this one” requires a vector per row and a distance between vectors. Building the vector is mostly mechanical; deciding how much each column contributes to the distance is the entire problem, and it is usually made by accident.
Why a row is hard to embed
Text embedding has an easy case that tabular data does not: a sentence is a sequence of tokens from one vocabulary, so a single model can map the whole thing. A table row is a heterogeneous tuple. One field is a count in the range 0–20, the next is an amount in the range 0–500,000, the third is one of eleven industry codes, the fourth is a date, the fifth is a free-text note.
Those fields are not commensurable. A difference of 1 in the count field and a difference of 1 in the amount field are wildly different events, and any distance metric that treats them as equal has made a claim about their relative importance without being asked. That claim — not the choice of cosine versus Euclidean, not the architecture — is what determines what your similarity search returns.
The concatenation baseline
The baseline that works surprisingly well: encode each column independently, concatenate, and use a standard metric. Numeric columns are scaled to a comparable range; categorical columns are one-hot encoded or replaced with a learned vector; dates become one or more numeric components such as an epoch offset plus cyclical day-of-week terms; free text goes through a text embedding model and contributes its own block of dimensions.
Every design decision in that sentence is a weighting decision in disguise. A 384-dimensional text block concatenated with six scaled numeric columns means the text contributes 384 of 390 dimensions to a Euclidean distance, and your “customer similarity” search is a note-similarity search with a rounding error attached. If the text block is meant to count for a third of the similarity, it has to be scaled to make that true — typically by unit-normalising each block separately and then multiplying by an explicit per-block weight.
A worked distance, twice
Two customer rows, three columns: monthly spend, number of seats, and plan tier.
spend seats plan row A 1,200 14 pro row B 1,650 15 pro Column statistics over the full table: spend: mean 1,400 std 900 min 0 max 500,000 seats: mean 12 std 6 min 1 max 400
Standardise each numeric column by subtracting the mean and dividing by the standard deviation, and one-hot the plan into three columns of which both rows share the same one:
standardised:
A = ((1200-1400)/900, (14-12)/6, 1,0,0) = (-0.222, 0.333, 1,0,0)
B = ((1650-1400)/900, (15-12)/6, 1,0,0) = ( 0.278, 0.500, 1,0,0)
squared differences:
spend (0.278 - -0.222)^2 = 0.500^2 = 0.2500
seats (0.500 - 0.333)^2 = 0.167^2 = 0.0279
plan 0
------
Euclidean distance = sqrt(0.2779) = 0.527
spend contributes 0.2500 / 0.2779 = 90.0% of the distanceNow min-max scale the same two columns to [0, 1] instead, using the minimum and maximum from the table:
min-max scaled:
A = (1200/500000, (14-1)/399, 1,0,0) = (0.00240, 0.03258, 1,0,0)
B = (1650/500000, (15-1)/399, 1,0,0) = (0.00330, 0.03509, 1,0,0)
squared differences:
spend (0.00090)^2 = 0.00000081
seats (0.00251)^2 = 0.00000630
----------
Euclidean distance = sqrt(0.00000711) = 0.00267
spend now contributes 0.00000081 / 0.00000711 = 11.4% of the distanceSame rows, same metric, same encoder. Spend went from 90% of the distance to 11%, and the nearest-neighbour lists these two scalings produce will disagree substantially. The cause is that min-max scaling divides by the range, and one extreme spend value of 500,000 compresses every ordinary customer into the bottom hundredth of the axis. This is the same distortion described in numeric feature scaling, with a directly visible consequence.
The lesson is not that one scaler is correct. It is that the scaler is the weighting, and if you have not decided what the relative weights should be, the scaler decided for you.
Gower distance, and what it fixes
Before learned encoders existed, the standard answer to mixed-type rows was the general similarity coefficient John C. Gower published in Biometrics in 1971. It defines a per-column similarity and averages them with explicit weights: for a numeric column, one minus the absolute difference divided by that column’s range; for a categorical column, one if the values match and zero otherwise; for a binary column, an asymmetric variant that does not count a shared absence as evidence of similarity.
Two properties make it worth knowing even now. The per-column contribution is bounded in [0, 1] by construction, so no column can dominate by virtue of its units — the failure the worked example above demonstrates. And the weights are a named parameter rather than an emergent property of your preprocessing, so the weighting decision is made explicitly. The cost is that it does not vectorise into a single dot product, so it does not slot into an approximate nearest-neighbour index the way a fixed-length vector does.
Learned row encoders
A learned encoder replaces the hand-built vector with one trained so that rows which are similar for your purpose land near each other. TabTransformer, introduced by Xin Huang and co-authors at Amazon in “TabTransformer: Tabular Data Modeling Using Contextual Embeddings”, is the canonical form: categorical columns become embeddings, transformer layers let those embeddings condition on each other, and the contextual output is concatenated with the numeric columns.
What the contextual part buys is column interaction. In a concatenated encoding, the vector for industry=healthcare is the same regardless of the rest of the row. In a contextual one it can differ for a two-seat healthcare account and a two-thousand-seat one, because the embedding attends to the other fields. Whether that is worth the training cost depends on whether such interactions exist in your data, and the honest answer is that on many tables they do not and the concatenation baseline with careful weighting is competitive.
The training signal is the other decision. Self-supervised objectives — mask a field and reconstruct it — give a general-purpose space. A supervised or contrastive objective on labelled pairs gives a space aligned to one definition of similarity, which is usually what a business actually wants and is why an embedding trained for churn prediction returns different neighbours from one trained for fraud.
Serving it as a search index
Once rows are fixed-length vectors, the retrieval problem is the ordinary one: an approximate nearest-neighbour index, a metric chosen to match how the vectors were built, and a re-ranking pass if precision at the top matters. Cosine ignores magnitude and Euclidean does not, and for row embeddings that difference is substantive rather than cosmetic — magnitude often encodes size of account, which you may or may not want folded into similarity. The trade-offs are the same ones covered in vector similarity metrics.
Two operational points that bite specifically here. Scaler parameters are model state: the means and standard deviations used to build the index must be persisted and reused at query time, or a query vector computed with recomputed statistics sits in a subtly different space from the index. And drift is a re-index event — when the distribution of spend shifts, the standardisation that defined the space is stale, and the neighbours degrade silently rather than failing. Recompute the statistics on a schedule and rebuild, or accept that the index means something different each quarter.