Skip to content

Extracting Structured Data From a Pawn Shop or Consignment Receipt

10 min read · updated August 11, 2026

A pawn ticket and a consignment agreement are often printed on the same shop’s stationery, in the same layout, with the same fields in the same places. They are legally opposite transactions, and an extraction schema that does not separate them will produce a loan_amount that is sometimes a reserve price.

Two documents on one form

A pawn is a secured loan. The shop advances money, takes the item as collateral, and the customer has a defined period in which to repay principal plus charges and take the item back. Title to the item does not move unless the customer fails to redeem. The document’s load-bearing fields are the amount advanced, the finance charge, the maturity date and the redemption deadline.

A consignment is an agency arrangement. No money changes hands at intake. The shop takes possession of the item, tries to sell it, and on a sale pays the consignor the sale price less an agreed commission. Title stays with the consignor throughout. The load-bearing fields are the commission rate, any reserve or minimum price, the consignment period and the terms for return of unsold goods.

Both documents have a date, a customer, an item description, a number and a signature, which is why they look alike and why a model asked to “extract the amount” will find one on either. On a pawn ticket that amount is money the customer received. On a consignment agreement the most prominent amount is frequently an asking price or a reserve — money nobody has paid anybody. Merged into one column, these are indistinguishable and quietly wrong.

So the first field the extractor produces is a discriminator: agreement_type, one of pawn, consignment, outright_purchase or unknown. The signals are textual and reliable — the presence of a finance charge or an annual percentage rate, the words “redeem”, “pledge” or “collateral” against “consign”, “commission” or “net to seller”. Extract the discriminator first, then apply a different schema branch to each, and make unknown a routing outcome rather than a default. Guessing here corrupts a financial figure, and a corrupted financial figure is not detectable downstream.

The redemption deadline is a legal date

On a pawn ticket, the redemption deadline is the date after which the customer loses the item. It is regulated at state level in the US: statutes set a minimum loan term, and many set a grace period after maturity during which the item still cannot be sold. Those periods differ between jurisdictions and are amended, so the deadline printed on a ticket in one state is not derivable from the rule you learned in another.

The rule that follows is short: if the document prints a deadline, extract the printed deadline verbatim and use it. Compute one only as a cross-check, and when the computed date disagrees with the printed one, trust the document and flag the disagreement. The printed date is the one that governs, and your arithmetic does not know about the state grace period, the shop’s own policy of rounding to the end of a business day, or a holiday rule.

Statutory pawn holding periods, grace periods and maximum finance charges are state law and are amended. Nothing on this page is legal advice, and a deadline you compute is not a substitute for the deadline on the ticket.

Ninety days and three months are not the same date

When you do compute a cross-check, the ambiguity is not in the addition. It is in what the term means. Assume, purely as a worked example, a loan dated 14 March 2026 with a stated term of 90 days and a stated grace period of 30 days.

loan date                        2026-03-14

term stated as "90 days":
   remainder of March (15th-31st)      17 days
   April                               30 days   (47)
   May                                 31 days   (78)
   June 1-12                           12 days   (90)
   maturity                         2026-06-12
   + 30 day grace                   2026-07-12

term stated as "3 months":
   maturity                         2026-06-14
   + 30 day grace                   2026-07-14

two days apart, from the same loan date

Two days is not a rounding difference on a document where the date decides who owns the item. And the calculation above already made a choice you should make explicitly: it counted the day after the loan date as day one. Counting the loan date itself as day one moves every result one day earlier. Both conventions exist in real contracts, and the document usually does not say which it used.

Month arithmetic has its own edge that bites a few days a year: 31 January plus one month is not a date. Implementations differ between clamping to 28 February and rolling forward to 3 March, and the difference is three days on exactly the kind of deadline where three days matters. Whatever your date library does here, it made a choice on your behalf that you should know about.

Record, for every deadline, whether it was printed or derived, and if derived, which convention produced it — provenance of this kind is what the general date field validation rule asks for:

"redemption_deadline": {
  "value": "2026-07-12",
  "source": "printed",
  "computed_cross_check": "2026-07-12",
  "agrees": true,
  "convention": "day_after_loan_date"
}

Item description is a matching problem

The item description on these documents is free text written at speed by a person holding the item, and it is the field on which the whole document depends, because the item is the collateral. It reads like 14K YG chain 22in 18.4g or Yamaha acoustic gtr w/ case, scratch on back. There is no controlled vocabulary and there will not be one.

Do not attempt to normalise it into a product taxonomy at extraction time. What you can extract reliably is the structured fragments the description happens to contain — a weight with its unit, a metal purity, a serial number, a length, a stated condition note — and you should extract those alongside the verbatim string rather than instead of it. The verbatim string is what the customer signed for; the fragments are what you can search on.

The serial number is the one fragment worth special handling, because it is the field that matches this item to a police property report, an insurance claim or a prior ticket. It is also the field most damaged by OCR: a serial is unsegmented alphanumerics with no checksum and no dictionary, the exact conditions under which a vision model produces a confident wrong reading. Where a serial exists, capture it as a string, preserve case, and never strip characters you think are separators, since a hyphen may be part of the manufacturer’s format. If you hold prior tickets, string-distance the new serial against them: a serial one character away from an existing record is much more likely to be the same item than a new one.

Weights and purities deserve a unit field rather than a parsed number. A chain recorded as 18.4g and one recorded as 0.59 ozt are the same object, and troy ounces are not avoirdupois ounces — a distinction that only matters when someone totals a portfolio, at which point it matters by about 10%.

One record, two shapes

Model this as a tagged union rather than one wide table with half the columns null, which is the choice weighed in nested versus flat extraction schemas. The shared fields are the shop, the ticket number, the date, the customer reference and the item block. Everything else hangs off the discriminator:

common:      shop, ticket_number, agreement_date, item[], signature_observed

pawn:        principal_advanced, finance_charge, total_to_redeem,
             maturity_date, redemption_deadline, deadline_source

consignment: commission_rate, reserve_price, asking_price,
             consignment_period_end, unsold_disposition

The practical benefit shows up the first time somebody runs a query. “Total advanced this quarter” sums a column that only exists on pawn records and therefore cannot accidentally include a consignment reserve price. A wide table makes that mistake possible; a union makes it a type error. For how much structure to invest in when the volume is a few hundred documents rather than a few million, see schema design at low volume, and for the self-consistency checks that apply once money is involved, the auction receipt page covers the reconciliation identity in detail.