Extracting Sender, Recipients and Thread Structure From an Email Export
11 min read · updated August 11, 2026
An email export is one of the few unstructured documents that contains its own structure. Every reply carries a pointer to its parent. If you are reconstructing a thread by matching subject lines or by reading the nested quote blocks top to bottom, you are ignoring the graph the file is handing you and solving a much harder problem badly.
Why subject matching is the wrong algorithm
Subject-line threading fails in both directions and both failures are common. It joins messages that are unrelated, because “Re: Update” is not a distinctive string and two different conversations in a large export will collide on it. It splits messages that belong together, because a participant renames the subject mid-thread, or replies from a client that localises the reply prefix — Re:, RE:, AW: in German, SV: in Swedish and Danish, VS: in Finnish, Antw: in Dutch. A stripper that only knows Re: and Fwd: puts a German reply in its own thread.
Subject matching also cannot tell you shape. A thread is a tree, not a list: three people reply to the same message and each of those branches continues separately. Sorting by subject and then by date renders that tree as a flat sequence in which a reply appears after messages it never saw. Anything you then extract from “context” — who agreed to what, who was asked — is derived from an order that did not happen.
The two headers that carry the graph
RFC 5322, the Internet Message Format standard published by the IETF, defines three identification fields in section 3.6.4: Message-ID, In-Reply-To and References. Every message should carry a globally unique Message-ID in angle brackets. A reply should carry In-Reply-To containing the parent’s message id, and References containing the parent’s References list followed by the parent’s own id — so References accumulates the ancestry back to the root, in order.
Message-ID: <[email protected]> In-Reply-To: <[email protected]> References: <[email protected]> <[email protected]> Subject: Re: Q3 renewal terms Date: Tue, 14 Jul 2026 09:12:04 +0100 From: "A. Okafor" <[email protected]> To: "M. Reyes" <[email protected]> Cc: [email protected]
Two things about that block matter for parsing. References is a folded header: RFC 5322 allows a long header to continue on the next line as long as the continuation begins with whitespace, so you must unfold before you tokenise, or you will read only the first id. And the last element of References is the parent, which means References alone is enough to build the tree even when In-Reply-To is absent. Use In-Reply-To when present, fall back to the last References entry, and use the full References list to attach an orphan to the nearest ancestor you do have. The IETF publishes RFC 5322.
Building the tree
- Parse each message to a record of id, parent candidates, ancestry list, date, from, to, cc and body parts. Normalise message ids by stripping the angle brackets and trimming whitespace; do not lowercase them, because the local part of a message id is case-sensitive and some generators use mixed case.
- Index every message by id. For each message, take
In-Reply-Toif it names an id you have; otherwise walkReferencesfrom the end backwards and take the first id you have. If none resolves, the message is a root for now. - Attach children to parents and sort each sibling group by the
Dateheader. Order within a sibling group is a timestamp question; order between generations is the graph, and the graph wins where they disagree. - Detect cycles before you recurse. A malformed export can contain a message that references itself or a pair that reference each other, and a naive depth-first walk over that hangs. Cap depth and mark the offender rather than failing the whole export.
Do not order the whole thread by Date. That header is written by the sending client from the sender’s own clock, so a machine with a wrong timezone or a skewed clock produces a reply dated before the message it answers. RFC 5322 also permits a zone offset of -0000, which specifically means the local zone is unknown rather than meaning UTC. The Received headers, prepended by each hop with the most recent first, are a better source for server-side arrival time if you need one — but the parent graph is better than either, because it is what actually happened.
Separating new text from quoted history
Once the tree is right, the body of each message still contains most of the thread again, quoted. If you feed that to anything downstream you will extract the same commitment five times, once per reply that quoted it.
There is no header that marks quoted text, so this is heuristic, but the heuristics are good and they stack. Lines prefixed with > are quoted, and the number of > characters is the quote depth — RFC 3676, which defines the format=flowed parameter for text/plain, formalises this prefix as the quote marker. An attribution line immediately above a quoted block — “On Mon, 13 Jul 2026 at 17:40, M. Reyes wrote:” — introduces it, though the exact wording is client-specific and localised, so match it loosely and only in the position where you expect it. A line consisting of exactly two hyphens and a space is the conventional signature separator, and everything after it in a plain-text part is a signature.
HTML mail is different and often easier. Several clients wrap the quoted portion in a container with a recognisable class or in a blockquote, and Outlook historically emits a horizontal rule followed by a block of From/Sent/To/Subject lines. Parse the HTML, drop the quote container, and take the remainder — that is more reliable than converting to text first and then regexing, because the conversion is what destroys the marker.
When the headers are missing or lying
- The export is a rendering, not a mailbox. A PDF or a printed thread has no headers at all. You are then genuinely doing layout work: the visual nesting and the attribution lines are all you have, and you should record that the ordering is inferred rather than read, because its reliability is completely different.
- Forwarded digests. A single message containing five messages has one
Message-IDand five attribution blocks. Split on the attribution lines and treat the results as reconstructed messages with a flag, not as first-class ones. - Missing References on Outlook-originated replies. Microsoft clients maintain their own conversation index in a
Thread-Indexheader, a base64 structure whose layout Microsoft documents in the Exchange protocol specifications, and historically some paths populated it more reliably thanReferences. If you have an export where the standard headers are thin,Thread-Indexis a genuine fallback — consult Microsoft’s own protocol documentation for the byte layout rather than any third-party description of it. - Mailing lists rewrite things. Some list software adds its own headers and a subject tag; a few historically rewrote
References. Threads that pass through a list are the ones worth spot-checking by hand. - Duplicate message ids. The same message appears in two mailboxes in the same export. Deduplicate on message id before building the tree, not after, and keep the folder path of each copy — it is often the only evidence of who received what.
The run
- Read the export with a real MIME parser rather than a regex over the raw bytes. Header folding, RFC 2047 encoded-words in display names, and transfer encodings all have to be handled correctly before any of the above works, and every language has a library that does it.
- Build the tree from the graph as described. Persist the resulting parent id on each message, with a field recording which header it came from, so that a later argument about ordering has an answer.
- Strip quoted history and signatures, keeping the original body. Store the stripped body as its own column; everything downstream reads that one.
- Only now involve a model, and only on the stripped new text of one message at a time, with the thread structure supplied as context. The model is for the language in the message, not for working out what replies to what — you already know that exactly.
With a correct tree and clean per-message text, the interesting extractions become tractable: who committed to what and what was attached and whether it is actually in the export.