3D Mesh Reconstruction From Images
10 min read · updated August 11, 2026
A mesh is a surface: vertices, edges, faces, and an inside. A point cloud is none of those. Every step from one to the other is an inference, and the useful skill is knowing which step invented the geometry you are looking at.
Six stages, and what each one assumes
- Sparse reconstruction recovers camera poses and a few thousand triangulated points. Assumes distinctive, repeatable texture. Covered in full on the structure-from-motion page.
- Multi-view stereo densifies: for each pixel in each image, search along the epipolar line in neighbouring images for the best photometric match. Assumes the surface is Lambertian — that it looks the same from every angle.
- Fusion and filtering merges the per-view depth maps into one cloud and discards points that only one view supports. Assumes redundancy: a point seen by three cameras is trusted, a point seen by one is not.
- Normal estimation assigns each point an outward surface direction. Assumes local planarity within the neighbourhood radius.
- Surface reconstruction fits a continuous surface and extracts a triangle mesh from it. Assumes the sampling is dense enough relative to the feature size it is asked to resolve.
- Simplification and texturing reduces the triangle count and projects images back onto the surface. Assumes the poses from stage 1 are accurate, because a 2-pixel pose error becomes a visible seam.
From sparse points to a dense cloud
Multi-view stereo turns a few thousand points into tens of millions. The core operation is a photoconsistency search: take a small patch around a pixel in a reference image, hypothesise a depth, project the patch into a neighbouring image at that depth, and score the agreement — normalised cross-correlation is the standard measure because it is invariant to brightness and contrast changes between cameras. Sweep the depth and keep the best score.
The assumption that breaks first is Lambertian reflectance. A polished floor, a car body or a window looks different from every camera, so the photoconsistency score peaks at a depth that has nothing to do with the surface. What appears in the cloud is a plausible-looking sheet of points floating behind the glass, at the depth of whatever was reflected. It is not noise and no statistical outlier filter removes it, because locally it is a perfectly coherent surface.
The second assumption is texture. A blank wall gives an identical score at every depth, so the estimate is arbitrary. Good implementations detect the flat score profile and return no depth rather than a random one, which is why dense clouds have holes exactly where the surfaces are simplest.
Normals, and the orientation problem
A normal is estimated by taking the k nearest neighbours of a point, computing the covariance of their positions, and taking the eigenvector with the smallest eigenvalue: the direction in which the neighbourhood varies least is the direction perpendicular to the local plane. The ratio of that smallest eigenvalue to the sum of all three is a useful by-product — it measures how planar the neighbourhood actually is, and a high value means the normal you just computed is not trustworthy.
The eigenvector gives you a line, not a direction: it is equally valid negated. Deciding which way is “out” is a separate, global problem, and it is the one that most often ruins a reconstruction. Two solutions are standard. If the cloud came from images or a scanner, you know where the sensor was for each point, and the normal simply has to face the sensor — this is reliable and should be used whenever the viewpoint survived into your file. If it did not, orientation is propagated across a neighbourhood graph by minimum spanning tree, flipping normals to agree with already-oriented neighbours. That works on a smooth closed object and fails at thin structures, where the two sides of a 3 mm leaf are neighbours in the graph and must not agree.
Poisson reconstruction
Poisson surface reconstruction, published by Michael Kazhdan, Matthew Bolitho and Hugues Hoppe in 2006, treats the oriented points as samples of the gradient of an indicator function — a function that is one inside the object and zero outside. Recovering the function from its gradient is a Poisson equation, solved over an adaptive octree; the surface is then the level set at the value where the input points sit, extracted with marching cubes. The original Poisson reconstruction paper describes the indicator-function formulation and the octree solver.
Two parameters carry the behaviour. Octree depth sets the finest resolution: depth d corresponds to a grid of 2d cells across the bounding box, so depth 9 is a 512-cell grid and depth 11 is 2,048. Anything thinner than one cell disappears entirely, and each extra depth multiplies memory and time by roughly eight. Point weight, the screening term added by Kazhdan and Hoppe in 2013, controls how hard the solution is pulled to interpolate the input points rather than smooth across them; raise it for clean scans, lower it for noisy ones.
The defining property of Poisson reconstruction — the reason it is chosen and the reason it surprises people — is that it always returns a watertight surface. There is no hole in the output because the indicator function is defined everywhere. Where you scanned nothing, it extrapolates smoothly and confidently, producing a bulging membrane across the gap. Every implementation therefore emits a per-vertex density value, and the standard practice is to trim the mesh at a density threshold afterwards, deleting the parts that were invented. Skipping the trim is the most common way a reconstruction ends up looking melted.
When Poisson is the wrong tool
- Ball pivoting rolls a virtual sphere of fixed radius over the points and creates a triangle whenever it rests on three. It only ever connects original points, so it never invents a vertex and never closes a hole — the opposite trade to Poisson. It needs roughly uniform density, because a radius that bridges the sparse region will also bridge across a genuine gap elsewhere.
- Delaunay-based and alpha-shape methods tetrahedralise the points and carve away tetrahedra that cameras can see through. This preserves sharp features better than a smoothed implicit surface, and it is the sensible choice for architectural scans with crisp edges that Poisson rounds off.
- Volumetric TSDF fusion integrates depth maps directly into a truncated signed distance field, which is what RGB-D scanning pipelines do in real time. If you have depth maps rather than a merged cloud, this skips normal estimation altogether — the sign comes from the ray direction.
- Neural implicit and Gaussian representations — NeRF-style radiance fields and 3D Gaussian splatting — optimise a continuous scene representation directly against the images and can beat the classical route on view synthesis of reflective or translucent material. They are a different pipeline with different failure modes and are not a drop-in replacement when the deliverable is a measurable triangle mesh.
Cleanup, and what a clean mesh is not
A raw reconstruction has non-manifold edges, tiny disconnected components, self-intersections and far more triangles than anything needs. Standard cleanup removes components below a vertex-count threshold, fills small holes, and simplifies with quadric edge collapse — Garland and Heckbert’s method, which chooses the edge whose removal least changes the surface’s quadric error and can usually take a mesh to a tenth of its triangle count with no visible change.
What none of this does is make the mesh accurate. Simplification is measured against the reconstructed surface, not against the object, so a mesh can be perfectly manifold, beautifully decimated and systematically two centimetres wrong because the scale of the reconstruction was never fixed against a known length. If the mesh is going to be measured, the accuracy check belongs on the point cloud before meshing — compare distances against surveyed control points, and gate on point density and noise before the surface stage rather than after it. For scans destined to become an as-built model rather than a render, that check is the whole job; see building a digital twin from a 3D scan.