Skip to content
// case study · 001

An 1860 lighthouse, rebuilt from its original drawings

Digital Montauk is a museum-grade digital twin of the Montauk Point Lighthouse — reconstructed strictly from the National Archives' 1860s construction drawings and period photographs, by one engineer directing a fleet of AI agents. It is our standing proof that generative systems can be held to a source of truth: audited, cited, and reproducible.

240
cited dimensions in the registry
92
numbered expert verdicts logged
24
agents in one adversarial audit
21
defects found, confirmed, fixed
12
archival sheets (NARA RG26)
70
licensed reference photographs
1860s construction drawing of the Montauk Point Lighthouse tower section, National Archives Record Group 26
the source — tower section, NARA RG26, traced at 74.5 px/ft
Procedural render of the digital twin, west elevation bench shot
the build — west bench, procedural geometry, calibrated color
// 01 · the problem

Beautiful AI output is not the same as correct AI output

Generative tools will happily produce a lighthouse that looks right. Ours has to be right — to the drawn dimension, at print scale, under the eye of people who know the real structure. That gap between plausible and correct is the same gap that stalls most commercial AI projects: MIT counts 95% of AI pilots delivering zero P&L impact, and the missing ingredient is rarely a better model. It is engineering discipline — a way to hold the output accountable to something true.

We built that discipline on a structure with an unusually demanding ground truth: a 166-year-old lighthouse with its original construction drawings preserved in the National Archives, and thousands of living witnesses who would spot a fake.

// 02 · source truth

Measurement is law

The foundation is twelve archival sheets from the National Archives' Record Group 26 — the lighthouse's 1860s construction set — plus a 70-file licensed reference corpus of period and modern photography. Each drawing is treated as an instrument, not an illustration: calibrated against its own scale bar, cross-checked by summing dimension chains until the sheet closes on itself, and read at full resolution with survey-style gridline overlays. Where a tear or repair crosses a measured span, we measure fragment by fragment. Where two views disagree, every independent instrument gets a vote and the count decides.

// 03 · the registry

Every dimension is a cited claim

Nothing enters the build as a number typed from memory. The project keeps a dimension registry — currently 240 entries — where every value carries its source sheet, its calibration method, the arithmetic that derived it, and a note on what it superseded. The procedural build scripts read the registry; they never hardcode. When a value is corrected, the correction cites the error it fixes. The registry is the audit trail that makes the whole model falsifiable — any dimension can be traced back to a line on a drawing.

// 04 · the loop

Ninety-two verdicts from a human eye

Instruments alone don't make something read as real. The project runs an eye-verdict loop: fixed camera benches render the identical shot every iteration, a human judge rules on the frame, and each numbered verdict becomes exactly one change in exactly one pull request — 92 verdicts logged so far, each one cross-checked against the drawings before any geometry moves. Roughly a third of the time, the eye flags something the archives turn out to have documented all along.

// 05 · adversarial audit

Twenty-four agents told to prove us wrong

In August 2026 we ran the entire tower builder through an adversarially-verified compliance trace: twenty-four AI agents, working in parallel, each tracing a section of code against the registry — followed by independent refuters whose only job was to kill each finding. The audit confirmed 21 real defects, ranked them by what a viewer could actually see, and every product-visible one was fixed within a day. This is what we mean by production discipline: the system that builds the model also audits the model, and the audit is designed to be hostile.

// 06 · instruments

Calibration, not vibes

Color is calibrated in a closed loop: we measure the ratio between painted zones in reference photographs, render the same ratio on a fixed bench, and iterate the material until the two agree — the current pass converges within 1.3%. A structural-similarity instrument pre-screens every change, localizing visual drift to the regions a fix was supposed to touch and flagging anything it wasn't. When a change claims "no visible difference," we have a number for that.

// 07 · why it matters

The same discipline, pointed at your operations

A lighthouse is our proving ground, but the method is the product: agents held to a source of truth, every output auditable, every claim cited, hostile verification before anything ships. That is exactly what separates the 5% of AI projects that reach production from the 95% that don't — and it is the discipline we bring to agentic-commerce work: storefronts that AI buying agents can actually transact with, and custom agents wired to your real operations with an audit trail you can stand behind.

// status

The twin is in its fidelity campaign now: the standard is a render a Montauk local cannot tell from a photograph, and nothing ships as product until it clears that bar. The registry, the verdict ledger, and the audit digests are maintained as the project's provenance record.

Book the 2-week diagnostic →

Two weeks. Your store, audited with this same discipline.