All watches

Changelog

Every change to WatchDB, newest first — each with its goal, method, and results.

Catalognew watches added from a catalogDatareferences, calibers, coverageTaggingwatches assigned to collectionsModelthe data / identity modelAppwebsite features you can see
Show internal changes (method, infra)

August 16, 2026

Data21:xx

King Seiko: decode the JDM codes, recover the vintage mechanicals

Goal

Populate the King Seiko sub-collections, which held ~0 identified mechanicals even though the watches are in our data.

Method
  • The catalogs print King Seikos as JDM codes (KSK, KSSK, KSCM, 45KSC, 56KAW) that the caseback builder never decoded, so they sat untyped under likely_collection King Seiko.
  • Build research/king_seiko.md from kingseiko.info (158 documented models, per-series reference tables).
  • Write pipeline/scripts/decode_king_seiko.py: code prefix maps to caliber, 3-digit case design from the row's own text, caseback = caliber + design + region 0.
  • Verify the 56-series against Boley (covered); trust the documented mapping for the 44/45/52-series (Boley has none), validated against kingseiko.info.
  • Dry-run, back up, apply; then tag the caliber-bearing 56KS rows with assign_collections.
Results
  • Decoded 25 coded rows.
  • King Seiko identified models: 44KS 0 to 3, 45KS 0 to 3, 56KS to 13; total 19 (was ~0 mechanical).
  • 199 King Seiko rows carry a code but no case design in the extracted text — the next recovery target.
  • 52KS and Vanac still 0 (code-encoded without a design, or ambiguous).
Rejected

Requiring a Boley match for all King Seiko. Boley covers only the 56-series (5621/5625/5626); the 44/45/52-series are pre-Boley hand-wounds, so we trust the documented code-to-caliber mapping plus the row's own case design.

watchdb.sqlite.bak-20260816-134242-decodeKS

Data20:xx

Recover the 62MAS — both 6217-8000 and 6217-8001 (collection complete)

Goal

Find the 62MAS, Seiko's first diver, which held 0 identified models.

Method
  • The catalog prints the caliber as 62MAS, not 6217, so the earlier caliber search missed it.
  • Two rows carry 62MAS 010 (800-800) at 17 jewels: id 356 in the 1967 catalog and id 616 in the 1968 No.1 catalog.
  • TheSeikoGuy documents the catalog-year mapping: 1966-67 shows the small crown 6217-8000, 1968 vol.1 the big crown 6217-8001.
  • Boley lists both references.
  • Set caliber 6217, caseback_ref, and collection 62MAS: id 356 to 6217-8000, id 616 to 6217-8001.
  • Widen the 62MAS window to 1965-1968 in the registry so the rule stays reproducible.
Results
  • 62MAS now holds both references, 6217-8000 and 6217-8001: 2 of 2, 100% complete.
  • First hunt-collection closed from 0 to done.

watchdb.sqlite.bak-20260816-133136-62mas

App19:xx

Status: 6159 Diver at 100%, two-decimal headline shares

Goal

Reflect the completed 6159 collection and show precise headline percentages.

Method
  • Set 6159 Diver est to 3 (its three references 7000, 7001, 7010) and Confirmed in research/collections-logic.md; regenerate collections.ts.
  • Show the Identified, Right name, and Right collection headline shares to two decimals on the Status page.
Results
  • 6159 Diver reads 3 / 3 · 100% on Status.
  • Headline shares now show two decimals (right name 66.56%, right collection 0.05%, identified 0.03%).
Data18:xx

Recover the missing 6159-7001 (1969 300m hi-beat diver)

Goal

Add the missing 6159 reference — the 1969 300m hi-beat monocoque diver that we lacked.

Method
  • Two caliber-6159 Seiko Diver rows (1969 and 1970) carry the catalog case code 011; the 010 code is the proven 6159-7000, so 011 is the increment.
  • Confirm 6159-7001 exists in Boley independently, with its own parts (crystal 325W14GA).
  • Domain-expert confirmation that 6159-7001 is a real JDM reference.
  • Set caseback_ref = 6159-7001 on ids 2148 and 2678.
Results
  • 6159 Diver now holds its three references: 6159-7000, 6159-7001, and 6159-7010.
  • ids 2148 and 2678 recovered from caseback-less to identified.

watchdb.sqlite.bak-20260816-132228-6159-7001

App17:xx

Hide the case-material field and filter

Goal

Stop showing case material until we trust the data.

Method
  • Hide the Material filter in the sidebar and the case material row on the watch page.
  • Keep both in code, commented, so they are easy to restore once the material data is trusted.
Results

Case material no longer appears in the filter sidebar or on the watch page.

App16:xx

One page container across the app + a prominent collection callout

Goal

Give every page one container, and make a watch's collection and details easy to read.

Method
  • Add a shared Container (max-w-5xl, centered, one padding) used by the watch, changelog, and status pages.
  • Narrow the browse index to max-w-6xl (was full-bleed) with the same centering and padding; drop the 5-column grid tier.
  • On the watch page, turn the collection paragraph into a callout: accent panel, icon, collection-name heading, readable foreground text, placed first under the name.
  • Show every specification in one simple card, one per line (identity, case, dial, and caliber specs together).
Results
  • The index, watch page, changelog, and status share one container width and padding.
  • The collection reads as a clear callout, not gray small print.
  • Every watch specification sits in one card, one per line.
Rejected

A single identical width for every page. We kept the browse index one step wider (max-w-6xl) than the reading pages (max-w-5xl) for grid density.

App15:xx

Show only identified watches; unify model identity + title format

Goal

Make counted, shown, and distinct mean one thing, and display only watches that have a real model name.

Method
  • Define one identity: the model name = caseback reference, else sales code, else pre-1966 case code (caseback first, so one watch keeps one identity even when a later catalog also prints its sales code; ERA_NAME_SQL in lib/db.ts).
  • Show a watch only if it has that name; hide the rest across the grid, the API feed, and the facet counts.
  • Dedupe the catalog by that same name, so shown count equals distinct count.
  • Change the title to <model name> — <collection> (lib/identity.ts, modelName).
  • Count the facet sidebar and the Status have-column by that same identity, so a collection shows the same number everywhere.
Results
  • The app shows 7,753 identified models.
  • 6159 Diver: 2 identified models, 6159-7000 and 6159-7010, each titled e.g. 6159-7010 — 6159 Diver.
  • Un-referenced watches no longer appear anywhere.
  • Counted, shown, and distinct now agree everywhere: 6159 Diver reads 2 on the grid, the facet sidebar, and Status.
Rejected

A year-gated rule (sales code for 1975+). It split one physical watch into two cards when a later catalog also printed its sales/order code (6159-7010 vs YAQ028). Caseback-first fixes it.

App14:xx

Catalog shows distinct models, random order, collection-aware names

Goal

Show the catalog as distinct watches in random order, each named by its collection.

Method
  • Grid dedupes to one card per distinct model, keyed by caseback reference, then modern sales code, then model_key; a window function keeps the best appearance (has a crop, has a reference, earliest).
  • Every appearance stays in the DB as provenance; the watch page still lists the other catalog appearances.
  • Default order is a fresh random seed per page load, threaded through infinite scroll so the pages stay consistent.
  • Card and page titles now prefer the assigned collection over a soft hint or a raw catalog heading (lib/identity.ts).
Results
  • The catalog shows 12,836 distinct wristwatch models instead of 35,412 appearances.
  • 6159 Diver: 14 appearances collapse to 5 distinct models.
  • Tagged watches read e.g. 6159 Diver 6159-7010, and caseback-less ones read 6159 Diver 1969 instead of Special Watches 1969.
  • Arriving on the home page shows a fresh random order, not oldest-first.
Rejected

Relabeling duplicate appearances as unknown in the DB. Appearances are real provenance (the same model in another catalog year, with its own crop), not errors; we dedupe at the query layer instead, which is reversible and keeps the history.

Tagging13:xx

First collection slice: 6159 Diver (Tuna) + the assign_collections tool

Goal

Build the re-runnable collection assigner and complete the first vertical slice.

Method
  • Add pipeline/scripts/assign_collections.py: registry-driven rules (caliber or name), idempotent, with a dry-run, a per-slice apply that backs up the DB first, and a conservation report by decade.
  • A model that matches two collections is left Unclassified, never guessed.
  • Verify the 6159 candidates against the catalog text, then apply.
Results
  • 6159 Diver: 7 models (14 rows) assigned, all caliber 6159, verified.
  • Dry-run map: base rules auto-assign 4,960 of 15,065 models (33%); 10,105 Unclassified; 110 ambiguous models flagged for tiebreakers.
  • Conservation balances in every decade.
  • Collections assigned in the DB: 1.
Rejected

Assuming the caliber spine is complete. Many Grand Seiko and King Seiko families hold 0 models by caliber, because their calibers were never extracted; those families need hunting or name recovery first, not tagging.

watchdb.sqlite.bak-20260816-122541-assign-6159Diver

App12:xx

Structured change log (CSV) + /changelog page

Goal

Keep one structured file we append to for every change, and show it clearly at /changelog.

Method
  • Store every change as a row in CHANGELOG.csv (date, time, type, title, goal, method, results, rejected, backup).
  • Parse the CSV at request time in apps/web/lib/changelog.ts.
  • Render each change as a card with a type badge and labeled Goal, Method, and Results sections.
  • Add a per-collection paragraph on each watch page from research/collections-paragraphs.md.
  • Rename research docs to kebab-case and add research/README.md.
  • Add CLAUDE.md with the logging rule and the change-type taxonomy.
Results
  • /changelog parses the CSV directly, so there is no regenerate step.
  • 81 of 81 collection paragraphs matched.
  • typecheck passed.
  • /changelog and /watch/176 returned HTTP 200.
Rejected

An intermediate generated TS file for the changelog. We replaced it by parsing the CSV at runtime, so appending a row needs no script.

August 15, 2026

Model21:xx

New identity model: three tiers, the 1956 floor, 83 MECE collections

Goal

Rebuild how a watch is identified, with exactly one home per watch.

Method
  • Define three identity tiers by year: 1956-1965 family + case code (case_code); 1966-1974 caliber + case built reference (caseback_ref); 1975-1998 sales code (sales_code).
  • Set a 1956 floor with a new in_catalog column (0 for year below 1956; hidden everywhere).
  • Define 83 MECE collections by caliber family.
  • Move the old line column to likely_collection; add an empty collection column.
  • Rebuild the Status page around three shares: Identified, Right name, Right collection.
Results
  • 55 pre-1956 rows hidden (the 1932 catalog).
  • Name coverage 67% (10,026 of 15,065 models); collection 0%; identified 0%.
  • Names by era: 1956-65 = 0%; 1966-74 = 12% (651 of 5,258); 1975-98 = 96% (9,375 of 9,762).
  • The collection filter now lists the real 81 collections.

backups .bak-20260815-likelycol and .bak-20260815-incatalog

Tagging20:xx

Collection audit: year-window check and tag fixes

Goal

Find watches tagged to a collection in a year that collection did not exist, and know per collection where we stand.

Method
  • Build a production window per collection from all-collections.md and the horological record.
  • Flag every watch whose catalog year sits outside its window.
  • Spot-check each flagged set against the raw text.
  • Move the confirmed mis-tags.
  • Set windows for 5 Actus and Skyliner.
  • Rebuild the Status page and add a missing-collections section.
Results
  • 325 watches flagged across 5 collections; 34 passed; 3 unchecked (no window).
  • Marvel after 1968 moved to Lord Marvel (28 rows).
  • Laurel from 1980 moved to Laurel Quartz (210 rows).
  • Marvel now 1956-1966 (14); Lord Marvel now 1967-1978 (57).
  • Still flagged for review: Skyliner 1969-74 (21) and 5 Actus 1983-85 (13).

backup .bak-20260815-collfix

Tagging19:xx

Section headers leaked as watch names; a new tagging source

Goal

Stop catalog section headers showing as watch names, and use them to fill the empty collection.

Method
  • Detect a section-header heading (it ends in Group or Series) and stop using it as a watch name (lib/identity.ts, isSectionHeader).
  • Read the collection out of the header and write line for untagged rows, only where the header names a real collection.
  • Add a /status page.
Results
  • watch 2573 now reads Seiko 1970; zero headers leak as names.
  • Tagged: 5 Actus 6 to 370; Avenue 0 to 668; Arctura 22; Kinetic 13.
  • Reclassified 174 pre-1962 rows to family + year.
Rejected

Generic category headers (Business Group, Dress Watch Group, Calendar Group). They name a category, not a collection. Inventing a collection from them would break the never-invent-a-fact rule.

backup .bak-20260815-collections

Data17:0x

LinkUp API: jewel lookup works, sales-code lookup fails

Goal

Use the LinkUp API to resolve ambiguous calibers by jewel count, and sales codes to references.

Method
  • Build a client pipeline/scripts/linkup.py that caches every answer and fixes the macOS certificate error.
  • Store every answer and citation in data/sources/linkup/audit.jsonl and the table linkup_audit.
  • Move 5 (jewels): match the printed jewel count to exactly one caliber (enrich_jewels.py).
  • Move 2 (sales code): probe the reference with strict Boley rules.
Results

Move 5 done: 65 of 179 ambiguous rows resolved; wristwatch references 25,753 to 25,818.

Rejected

Move 2 stopped. Deep depth ($0.055 per query) invented text and returned generic references; standard depth found 0 of 20 rows; JDM order codes are not indexed on the web. We wrote nothing.

backup .bak-20260815-jewels

Data16:4x

Reclassify pre-1962; apply the held sales-code tier

Goal

Fix the tier of pre-1962 watches and apply the held sales-code references.

Method
  • Set all 112 pre-1962 wristwatches to the tier pre_numbering, because the caliber-case system did not exist then.
  • Measure coverage only over watches that can have a reference.
  • Write the 168 held sales-code references (recover_salescode.py --include-unconfirmed).
Results
  • 31 rows corrected from the wrong tier none.
  • Coverage 70.4% (25,753 of 36,604).
  • Wristwatch references 25,585 to 25,753.

backup .bak-20260815-moves34

Data16:3x

Sales-code self-join (safe tier)

Goal

Recover references by copying within the DB, using only safe matches.

Method
  • recover_salescode.py. A sales code is a unique model id; if the same code carries a reference on another page, copy that reference.
  • Use only codes that map to one reference.
  • Write only the matches that are also in Boley.
Results
  • 4,920 stuck rows have a sales code; 423 resolve by self-join (255 are also in Boley; 168 are self-consistent only).
  • Wrote the safe 255; all gained parts.
  • Coverage 25,330 to 25,585 (69.0% to 69.7%).
Rejected

Held the 168 self-consistent-only rows, because they are not proven outside our own data.

backup .bak-20260815-salescode

Data16:1x

Unified recovery pass (code grammar + JDM-0 + OCR repair)

Goal

Recover 1966-1975 references that are stored in the text but not decoded.

Method
  • recover_refs.py: one Boley load, one DB scan, one write; every method checked against Boley's 38,142 references.
  • code_grammar: the LM/KS/GS formula (56LMW to 5606) plus family+case that is unique in Boley.
  • variant_jdm0: a Boley group resolved by the JDM region-0 rule.
  • ocr_repair: a garbled caliber next to its own case bracket.
  • Attach parts in the same pass with concurrent workers.
Results
  • 279 references, all confirmed in Boley, all gained parts.
  • variant_jdm0 holds fully: 0 of 238 break; 59 resolve; 179 stay ambiguous.
  • Coverage 25,051 to 25,330 (68.2% to 69.0%).
  • Boley-linked 18,598 to 18,877.
  • The pass takes about 1 second.
Rejected

The numeric code form (61-5A). Family 61 has many calibers, so it is not safe.

backup .bak-20260815-recover

Data15:19

Caseback backfill from stored text

Goal

Raise strict caseback_ref coverage for wristwatches (start 23,659 of 36,716, 64%).

Method
  • backfill_caseback.py, dry-run first, then --apply.
  • Recover a missing caliber or case only from evidence tied to the same row: the 56LMW grammar, the row's own refs, or a bracket tied to the row's caliber.
  • Re-link parts with boley_relink.py.
Results
  • 1,392 references.
  • Coverage 23,659 to 25,051 (64% to 68.2%).
  • 1,282 watches gained parts; Boley-linked 17,316 to 18,598.
  • Online check of 10 new references: every caliber real; 7 exact references confirmed; 3 not found exactly but each real; 0 fakes.
Rejected

A loose scan of the whole page. On multi-watch pages it took a neighbour's case number (6117 010, 012, 022 all became 6117-6010). It offered about 2,000 rows but made wrong ids. Precision before recall.

backup .bak-20260815-caseback

August 8, 2026

App17:48

Simpler client UI (for collectors)

Goal

Remove jargon from the client so collectors read it easily.

Method
  • Show a reference (5606-7000) or a name (Marvel 1956) as the title; neither shows Unidentified.
  • Show case materials as words and movements as Automatic, Quartz, or Hand-wound (lib/labels.ts).
  • Remove the per-field confidence badges, the Identity filter, and the Caliber filter.
  • Keep Collection, Movement, Dial colour, Case colour, Shape, and Material.
Results
  • Cleaner card and detail titles.
  • The detail page shows Reference, Collection, Caliber, Year, and plain attributes.
  • The logo links home.
Data17:37

boley.de case-parts harvest + cross-reference (backend only)

Goal

Get an independent reference database to confirm our reconstructed references.

Method
  • harvest_boley.py (polite, concurrent, checkpointed) pulls the full boley.de Seiko case-parts database.
  • A build_webdb.py post-pass checks our caseback_ref against Boley and attaches part numbers.
  • Strip the source from every field in lib/db.ts so it never reaches the client.
Results
  • 38,142 references and 715,476 part rows harvested.
  • 17,260 confirmed (likely to verified); 56 corrected; parts attached to 17,316 watches.
  • Checked: 0 instances of boley in the client HTML.
Model17:13

Caseback reference (caseback_ref): the collector identity

Goal

Store the caseback reference collectors use (5606-7000), not the one-digit-short catalog form.

Method
  • Derive caseback_ref = caliber + 3-digit case design + the JDM region digit 0 (derive_identity.caseback_ref).
  • Make it the main reference in vintage titles.
  • Make it searchable by exact match and FTS.
Results
  • It reproduces the exact casebacks (5606-700 to 5606-7000).
  • 25,159 watches carry one (4,620 distinct).
  • Confidence is capped at likely, because the region digit is inferred.
App17:00

One README; infinite-scroll app shell

Goal

Merge the scattered methodology docs into one README and rebuild the home page.

Method
  • Merge every methodology doc into one README.md; remove the scattered files and the client /methodology page.
  • Rebuild the home page as a fixed app shell with a pinned header and sidebar.
  • Replace pagination with infinite scroll (/api/watches feed + InfiniteGrid).
Results
  • One clean README.
  • Only the catalog grid scrolls.
  • Commit 2671bf5.
Model16:50

Monorepo migration + one canonical model_key per watch

Goal

Give each watch one canonical identity and move the repo to a monorepo.

Method
  • Add one canonical model_key per watch (with model_key_source, family_id, identity_tier).
  • Merge on the validated sales_code column; add a junk-ref guard; use a linear 3-phase build order; add audit_db.py, a read-only truth audit.
  • Date the JDM code eras from our own catalogs.
  • Convert to a pnpm/Turborepo monorepo: apps/web on the packages/ui design system, and pipeline/ outside the JS workspace.
Results
  • 40,332 appearances reconciled into 16,327 distinct SKUs and 4,941 case families (before, an inflated 20,276).
  • 200 family_id conflicts to 0; about 2,600 under-merges recovered.
  • Commit 2df9ddd.

August 7, 2026

App18:30

First browser: wristwatch UI, evidence view, EN/JP toggle

Goal

Show the extracted data to a reader, with honest confidence.

Method
  • Build a wristwatch-scoped browser with an identity block and an evidence view.
  • Show a confidence badge on each field.
  • Hide the price.
  • Add an English/Japanese toggle for labels, values, and line names.
  • Add a per-catalog confidence coverage dashboard.
  • Scope the facets to the subject and clean the URLs.
Results
  • The first browsable WatchDB.
  • We simplified it later, on 2026-08-08, to the collector model.
Model18:09

Schema gate + per-field confidence tiers

Goal

Keep the data honest, with one vocabulary and a confidence tier per field.

Method
  • Add a schema gate (schema.py): single-source vocabularies, validate_row, and a schema.json export.
  • Route rows that fail the schema to a violations table.
  • Derive sales_code and the marketing line as distinct identity fields.
  • Set a rule-based confidence tier per field, cross-checked against the caliber DB, the era, and the vision pass.
Results
  • Every field carries a confidence tier and a source.
  • Only schema-valid values reach the app.

August 6, 2026

Catalog14:53

Catalog extraction: PDFs to one row per watch (Gemini on Google Cloud)

Goal

Turn scanned Seiko catalog PDFs into one data row per watch photo.

Method
  • Download about 129 Seiko catalog PDFs from the Internet Archive (01_download_catalogs.py).
  • Render each page to PNG with PyMuPDF at 300 DPI.
  • Read each page with Gemini 2.5 Flash on Google Cloud Vertex AI (aiplatform.googleapis.com), with an OpenRouter fallback.
  • Per page, detect each product unit by bounding box, transcribe the full page text, read each caption on its own text crop, link captions to photos with a numbered-box overlay, and cut three crops per watch.
  • Run a second vision pass with Gemini 2.5 Flash-Lite (vision_colours.py) to read dial colour, case colour, shape, and strap from each photo.
Results
  • One row per watch photo across the catalogs, each with its verbatim caption and crops.
  • The caption gives sales code, jewels, material, caliber, case number, and line; the photo gives colour and shape.
  • This is the origin of every watch in the database.