Migration — Staying Current
This content is for v0.4.1. Switch to the latest version for up-to-date documentation.
Paxman follows Semantic Versioning. The capability set, the contract surface, and the data tables grow across releases — your code should be ready to move forward without surprise.
In plain language: small releases add things; breaking releases can change things. This page tells you which is which and what to do when you upgrade.
Versioning
Section titled “Versioning”| Version bump | Meaning for your code | Example |
|---|---|---|
PATCH (0.0.X) | Contract compatibility preserved. Docs and internal fixes; authority data corrections may change recognition, status, or canonicalized_value when specs evolve — re-run golden samples even on PATCH. | 0.1.0 → 0.1.1 may update CLDR/IDNA tables |
MINOR (0.X.0) | Contract compatibility preserved (existing contracts still validate). Data-driven results may change when authority tables grow — pin paxman version and use contract.year (filters publication_year <= year) where point-in-time reproducibility matters, store version_stamp, and re-run golden samples. | 0.1.x → 0.2.0 adds capabilities and data |
MAJOR (X.0.0) | Breaking contract or flag semantics. Read the release notes — names, defaults, or canonical forms may change. | 0.x → 1.0.0 |
Contract compatibility (which contracts are accepted) is stable across PATCH and MINOR; result stability (which status/canonicalized_value you get) depends on data and is not promised when spec tables change. year filters rules by publication_year <= year; only pinned_rules and excluded_rules identify rules. Provenance and spec-version changes alone do not imply a MAJOR bump.
Determinism is per-installed-build: the version_stamp.paxman_version on every ExecutionResult records exactly which build produced the answer, so you can audit what changed across an upgrade. Since 0.2.0, version_stamp.recognition_revision (hash of the compiled matcher set + snapshot SHAs, per ADR-0009 §13) is the same-snapshot diff signal — if recognition_revision changes, recognition behavior changed for at least one capability even when paxman_version is unchanged (e.g., a lexicon token table update).
0.3.2 — ISBN/ISSN audit + kernel label fix (patch, non-breaking)
Section titled “0.3.2 — ISBN/ISSN audit + kernel label fix (patch, non-breaking)”Scope: patch 0.3.1 → 0.3.2 — contract compatibility preserved, no new capability, no flag change. Two audit truncation guards plus one kernel span fix; docs now cover ISSN as a first-class page.
What changed:
| Area | Before (0.3.1) | After (0.3.2) |
|---|---|---|
ISBN truncated hyphen continuation (1 0-306-40615-2, cut 978-… groups) | matched a truncated prefix | rejected → MISSING (trailing (?![-]\d) guard, both grammars) |
ISSN truncated hyphen-digit (0317-8471-2) | matched truncated 0317-8471 | rejected → MISSING (same guard class) |
IBAN bare-colon label (IBAN:AA0000000000000 after other text) | span covered core only (31, 46) — new/legacy parity gap | span absorbs label (26, 46, 'IBAN:AA0000000000000'), matching legacy |
ISBN helpers | _find_length duplicated in capability + rule | single find_registrant_length in rules/data/range_message.py (+ alias) — no behavior change |
| Docs | no docs/user/capabilities/issn.md | ISSN user page added; ISBN recognition-vs-validation table clarified |
Why this is correct: truncated prefixes are not valid identifiers (ISO 2108 / ISO 3297 check structure), so MISSING beats a truncated SUCCESS; the IBAN label belongs to the mention per the frozen legacy reference. normalize() now honors “Rules never raise” on direct calls for the touched ISBN rules.
How to detect: result.version_stamp.paxman_version 0.3.2 vs 0.3.1; recognition_revision bumps for the new trailing guards and the LabelMatcher scan change.
0.3.1 — IP audit fixes (patch, non-breaking)
Section titled “0.3.1 — IP audit fixes (patch, non-breaking)”Scope: patch 0.3.0 → 0.3.1 — contract compatibility preserved, no new capability, no flag change. One grammar completeness fix plus defensive hardening; docs now surface the new behavior.
What changed:
| Area | Before (0.3.0) | After (0.3.1) |
|---|---|---|
IPv6 mixed with embedded IPv4 (::ffff:192.0.2.1, 64:ff9b::192.0.2.1, ::192.0.2.1) | truncated to ::ffff:192 + second candidate 192.0.2.1 → AMBIGUOUS ['::ffff:192','192.0.2.1'] (truncated value) | single IPv6 candidate ::ffff:192.0.2.1 (plus trailing 192.0.2.1 via IPv4 \b) → AMBIGUOUS ['::ffff:192.0.2.1','192.0.2.1'] — prefer the IPv6 value (overlap documented, cross-grammar dedup deferred without ADR) |
IPv4 leading-zero (010.020.030.040) | recognized and normalized to 10.20.30.40 (already) — now documented | same, documented in docs/user/capabilities/ip.md and README.md |
IPNotation | @dataclass(frozen=True) | @dataclass(frozen=True, slots=True) + field docs — no behavior change |
normalize() | raised ValueError/AddressValueError on 999.999.999.999 / not-an-ip | try/except ValueError → returns input unchanged (never raises) |
| Provenance URLs | https://tools.ietf.org/html/rfc791 | https://datatracker.ietf.org/doc/html/rfc791 (same for RFC 5952) |
| Docs | IP page listed only compressed/expanded IPv6 | now lists mixed LS32 (RFC 4291 §2.2), leading-zero, overlap and triple-colon notes |
Why this is correct: 0.3.0 under-recognized Las32 mixed addresses that ipaddress.IPv6Address already accepts and RFC 4291 defines; 0.3.1 aligns recognition with the validated spec. The overlapping 192.0.2.1 candidate remains until an engine cross-grammar dedup policy exists — callers should pick the IPv6 value (see docs/user/capabilities/ip.md overlap note). normalize() now honors “Rules never raise” even on direct calls.
How to detect: result.version_stamp.paxman_version 0.3.1 vs 0.3.0; recognition_revision also bumps for the new mixed branches.
Migration snippet:
from paxman.api.bootstrap import register_all_shippedfrom paxman.api import canonicalizefrom paxman.capabilities.IP.contract import IPContract
register_all_shipped()r = canonicalize("::ffff:192.0.2.1", IPContract())# 0.3.0: AMBIGUOUS ['::ffff:192', '192.0.2.1'] (truncated)# 0.3.1: AMBIGUOUS ['::ffff:192.0.2.1', '192.0.2.1'] (correct, prefer IPv6)assert r.status.value == "ambiguous"assert r.candidates[0].value == "::ffff:192.0.2.1"0.2.0 — Recognition Kernel (breaking, scoped)
Section titled “0.2.0 — Recognition Kernel (breaking, scoped)”Scope: This is a pre-1.0 minor bump (0.1.0 → 0.2.0) with one intentional breaking change: the F1 correctness fix for whole-input vocabulary matching (ADR-0009). All other inputs are byte-identical under the parity gate.
What changed: Country name_recognition moved from whole-input lookup to an in-text word-anchored trie on the CountryNameFold view. Short-code grammars (alpha2) now honestly compete with the name grammar instead of silently winning when the name was invisible.
| Input class | Before (0.1.x pipeline) | After (0.2.0 kernel) |
|---|---|---|
Exact name, whole input ("United States") | SUCCESS "US" | SUCCESS "US" — unchanged |
Name embedded in prose ("Ship to United States please") | SUCCESS "TO" — wrong (Tonga) | MultipleMentionsError under a single_value contract; both mentions via paxman.scan() |
Short code as ordinary word ("to" in prose) | recognized and validated as alpha-2 — silent win | recognized; competes with the name mention — no silent win |
| All other inputs | — | byte-identical (parity gate) |
Migration snippet:
# Exact value — unchanged:import paxmanfrom paxman.capabilities.Country import Country
paxman.register_all_shipped()contract = Country.create_contract()paxman.canonicalize("United States", contract) # SUCCESS "US"
# Prose with embedded values — the new honest paths:from paxman.core.errors import MultipleMentionsError
try: result = paxman.canonicalize("Ship to United States please", contract)except MultipleMentionsError: # scan() shares one ScanContext substrate across all contracts in the batch mentions = paxman.scan("Ship to United States please", [contract]) # mentions.mentions["country"] == [ # Mention(span=(5, 7), grammar="alpha2_recognition", notation=...), # Mention(span=(8, 21), grammar="name_recognition", notation=...), # ] # Segment first — docs/recipes/segmentation.md remains valid for m in mentions.mentions["country"]: print(m.span, m.grammar, m.notation)
# Or segment-first (recipe) — `paxman.scan` is preferred; no direct# `normalize_name` import is needed.How to detect the change: Compare result.version_stamp.recognition_revision across builds. The F1 migration changes the compiled matcher set, so recognition_revision changes even if you stay on the same paxman_version snapshot. Store both paxman_version and recognition_revision for audit trails.
Why this is correct: 0.1.x returned a confident, provenance-backed, wrong answer on ordinary prose ("Ship to United States please" → Tonga). 0.2.0 surfaces the competition honestly; scan() turns the caller-owned split-then-canonicalize loop into an API (see docs/recipes/segmentation.md and paxman scan --help).
Other 0.2.0 additions (non-breaking):
paxman.scan(text, contracts)batch API +Mention/ScanResultmodel +paxman scanCLI (one substrate pass, seedocs/user/api-reference.md).ScanContextlazy views,MatcherSpec/LexiconMatchertrie (SIUnit 2.4–6.5× win at 650/820 tokens),BoundarySpecpresets,AnchorSetT0 prefilter.- Snapshot rails (
paxman/shared_data/*_snapshot.json+tools/regenerate_*+ CI drift gate) and derived recognition keys (BIC country codes, Language IANA subset) per ADR-0009 §14.
0.2.0 — Two-array offset maps (A4 Rev.4, breaking in spans only)
Section titled “0.2.0 — Two-array offset maps (A4 Rev.4, breaking in spans only)”Pre-0.2.0 the recognition kernel used a single-array D3 invariant
offsets[s] -> offsets[e] with a len(text) sentinel. When a
normalizer dropped source characters (CountryNameFold strips
punctuation, StripSeparators drops ()-., IDNAFold drops tabs) the
translated end absorbed the dropped tail: "United States."
recognized as (0, 14) with raw_text == "United States.".
ADR-0009 Rev.4 amends D3 to two arrays:
View.original_span(s, e) -> (starts[s], ends[e-1]) when mapped,
(s, e) when None; empty (0, 0). Each normalizer now returns
(subject, starts, ends) with len(starts)==len(ends)==len(subject)
and 0 <= starts[i] < ends[i] <= len(text).
Visible change (Option 1, word-boundary-aligned mentions):
| Input | Before | After |
|---|---|---|
"United States." name mention | (0, 14) raw_text="United States." | (0, 13) raw_text="United States" |
"United States of America," | (0, 25) includes trailing , | (0, 24) trimmed |
"+1 (555) 123-4567" via StripSeparators | ends[-1]==len(text) sentinel | ends per-char s+1, original_span(0,n)==(0,17) exact |
All length-preserving views (CaseFold etc.) | None | None, None — zero-cost unchanged |
Whole-input canonical values are unchanged (rules normalize);
only span/raw_text presentation shifts, and scan() mentions no
longer carry trailing dropped punctuation. raw_text == text[start:end]
is now an engine invariant enforced for every emitted match.
Migrate: if you stored span for later slicing, re-derive it from the
new ExecutionResult.span/Mention.span; do not add +1 for dropped
chars. Golden samples that asserted (0, 14) for "United States."
should assert (0, 13).
0.4.0 — Whole-input suppression exemption (A0, #122)
Section titled “0.4.0 — Whole-input suppression exemption (A0, #122)”Calling canonicalize() with a contract asserts the kind — “a canonical
value is derivable from this input” — so suppressing the whole input
contradicted the asserted intent (MISSING indistinguishable from
canonicalize("")). Under suppress_common_words=True, a suppressible
word-bounded hit that covers the entire trimmed input is now never
suppressed (ADR-0009 Rev.5, §16 amendment). Embedded mentions stay
suppressed — scan() prose behavior is unchanged.
Behavior change (only under suppress_common_words=True; flag-off
results are byte-identical):
| Input | Contract | Before (0.2.0–0.3.x) | After (0.4.0) |
|---|---|---|---|
to / TO / to | Country, suppress on | MISSING | SUCCESS "TO" |
ALL | Currency, suppress on | MISSING | SUCCESS "ALL" |
en | Language, suppress on | MISSING | SUCCESS "en" |
in/ | Country, suppress on | MISSING | MISSING (only the whole input is exempt), suppressed_count=1 |
to and usa | Country, suppress on | SUCCESS "US" | unchanged (embedded to/and and the α3 usa hit stay suppressed — usa ∈ COMMON_WORDS; survival is via the non-suppressible name_recognition hit at the same span) |
cd | SIUnit, suppress on | SUCCESS "cd" | unchanged (no SIUnit matcher is suppressible) |
This supersedes the whole-input row of the 0.2.0 suppression note below
(canonicalize("to", … suppress on) → MISSING no longer holds); the
rest of that note (table, matchers, scan guidance) still applies.
New ExecutionResult signal — suppressed_count: int = 0 and
suppressed_spans: tuple[tuple[int, int], ...] = (), populated whenever
suppression fires (on MISSING and INVALID, not just MISSING;
0/() when the flag is off), so MISSING + suppressed_count == 1
(“recognized but suppressed”) is distinguishable from MISSING +
suppressed_count == 0 (“nothing recognized”):
result = paxman.canonicalize("in/", Country.create_contract(suppress_common_words=True))assert result.status == Resolution.MISSINGassert result.suppressed_count == 1assert result.suppressed_spans == ((0, 2),)A1 rejected: the x→0 fallback (keep the unsuppressed set when
suppression would leave zero mentions, e.g. "to and is") is evaluated
and rejected in #122 — suppression-to-MISSING there is the desired
noise reduction, now observable via the signal instead of silent.
Migrate: if you worked around whole-input suppression (flag-off
contracts for bare codes, special-casing MISSING for to/ALL/en),
you can drop the workaround and pass the suppression contract straight
through — whole-input canonical values now re-enter as fixed points
under suppression (ADR-0010 property suite, #123 cross-link). See
ADR-0009 Rev.5.
0.4.0 — Phone national de-offered, split successor (breaking, ADR-0011 Phase 2)
Section titled “0.4.0 — Phone national de-offered, split successor (breaking, ADR-0011 Phase 2)”PhoneContract(output_format="national") now raises ContractError with a
migration message naming split (was SUCCESS rendering the bare NSN for
NANP values, E.164 otherwise). national dropped the country code that
recognition/validation depend on and could not re-enter under the default
contract (param dependence; value-dependent shape).
| Before (0.3.x) | After (0.4.0) |
|---|---|
Phone.create_contract(output_format="national") → SUCCESS "2125551234" | raises ContractError (migrate to split) |
Phone.create_contract(output_format="split") → ContractError | SUCCESS "+1 2125551234" (uniform +CC NSN, param-free re-entry) |
Migrate: replace output_format="national" with
output_format="split" for storage and round-tripping; e164/rfc3966
unchanged; default_country remains input-only for national-shaped input.
0.4.0 — Language alpha2 carries subtags (changed, ADR-0011 Phase 3)
Section titled “0.4.0 — Language alpha2 carries subtags (changed, ADR-0011 Phase 3)”alpha2 now maps the primary subtag and carries the remainder verbatim
(en-US → en-US, was en). alpha3/alpha3-bib render the primary only
for extended tags (de-CH-1901 → deu/ger) and stay waived projections
that re-enter as fixed points. Bare-code rendering unchanged.
Migrate: if you asserted alpha2 strips subtags, update goldens to the
carried form; use name only for display (waived projection).
0.4.0 — MacAddress bit_reversed de-offered (breaking, ADR-0010)
Section titled “0.4.0 — MacAddress bit_reversed de-offered (breaking, ADR-0010)”MacAddressContract(output_format="bit_reversed") now raises
ContractError (was per-octet bit-swap SUCCESS). The view is an involution
(f(f(x)) == x), not a fixed point, so it is no longer offered.
Migrate: use default colon for storage; compute Token-Ring display locally.
0.4.0 — New capabilities: Element + Coordinates (additive)
Section titled “0.4.0 — New capabilities: Element + Coordinates (additive)”Element (IUPAC symbol canonical, name format) and Coordinates
(lat-first decimal canonical, decimal/iso6709/geo_uri/geojson_pair/dms/dm
formats) are registered by register_all_shipped() (now 18 shipped).
MacAddress registration is also activated. No migration required —
existing contracts are byte-identical.
0.2.0 — Common-word suppression for scan (B1, ADR-0009 §16)
Section titled “0.2.0 — Common-word suppression for scan (B1, ADR-0009 §16)”ADR-0009 §16 was deferred as non-binding; it now ships off by default as the
suppression table for usable scan() on prose (R7: ~80% of prose scan()
hits were short-code noise like to→Tonga). The change is additive and
off-by-default — byte-identical for every existing caller.
Contract: CapabilityContract.suppress_common_words: bool = False
(after extra_grammars; frozen no-slots). Every capability’s
create_contract(..., suppress_common_words: bool = False) forwards it.
Default False preserves existing canonicalize() / scan() results;
True removes word-bounded short-code recognitions whose lowercased span
is in the curated table — never canonicalizing, only suppressing
recognition (provenance-neutral by construction).
Table: paxman/core/grammar/data/common_words.py:COMMON_WORDS
frozenset[str] curated via
Google 1000 (https://github.com/first20hours/google-10000-english, google-10000-english.txt first 1000 lines) ∩ (ISO 3166 α2/α3 + ISO 4217 + ISO 639-1/2/3) lowercased, reviewable, frozen with assert len(COMMON_WORDS)==67 and assert "USD" not in COMMON_WORDS. USD is
deliberately not suppressed (not in Google 1000); currency scan() keeps
USD while to/in etc. are removed for country/currency/language
code shapes.
Matchers: short-code matchers marked suppressible=True (declaration,
not per-grammar code): Country alpha2_recognition / alpha3_recognition
/ numeric_recognition, Currency code_recognition, Language
language_code_recognition. Boundary is already word-bounded
(BoundarySpec.WORD / WORD_SIGN — required), so suppression only fires
on word-bounded hits.
Engine: paxman/core/grammar/engine_loop.py insertion between
view.original_span and emit: when contract.suppress_common_words
and matcher.suppressible and text[o_s:o_e].lower() in COMMON_WORDS
→ continue (skip emit).
Scan vs canonicalize:
scan("Ship to the United States of America, total 45.50 USD, weight 3.5 kg", [Country.create_contract(suppress_common_words=True)])keeps only the name mentionUnited States of Americafor the short-code shapes (plus numeric/already-word-bounded non-common-word hits); with the flag off the full current snapshot is preserved.Currencyscan keepsUSDin both modes.canonicalize("to", Country.create_contract(suppress_common_words=False))→SUCCESS "TO"(Tonga correct for bare code); with the flag on →MISSING(suppressed recognition, never validated).
CLI: paxman scan keeps default contracts (flag off) but now accepts
--suppress-common-words — thin create_contract(suppress_common_words=True)
construction only (see paxman scan --help). API remains the seam for
canonicalize().
ADR: Rev.4 §16 “deferred, non-binding” is superseded — the table ships in
0.2.0 off by default. recognition_revision bumps for the new matcher marker
(one-time, pre-release).
Migrate: no change required. Opt in only for scan() on prose where Tonga
noise matters; keep bare-code canonicalize() contracts flag-off.
What can appear in a minor release
Section titled “What can appear in a minor release”You do not need to change code for these — they are additive and backward compatible:
- New capabilities (the set of importable names under
paxman.capabilitiesgrows — never treat a current count as final). - New contract flags of the form
include_*,allow_*, ordefault_*(always optional, defaults preserve shipped behavior). - New offered
output_formatalternatives (the default rendering stays the same; pinoutput_formatif you rely on a specific rendering). - Expanded authority tables (e.g. new CLDR entries, additional URL IDNA mappings) where the spec itself grew.
Keep your registration future-proof by preferring the explicit form when you care about the surface:
# Future-proof — only the capabilities you namepaxman.register_capability(Email())paxman.register_capability(Date())
# Convenience — everything shipped in this buildpaxman.register_all_shipped() # convenient, but the set it registers grows over timeEither approach is supported; pick explicit when you want the upgrade to be a conscious decision, and bootstrap when you want the new capabilities automatically.
What signals a careful upgrade
Section titled “What signals a careful upgrade”These are major-bump signals — check the release notes and review the checklist below:
- A
DEFAULT_OUTPUT_FORMATorOFFERED_OUTPUT_FORMATSchange — the string behindcanonicalized_valuefor the same input may differ even thoughstatusstaysSUCCESS. - A rule’s provenance year or spec version changes — treat as an audit signal:
contract.yearboundaries move andinclude_historicalvs active coverage may shift. Alone it is not a MAJOR trigger (see above). - A capability renamed, merged, or removed.
Upgrade checklist (copy-paste)
Section titled “Upgrade checklist (copy-paste)”When moving to a new MINOR, glance; when moving to a new MAJOR, work through it:
- Read the release notes — skim new capabilities, new contract flags, and any new
output_formatvalues. Decide whether a new capability belongs in your registration. - Review contracts — new optional flags default to shipped-preserving values. Verify that rule names in
pinned_rulesandexcluded_rulesstill exist (a stale name raisesContractError), and separately confirm thatyear(filterspublication_year <= year) still expresses the temporal window you intend —yearis not a rule name. - Pin output if it matters — if your downstream expects a specific rendering (e.g.
Phonerfc3966), construct the contract withoutput_format="rfc3966"rather than relying on the current default. - Re-run your golden samples — keep a small file of
(text, contract) → canonicalized_valuesamples for the capabilities you use, assert them in CI, and compare after the upgrade. Determinism means a change is intentional, not noise. - Log or store
version_stamp— for audit trails, persistresult.version_stamp.paxman_versionalongsidecanonicalized_valueso you can explain which build produced which answer. - Segmentation review — if you added a new capability or flag, confirm the caller-owned split-then-canonicalize loop (see Segmentation) still routes each piece to the right capability/contract.
Re-entry: a SUCCESS canonical value V is safe to feed back as canonicalize(V, C) → SUCCESS V for any output_format under the same default contract (ADR-0010, #123). Custom pinned_rules/excluded_rules/year, and embedded or otherwise non-exempt suppressed matches, remain conditional.
Minimal golden-sample harness:
import paxmanfrom paxman.capabilities import Email, Countryfrom paxman.core.domain import Resolution
paxman.register_all_shipped()
checks = [ ("user@Example.COM", Email.create_contract(), "user@example.com"), ("United States", Country.create_contract(), "US"),]
for text, contract, expected in checks: r = paxman.canonicalize(text, contract) assert r.status == Resolution.SUCCESS and r.canonicalized_value == expected, ( text, r, )Temporal filtering and data drift
Section titled “Temporal filtering and data drift”Spec tables (CLDR, ISBN Range Message, URL IDNA) are regenerated from snapshots and live inside the library. When a spec evolves (new country names, new currency symbols, new IDNA mappings), the release notes will note it. Use contract.year to pin to specs published up to a given year when reproducibility against a point-in-time authority matters; combine it with a pinned paxman version in your environment for full reproducibility.
See Contracts for year filtering, Provenance for publication_year on each citation, and Execution Result for version_stamp.
See also
Section titled “See also”- API Reference — registration, contracts, and statuses
- Concepts — Pipeline — why statuses are stable (recognition → validation → resolution)
- Extending — keeping community grammars/rules compatible across upgrades
- Segmentation — caller-owned splitting for multi-entity text