Extending Paxman — Community Grammars & Rules
Paxman ships with capabilities that already cover their domain, but every capability is closed for modification yet open for extension: you can add new recognition and validation without forking the library. A new grammar finds a new form; a new rule says which spec accepts it.
In plain language: if Paxman handles
2024-01-01but your data also contains2024.01.01, you teach it to spot the dot form and which spec says that dot form is valid — without changing Paxman’s own papers.
When to extend vs when not to
Section titled “When to extend vs when not to”| Extend | Don’t extend — use a contract flag instead |
|---|---|
Your data has a new syntactic form not yet recognized (e.g. YYYY.MM.DD for Date) | An existing form is off by default — turn it on (e.g. include_localized for Country, include_obfuscated for Email) |
| A different authority should validate an already-recognized form | Exclude or pin pinned_rules to the authority you want |
If a flag already exists for your case, prefer it — extension is for genuinely new shapes or specs.
The seam in one diagram
Section titled “The seam in one diagram”Minimal example — dot dates for the Date capability
Section titled “Minimal example — dot dates for the Date capability”Suppose your pipeline sees 2024.01.01 alongside the shipped ISO forms. The shipped Date capability does not include it, so it comes back MISSING. You add a grammar + rule and opt a contract in.
import refrom datetime import datetime
import paxmanfrom paxman.capabilities import Datefrom paxman.capabilities.Date.notation import DateNotationfrom paxman.core.contract import Contractfrom paxman.core.domain import Grammar, Provenance, RecognitionMatch, Rule, RuleStrategy
# 1. Register the shipped capability you are extending — before the first callpaxman.register_capability(Date())
# 2. Grammar — pure syntax, span-bearing, no spec judgmentclass DotDateGrammar(Grammar[DateNotation]): name = "dot_date_recognition" semantics = "dot_date_recognition" _PATTERN = re.compile(r"\b(\d{4})\.(\d{2})\.(\d{2})\b")
def recognize(self, text: str) -> list[RecognitionMatch[DateNotation]]: return [ RecognitionMatch( notation=DateNotation(N1=m.group(1), N2=m.group(2), N3=m.group(3)), start=m.start(), end=m.end(), raw_text=m.group(0), ) for m in self._PATTERN.finditer(text) ]
# 3. Rule — validation + canonicalization + provenanceclass DotDateRule(Rule[DateNotation]): name = "dot_date_rule" strategy = RuleStrategy.PARSER provenance = Provenance( authority="Example", specification_name="Dot-date convention (illustrative)", kind="convention", reference_url="", version=None, lifecycle="active", publication_year=2026, ) citation = "Illustrative only — not an ISO 8601 citation" target_semantics = frozenset({"dot_date_recognition"}) requires_features = frozenset() # no contract flag needed for this rule
def matches(self, notation: DateNotation, contract: Contract) -> bool: try: datetime(int(notation.N1), int(notation.N2), int(notation.N3)) return True except ValueError: return False
def normalize(self, notation: DateNotation, contract: Contract) -> str: return f"{int(notation.N1):04d}-{int(notation.N2):02d}-{int(notation.N3):02d}"
# 4. Register — also before the first callpaxman.register_grammar("date", DotDateGrammar)paxman.register_rule("date", DotDateRule)
# Note: the DotDateRule provenance above is illustrative — `YYYY.MM.DD` is an# application-specific convention, not an ISO 8601 form. Never attribute a# made-up format to a real authority; cite your own convention instead.
# 5. Opt in — every community grammar/rule is opt-in via the contractcontract = Date.create_contract(extra_grammars=("dot_date_recognition",))result = paxman.canonicalize("2024.01.01", contract)print(result.canonicalized_value) # "2024-01-01"That is the full lifecycle: register before the first call → opt in via extra_grammars. Without the opt-in, shipped behavior is unchanged; with it, the same canonicalize() seam returns the same ExecutionResult shape, now including your evidence.
Rules of the seam
Section titled “Rules of the seam”Register before the first call — the registries freeze
Section titled “Register before the first call — the registries freeze”import paxmanfrom paxman.capabilities import Date
paxman.register_capability(Date())paxman.register_grammar("date", DotDateGrammar)paxman.register_rule("date", DotDateRule)
# First call freezes capability + extension registries togethercontract = Date.create_contract(extra_grammars=("dot_date_recognition",))paxman.canonicalize("2024.01.01", contract)
# Anything after the freeze raises CapabilityErrorpaxman.register_grammar("date", AnotherGrammar) # errorPlace all register_* calls in your application startup or notebook setup cell, before any canonicalize().
Opt-in only — nothing runs unless named
Section titled “Opt-in only — nothing runs unless named”A community grammar runs only when named in contract.extra_grammars. A community rule runs only when the contract’s extra_grammars resolve to one of its target_semantics. This keeps shipped behavior deterministic per contract — an unaware contract never sees community logic, and an aware contract sees exactly what it opted into.
import paxmanfrom paxman.capabilities import Date
paxman.register_all_shipped()
# Dormant contract — dot dates not opted in, shipped behavior onlypaxman.canonicalize("2024-01-01", Date.create_contract()).canonicalized_value# "2024-01-01"
# Opted-in contract — dot dates participatepaxman.canonicalize( "2024.01.01", Date.create_contract(extra_grammars=("dot_date_recognition",))).canonicalized_value# "2024-01-01"Unknown names are silent for grammars, fail-fast for rules
Section titled “Unknown names are silent for grammars, fail-fast for rules”- An unknown grammar name in
extra_grammarsis silently skipped for grammar activation — the contract still runs identically (deterministically); the name is kept as the semantics key for rule activation. - A name that matches no grammar but does match a known semantics id still activates that semantics’s rules; a rule that was opted in via an id no grammar claims fails fast with
ContractError. A rule that was never opted in stays dormant regardless.
This errs on the side of predictable results over noisy warnings for grammar names, while keeping rule activation explicit.
Names must be unique
Section titled “Names must be unique”A community grammar whose name collides with a shipped grammar name for that capability fails fast with CapabilityError at composition time.
How target_semantics and requires_features interact
Section titled “How target_semantics and requires_features interact”target_semantics— which grammars’ notations this rule judges (non-emptyfrozenset[str]). A rule only meets the recognitions from those grammars.requires_features— which contract flags must beTruefor this rule to run (e.g.frozenset({"include_localized"})). A rule only runs in addition to theextra_grammarsopt-in when its required features are present andTrue; otherwise it is dropped and a recognized input that needed it becomesINVALID.
class LocalizedCountryRule(Rule[CountryNotation]): target_semantics = frozenset({"name_recognition"}) requires_features = frozenset( {"include_localized"} ) # only when contract.include_localized is TrueWhat grammars and rules must look like
Section titled “What grammars and rules must look like”The engine enforces metadata at import time — missing or mistyped fields raise TypeError immediately instead of producing wrong results later.
Grammar — subclass Grammar[NotationT]:
name: str— snake_case, unique per capability.semantics: str— non-empty; the meaning id the grammar’s notations carry. Use a new id for new meaning; reuse an existing capability’s semantics id when two grammars share meaning (they then coalesce).recognize(self, text: str) -> list[RecognitionMatch[NotationT]]— span-bearing[start, end)+raw_text; never bare notations.
Rule — subclass Rule[NotationT]:
name— e.g."Section 4.3.1-calendar-date", unique per capability.strategy—RuleStrategy.REGEX/LOOKUP_TABLE/PARSER. Qualification (ADR-0012): aPARSERcandidate survives only withLOOKUP_TABLEcorroboration on the same recognition — a communityPARSERrule on LOOKUP-backed semantics needs aLOOKUP_TABLEcompanion validating the same grammar, or its candidates are disqualified.provenance—Provenance(...)— the authority citation.citation— e.g."Section 4.3.1 (calendar date)".target_semantics: frozenset[str]— non-empty.requires_features: frozenset[str]— may be empty.matches(self, notation, contract) -> boolandnormalize(self, notation, contract) -> str.
Practical recipe — keeping it tidy
Section titled “Practical recipe — keeping it tidy”# app/startup.py — run once at startup, before any canonicalize()import paxmanfrom paxman.capabilities import Datefrom my_extension.dot_date import DotDateGrammar, DotDateRule
paxman.register_capability(Date())paxman.register_grammar("date", DotDateGrammar)paxman.register_rule("date", DotDateRule)
# app/processing.py — per call, opt in via contractfrom paxman.capabilities import Date
contract_shipped = Date.create_contract() # no dot datescontract_with_dot = Date.create_contract( extra_grammars=("dot_date_recognition",)) # with dot dates
# Notebook — run the startup cell once, then create contracts per cell as aboveFor multi-entity text, combine this with the Segmentation Recipe — segment first, then canonicalize() each piece with the opted-in contract.
See also
Section titled “See also”- Contracts —
extra_grammarson every contract - API Reference —
register_grammar/register_rule - Pipeline — where grammars and rules run
- Errors —
CapabilityError/ContractErrorfor bad names or late registration