Skip to content

Capabilities — Overview

This content is for v0.4.0. Switch to the latest version for up-to-date documentation.

Each page below is a self-contained guide for one kind of identifier Paxman can canonicalize. Read the one that matches your data, copy the notebook snippet, and adapt the contract flags to your needs.

All capabilities share the same call shape — only the import, the factory, and the domain vocabulary change:

import paxman
from paxman.capabilities import X # Email, Country, ...
paxman.register_all_shipped() # once, before first use
contract = X.create_contract(...) # domain flags here
result = paxman.canonicalize(text, contract)

For the shared concepts behind these pages see Contracts, Pipeline, Execution Result, and the API Reference.


Your data looks like…Read
user@example.com, user at example dot comEmail
2026-01-15, 01/02/2026, 2026/01/15Date
US, United States, AlemaniaCountry
USD, $, euro, ¥ (identifiers without amounts)Currency
192.168.1.1, 2001:db8::1IP
9780306406157, 0306406152ISBN
USD 500, $500, 1.000,50 EUR (currency with amount)Money
+1 555 123 4567, (555) 234-5678, tel:+15551234567Phone
kg, m/s², megahertz, kPaSI Unit
https://example.com, http://münchen.deURL

The set above reflects the current release. New capabilities are added in minor releases — check paxman.capabilities or the latest release notes if you don’t see what you need.


Every capability page answers the same questions in the same order:

  1. What it canonicalizes and what it explicitly does not.
  2. Recognized forms — what patterns match, with what grammars.
  3. Canonical output & output_format — default and offered renderings.
  4. Contract flags — which knobs change recognition and validation.
  5. Statuses — concrete SUCCESS / MISSING / INVALID / AMBIGUOUS examples.
  6. Notebook snippet — runnable cleaning loop for a column.
  7. Provenance — which specifications vouch for the answer.

Start with the capability that matches your column; if you need more than one, register both and loop per cell (see the Segmentation Recipe for text that mixes kinds).


Paxman resolves one presumed entity per canonicalize() call (see Pipeline). Text that contains two different entities with different canonical values raises MultipleMentionsError rather than returning a merged answer — split first, then loop.

from paxman.core.errors import MultipleMentionsError
try:
result = paxman.canonicalize("alice@example.com, bob@example.org", contract)
except MultipleMentionsError:
# split the input and canonicalize each piece — see the segmentation recipe
...