About
What this is
A research tool for the printed calendars of medieval England's royal records — currently the Patent Rolls, the Close Rolls, the Fine Rolls, the Papal Registers and the Inquisitions Miscellaneous — from the reign of Henry III into the 15th century (1216–1486). It turns the calendared entries into a single chronological, searchable corpus and adds layers the printed calendars never had: every recorded act separated and classified, people and families tracked across entries, places mapped, and dates, themes and relationships made queryable.
The text is read by machine (OCR) from public-domain Internet Archive scans of the printed calendar volumes. Every entry carries a link to the scanned page it was read from, and a confidence score, so any word can be checked against the original. The printed General Indexes — the original editors' own name and place authority — are parsed and used as the standard against which spellings and identities are matched.
One particular aim is to recover the women of the rolls and follow their identities over time. A woman is often recorded under several names across her life: her father's surname, then her husband's (sometimes more than one). The site links those forms so she can be found under any of them and read as one life rather than several fragments.
Author
An independent digital humanities project by Creagh Factor. For suggestions, corrections, or collaboration, use the “Flag a data issue” form on any entry page or the contact address below.
Method
An offline, rule-based pipeline turns the calendar text into the corpus behind this site. The stages are deliberate and re-runnable:
- Reading. Public-domain volumes are read by machine (Tesseract OCR) from Internet Archive scans: the marginal date-and-place column is separated from the body by a per-volume gutter, every entry keeps its scanned-page link, and a date that cannot be read with confidence is left uncertain — never guessed.
- Dates. Each entry's date is normalized to a Julian ISO string, reconstructing from regnal year and feast day where needed, and flagged exact, inferred or uncertain.
- Places. Raw place strings are canonicalized against a gazetteer; counties follow the printed General Index where it gives them, and foreign places are treated as their own "county". Geocodes carry a provenance and a confidence, and low-confidence ones are flagged rather than hidden.
- Acts. A single calendar entry often records several discrete acts; each is segmented and classified by its opening formula (grant, pardon, protection, commission, and so on), with "the like" clauses inheriting the act before them.
- People, titles and kin. Personal names, offices and relationship clauses are extracted by pattern, then names that are spelling or Latin/vernacular variants of one person are clustered — with the printed General Indexes' own cross-references ("Bello Campo. See Beauchamp") as the deciding authority wherever the editors recorded one. Offices are recorded as "office of place" where stated.
- Women's identities. A married woman is linked under both her natal and marital surnames so she reads as one life across the rolls.
- Related enrolments. Entries that are repeats of one another are linked (see below).
- Search. Full-text search ranks by relevance; when there is no exact hit, fuzzy and phonetic indexes supply clearly-labelled close matches in medieval spelling.
Related enrolments
The rolls repeat themselves: the same instrument is often enrolled twice, re-enrolled on another roll, or re-issued verbatim later. Where two entries record the same matter, the entry page links them and labels the relationship a duplicate enrolment, a re-enrolled, reworded version, or a later re-issue.
The link is deterministic and shown with its evidence. Two entries are only compared if they share several rare words — the distinctive names and places that fingerprint an entry — so the thousands of near-identical but genuinely different formulaic acts (a protection for this man, a protection for that one) are never joined. Candidates are then confirmed by text similarity, and the kind follows from how similar they are and how far apart in time. Where the original editors themselves noted a duplicate ("Vacated because otherwise below"), the link is marked editor-noted; where they noted one but its twin is not in this corpus (it may sit in a roll not yet digitized), that is recorded as an open gap rather than hidden.
Everything extracted is auto-generated and unverified unless a profile is explicitly marked ✓ verified; the entry text itself is always the digitized calendar, shown verbatim.
Use of AI
AI-assisted coding tools were used in building this site's software. No large language model touches the data itself: every entry shown is the digitized calendar text, and all extraction and linking (dates, places, people, relationships) is done by deterministic, rule-based code that can be inspected and re-run. Nothing in the corpus is AI-generated or AI-paraphrased.
Relationship to the archives
This is an independent, non-commercial academic project. It is not affiliated with, or endorsed by, the Internet Archive or The National Archives. The printed calendars remain the canonical scholarly editions and the original rolls at The National Archives the ultimate authority; every entry links to the public-domain scan it was read from.
Contact
Questions, corrections, collaboration: creaghfactor gmail com, or use the “Flag a data issue” form on any entry page.
Data quality & statistics
Corpus built 2026-06-15T09:53:23.738519+00:00 · 326,005 entries · 1216-11-15 to 1486-03-10. These figures are regenerated with every rebuild, so they always describe the data currently being served.
| pipeline version | 0.1.0 |
|---|---|
| build date | 2026-06-15T09:53:23.738519+00:00 |
| source file count | 4496 |
| files parsed | 4496 |
| files failed | 0 |
| failed files | [] |
| entry count | 108564 |
| entries per monarch | {'edw2': 24268, 'edw3': 44735, 'edw1': 22781, 'hen3': 16780} |
| entries per roll type | {'patent': 72956, 'close': 35608} |
| entries per volume | {'patent/edw2_vol1': 4909, 'patent/edw2_vol2': 4229, 'patent/edw2_vol3': 4001, 'patent/edw2_vol4': 2927, 'patent/edw2_vol5': 2593, 'patent/edw3_vol1': 5371, 'patent/edw3_vol2': 5146, 'patent/edw3_vol3': 4438, 'patent/edw3_vol4': 3918, 'patent/edw3_vol5': 3208, 'patent/edw1_vol1': 2363, 'patent/edw1_vol2': 3800, 'patent/edw1_vol3': 6366, 'patent/edw1_vol4': 5739, 'patent/hen3_vol1': 218, 'patent/hen3_vol2': 199, 'patent/hen3_vol3': 3542, 'patent/hen3_vol4': 3794, 'patent/hen3_vol5': 3045, 'patent/hen3_vol6': 3150, 'close/edw1_vol1': 949, 'close/edw1_vol2': 564, 'close/edw1_vol3': 849, 'close/edw1_vol4': 1289, 'close/edw1_vol5': 862, 'close/edw2_vol1': 934, 'close/edw2_vol2': 1504, 'close/edw2_vol3': 1341, 'close/edw2_vol4': 1830, 'close/edw3_vol1': 2192, 'close/edw3_vol10': 1178, 'close/edw3_vol11': 1238, 'close/edw3_vol12': 838, 'close/edw3_vol13': 940, 'close/edw3_vol14': 996, 'close/edw3_vol2': 1756, 'close/edw3_vol3': 1984, 'close/edw3_vol4': 2300, 'close/edw3_vol5': 2148, 'close/edw3_vol6': 2315, 'close/edw3_vol7': 1549, 'close/edw3_vol8': 2110, 'close/edw3_vol9': 1110, 'close/hen3_vol1': 234, 'close/hen3_vol10': 148, 'close/hen3_vol11': 146, 'close/hen3_vol12': 158, 'close/hen3_vol13': 167, 'close/hen3_vol14': 162, 'close/hen3_vol15': 0, 'close/hen3_vol2': 207, 'close/hen3_vol3': 524, 'close/hen3_vol4': 237, 'close/hen3_vol5': 230, 'close/hen3_vol6': 222, 'close/hen3_vol7': 161, 'close/hen3_vol8': 113, 'close/hen3_vol9': 123} |
| date confidence | {'exact': 76967, 'inferred': 27854, 'uncertain': 3743} |
| pct entries with date | 96.9 |
| locations total | 1750 |
| locations geocoded | 407 |
| pct entries geocoded | 74.44 |
| dedupe dropped | 40413 |
| cross volume hash collisions | 25 |
| reign bound violations | 376 |
| invalid dates | 35 |
| feast day from raw text | 0 |
| feast day from date match | 27821 |
| day resolved from feast | 0 |
| persons total | 150000 |
| persons forced includes | 23 |
| problems | {} |
Known limitations
- Dates are Julian-calendar ISO strings as dated in the sources; many are inferred from regnal years and feast days rather than stated outright.
- Coverage is uneven and layered: volumes were digitized as public-domain scans became available, so some reigns are densely covered and others not yet. Machine-read entries carry a confidence score and a scanned-page image; see Coverage by reign for exactly where each series stops.
- Membrane references are present for about three-quarters of entries, sparse where the source lacked them.
- Place geocoding is incomplete; low-confidence geocodes are flagged, not hidden.
- People and relationships are auto-extracted; only profiles marked ✓ verified have been checked by hand.
- Near-duplicate entries from overlapping page ranges may persist in places.