#pagexml
Anyone can download their model data from Transkribus as PageXML and the images. I did it and uploaded into our instance of eScriptorium and trained the model. Both models are equally good and equally bad. We can’t share images because copyright issues. But it’s researcher’s decision.
March 16, 2026 at 8:23 PM
Hi :) I want to know how we can help people who want to publish their ground truth data to formalize them into reusable packages (like catalogued data on HTR United): that means having the ability to read informations about projects, images, PageXML from outside the platform.
March 16, 2026 at 2:59 PM
19/24 of our DH Vibe Coding Calendar. re approaching the finish line! Today: Ever wanted to replace some characters in PageXML? Try our vibe-coded PageXML Transformer! It might even work! #atr #htr #vibecoding #dh advent-calendar.humanities.tools
December 19, 2025 at 11:18 AM
16/24 of our DH Vibe Coding Advent Calendar :). Today LayoutXML-to-trOCR tool created by @burnedlegate.bsky.social to easily convert PageXML and ALTO into training datasets for trOCR models.

advent-calendar.humanities.tools #htr #atr #ocr #trOCR #machinelearning #ai
December 16, 2025 at 12:59 PM
I am so frustrated. Searching in the traditional way just don't work anymore. Enter something like "editing Pagexml" and you will get some random stuff from the unwanted AI-dummy and nothing to pagexml on the upper part of the list of results. So I literally went to places like github and pip.
December 14, 2025 at 11:21 AM
12/12 of our DH Vibe Coding Calendar: convert pagexml to an edition! :) Instantly! Made by @martin.rocek.dev .

#dh #atr #htr #pagexml #humanities
December 12, 2025 at 12:50 PM
Interesting question. If using transkribus I'd probably do correction there, then again in TEI. I don't edit anything directly in PageXML personally.
September 21, 2025 at 10:17 AM
would you recommend exporting to PageXML (would have to convert HCOR bc Internet Archive doesn't have that format) *before* doing OCR post-correction using PageXML format? or after?
September 21, 2025 at 10:12 AM
@triproftri.bsky.social

The better conversion to TEI is to export as PageXML and convert to TEI using github.com/dariok/page2... as this enables more choice in parameters and customisation.

The #TEI community is here (and mailing list and slack, etc) for TEI assistance. @teiconsortium.bsky.social
GitHub - dariok/page2tei
Contribute to dariok/page2tei development by creating an account on GitHub.
github.com
September 21, 2025 at 10:01 AM
Open Archieven toont scans en transcripties in een nieuwe viewer op basis van een IIIF manifests met annotaties, zie bijv. www.openarchieven.nl/transcriptie...

#IIIF #Tify #PageXML
Transcriptie NL-HaNA_1.04.02_8275_0120
Open Archieven maakt transcripties van scans van historische documenten doorzoekbaar en biedt de mogelijkheid er een samenvatting van te maken in hedendaagse taal.
www.openarchieven.nl
June 9, 2025 at 8:58 PM
<a href="https://bsky.app/profile/did:plc:zs2rsvhhqmqrmmiq3dhvfbks" class="hover:underline text-blue-600 dark:text-sky-400 no-card-link" target="_blank" rel="noopener" data-link="bsky-mention">@stefan_hessbrueggen wenn's auf PageXML-Basis ist, hätte ich da auch was für Dich ;-)
February 20, 2025 at 1:24 PM
🤔 I think only ALTO is supported for now...

The adaptation to PageXML might be relatively simple to make (and I might need to work on it myself in the coming weeks)
February 9, 2024 at 11:12 AM