The NOISE corpus was created through a series of iterative data-processing and curation workflows adapted to the characteristics of the underlying historical sources. Trade indexes, membership records, advertisements, blacklist cards, and diplomatic personnel registers differ considerably in structure, consistency, and visual complexity, requiring different approaches to extraction and processing. Comparatively regular material was processed primarily through parser- and dictionary-based workflows, while more heterogeneous or visually complex sources increasingly required AI-assisted extraction. Across all workflows, automated procedures remained embedded within human-supervised processes, with scholarly review retained throughout extraction, correction, normalisation, and grouping.
The advertisement corpus illustrates this combination particularly clearly. Large language models were used to identify relevant pages and extract structured information from advertisements, while computer-vision methods supported image segmentation and duplicate detection. The resulting records passed through repeated stages of grouping, correction, and validation. Automation substantially increased the amount of heterogeneous visual material that could be processed, while decisions concerning historical identity, equivalence, and ambiguity continued to depend on manual review and knowledge of the source context. Similar combinations of automated extraction and scholarly supervision were applied across the corpus according to the structure and complexity of the respective source material.

Extraction constituted only the first stage of data curation. Names, locations, classifications, and recurring actors subsequently had to be normalised and connected across documents, years, and datasets. Historical variation makes this particularly challenging: the same company or person may occur under different languages, abbreviations, spellings, or legal forms, while superficially similar names may refer to entirely different historical actors. Entity resolution consequently combines automated candidate generation and variation lists with iterative manual correction rather than relying on similarity-based matching alone.
Where appropriate, records are linked to established authority resources and classifications, including GeoNames, WZ 2008, ISIC Rev. 4, Wikidata, and the Gemeinsame Normdatei (GND). Documentary variants and unresolved ambiguities remain preserved rather than being silently replaced by a single standardised form. Source data, entity mappings, normalisation dictionaries, geographical authorities, and economic classifications are maintained separately and consolidated through reproducible rebuilds, allowing corrections and extensions to be propagated consistently across the corpus while preserving the connection between structured data and its documentary origins.