Reconstructing Peirce's Trails of Thought: Multimodal Computational Analysis of Non-Linear Manuscript CompositionThe digitization of manuscript collections opens new possibilities for analyzing complex, multimodal documents.

Abstract

The digitization of manuscript collections opens new possibilities for analyzing complex, multimodal documents. Charles S. Peirce's manuscripts at Harvard's Houghton Library are a paradigmatic case, combining prose, logical notation, and diagrams, with frequent rewriting and backtracking across hundreds of pages. In the 1980s a manual reconstruction mapped the compositional topology of one such manuscript, later named an S-diagram. This study investigates whether computational methodologies can recover this compositional structure and if a reconstruction accounting for both textual and visual content confirms and extends the original manual effort. Our pipeline applies lexical and semantic similarity analyses to the text, combining them with a visual classification of each page produced by Vision-Language Models (VLMs). Integrating these approaches is essential to capture what we term semiotic continuity, which refers to the seamless integration of prose, notation, and diagrams within the same page. We validate the method against a formalization of the manual reconstruction, archival sheet variants, cross-manuscript links, and the convergence of the two similarity methods. The workflow reproduces the manual mapping in its annotated zone and extends it into regions the original did not cover. Ultimately, this framework can be applied to any manuscript collection characterized by extensive rewriting and multimodal composition.

1 Introduction

In 1903, Charles Sanders Peirce1 observed: “All that you can find in print of my work on logic are simply scattered outcroppings here and there of a rich vein which remains unpublished. Most of it I suppose has been written down; but no human being could ever put together the fragments. I could not myself do so.” (MS 302). More than a century later, this assessment remains largely accurate. The Charles S. Peirce Papers (MS Am 1632), housed at Harvard's Houghton Library, comprise approximately 100,000 manuscript pages, of which a subset of 233 items (15,695 facsimile images) has been digitized and made available through the International Image Interoperability Framework (IIIF) [11]. Richard Robin's Annotated Catalogue (1967) organizes the collection into twelve thematic categories, but this structure does not reflect the chronological development of Peirce's thought [15].

The challenge these manuscripts pose for scholarly access goes beyond volume. Christian Kloesel, who edited the Writings of Charles Sanders Peirce for two decades, describes a compositional method in which Peirce would advance along one line of thought until reaching an impasse, then double back any number of pages to a logical forking point and proceed in a different direction [8]. As a consequence, the conventional distinction between draft and final version becomes inapplicable to such a corpus [6]. Moreover, given the author's commitment to visual reasoning, the iconic richness of the manuscripts remains largely inaccessible in many existing printed editions [17, 18].

This compositional habit has practical consequences for scholarship insofar as topics introduced in one manuscript reappear in others whose nominal subject is unrelated, and manuscript labels are rarely a reliable guide to content. The standard reference edition, the Collected Papers [12], presents several limitations given its topical arrangement which often places parts of the same manuscript in different volumes, and occasionally splices together fragments composed in different years [6, 8].

In the 1980s, Shea Zellweger physically examined and reordered the approximately 900 pages constituting Peirce's “The Simplest Mathematics” (MSS 429, 430, 431) at the Peirce Edition Project. His hand-drawn topological diagram of MS 431 is the only known attempt to map the compositional structure of a Peirce manuscript. It is organized around a vertical spine of sheet numbers in compositional order, with closed ovals marking zones of rewriting, circled numbers marking dated writing sessions, and lateral annotations tracking cross-references to other manuscripts. Keeler later proposed that scholars should be enabled to create such Shea diagrams (S-diagrams) computationally and link them to digitized manuscript pages, so that alternative reconstructions can be compared and disagreements localized [7].

In this paper, we present a computational method for generating and evaluating S-diagrams, building on Harvard's IIIF digitization of MS 431 and the 1963–64 microfilm edition (Reel 10). We apply two complementary similarity analyses, lexical and semantic, and combine them with visual page classification via a Vision Language Model (VLM). The analysis reveals a structural property we term semiotic continuity, the irreducible entanglement of prose, symbolic notation, and diagrammatic content on the same page, which makes textual approaches insufficient. Finally, the validation is assessed through Zellweger's hand-drawn diagram, Harvard's archival variant annotations, cross-manuscript similarity links, and the structural convergence of the two similarity methods.

2 Background

2.1 Computational Approaches to Peirce's Manuscripts

Existing computational work on Peirce's manuscripts has so far addressed the textual [13] and the visual [11] dimensions of the corpus separately. The textual line investigates the linguistic content of single manuscripts through a TF-IDF-based study of Prolegomena to an Apology for Pragmaticism [13], arguing that even within a single manuscript Peirce develops parallel argumentative threads detectable through lexical patterns. The visual line addresses the iconic and diagrammatic content of the manuscripts through a workflow combining layout segmentation, IIIF annotation, and VLMs applied to Peirce's diagrammatic pages to evaluate whether such models can engage with diagrams as operative semiotic forms [11].

Building on these premises, we use the term semiotic continuity to name this property of the corpus, drawing on Peirce's own doctrine of synechism, the thesis that continuity is the fundamental category of reality, including the reality of signs. In the manuscripts this principle materializes as a refusal to separate prose from iconic elements and diagrammatic demonstration. The direct methodological consequence of employing this concept for analyzing Peirce's manuscripts is that detection of compositional structure cannot rely on textual similarity alone, because two pages may be lexically distant but structurally related. Similarly, visual classification alone cannot distinguish a page that introduces a new diagram from one that revises an earlier one. The dual-method design adopted in this paper follows from this property rather than being imposed on the corpus from outside.

3 Method

3.1 Corpus

We work on three manuscripts from Peirce's 1902 The Simplest Mathematics, originally intended as Chapter III of the projected Minute Logic: MS 429 (typescript); MS 430, an autograph draft; and MS 431, the autograph from which Zellweger derived his S-diagram. The IIIF digitization of MS 431 stops at sheet 100, while Zellweger's diagram extends to sheet 200. The missing pages are available on microfilms produced in 1963–64 (frames 560–961).

We produced a diplomatic transcription of the full corpus through Google Gemini 3 Pro [4], encoding the output in TEI. Peirce frequently reused the same sheet number for multiple drafts of the same page, so deduplicating by label is unreliable. Instead, for each microfilm page sharing a sheet label with an IIIF page, we measured textual similarity, treating pairs above 0.85 as duplicates (keeping the IIIF version) and pairs at or below 0.85 as compositional variants (keeping both). This produced a unified corpus of 590 records ordered by sheet number, of which 484 contain text. The corpus spans sheets 2–198 with one gap at sheet 111.

3.2 Similarity and Detection

The choice of two textual measures also follows best practices in document retrieval, where combining lexical representations with semantic embeddings outputs more robust similarity [2, 9]. We use scikit-learn's TF-IDF for the sparse representation and all-MiniLM-L6-v2 [14] for sentence embeddings. The integration of visual and textual modalities has produced computational tools applicable to digitized cultural heritage collections, with tasks ranging from iconographic captioning [1] to Visual Question Answering and multimodal retrieval [3]. Google Gemini 3 Pro [4] has been chosen to perform the visual classification both for its price and performance [5, 16].

Following this approach, we analyze the corpus along two axes, textual and visual, and combine them:

    The visual axis is a classification of each page into one of four classes based on their content: blank, prose, notation, or mixed (e.g. prose and notation or diagrams on the same page). The VLM performs the classification on the page image alone. The output measures semiotic continuity at the page level and qualifies the textual events described below.\

    The textual axis is built on two similarity measures. The first compares pages by word overlap, treating shared vocabulary as a sign of compositional reuse. The second compares pages semantically, identifying pairs that address the same content even when phrased differently.\

From each similarity measure we define and extract a taxonomy of five compositional events: (i) a discontinuity is a sharp drop in similarity between two consecutive pages, marking a textual break; (ii) a backtrack is a high similarity between two pages that are not consecutive, indicating a return to earlier material; (iii) a hub is a page that receives repeated backtracks from distant locations, identifying a passage the author kept rewriting from later in the manuscript; (iv) a fork point is a page that combines a discontinuity ahead with a backtrack to an earlier page, marking the start of a new line of thought built on previous material; finally, (v) a hot zone is a window of consecutive pages with a high local density of discontinuities and backtracks, identifying regions of rewriting. The cutoffs that separate ”low” from ”high” similarity are derived from the data using percentiles instead of the standard mean-and-standard-deviation rule, which fails on the bimodal similarity distribution of multi-draft corpora.

A discontinuity is then qualified using the visual classification. If both sides belong to the same modality, we define it as a compositional fork. Conversely, if a transition between prose and visual content is registered, we call it a modality switch.

3.3 Validation

Finally, we validate the pipeline against four reference points.

    The first is Zellweger's S-diagram of MS 431. We transcribed the ground truth into a JSON representation, extracting 20 backtrack edges encoded in his annotations (e.g., 82(45) yields the edge 82 → 45).\

    The second is Harvard's archival annotation of compositional variants. The Houghton Library identifies 24 sheets in MS 431 as physical variants, marked with letter suffixes in the catalog. This annotation is independent of textual content. We test whether these sheets fall within our detected hot regions.\

    The third is cross-manuscript similarity. MSS 429, 430, and 431 are three stages of the same chapter, so pages discussing the same content occur across them. We compute textual similarity between MS 431 and each of MSS 429 and 430 with both methods, retaining the top fifty pairs in each case.\

    Fourth, the two similarity measures capture different aspects of textual proximity and could in principle disagree on every detection. Their overlap on the same compositional events therefore signals that the detected structure is real.\

4 Results

First, the visual classification confirms our hypothesis of semiotic continuity at the page level. Of the 484 non-blank pages in the corpus, 64% combine prose and notation or other kinds of visual content on the same page, 34% are prose, while 2% are notation.

Second, the textual analysis identifies the five compositional events defined in Section 3, computed independently for each similarity measure. Each method detects 71 discontinuities, 34 of which are shared (22 distinct sheets at sheet-level aggregation). Counts of fork points are comparable, with substantial overlap and hot zones converge across methods, as reported in Table 1.

Table 1: Convergence between lexical and semantic similarity methods.

Detected events

Shared

Discontinuities (71 each)

34 / 71

Hot zones (10 each)

9 / 10 with IoU > 0.5

Archival variants (24 sheets)

24 / 24 within hot regions

Sixteen of the 34 shared discontinuities fall within the lower zone Zellweger labeled “ABOUT 30 PAGES” (sheets 38–44). Zellweger established by physical inspection of the sheets that this cluster is a zone of repeated rewriting, and our pipeline recovers it without supervision. A further 9 fall within the “Trichotomic Mathematics” zone (sheets 155–171), in the half of the manuscript his diagram sketches but does not annotate in detail, and are to our knowledge not previously recorded.

Figure 1: Spine map of MS 431. Ticks mark detected discontinuities: those crossing the spine were found by both methods, those above by the lexical method only, those below by the semantic method only. Green bands are hot zones, and the arches join Zellweger's annotated rewrites to their source pages.

Horizontal spine map showing detected discontinuities, hot zones, and Zellweger ground-truth backtrack edges along MS 431.

We finally compared the detected backtracks with the 20 edges encoded from Zellweger's diagram. The lexical method recovers 8 out of 20 edges and the semantic method recovers 12, bringing their union to an 80% success rate. The 5 edges missed by both approaches involve sheet 92, a page that synthesizes a nine-page range (sheets 57–65) and consequently disperses the textual signal across multiple sources. The cross-manuscript analysis also recovers links between MS 431 and MS 429 in the upper zone, mapping sheets from the “Trichotomic Mathematics” section onto sheets 221–247 of MS 429 and sharing 24 of the top 50 pairs between the two methods.

The corpus, similarity matrices, detection outputs, and ground-truth transcription are available on Zenodo [10].

5 Discussion

The method recovers most of Zellweger's manual annotations and extends his mapping into the upper zone of the diagram. The complementarity of the two similarity measures is a finding in itself, as the lexical approach detects reused vocabulary during revisions, the semantic method picks up reformulated content. Relying on just one of these techniques would miss part of the data, while the connections missed by both reveal a structural limit of textual similarity on multi-source synthesis.

A digital edition tends to treat the combination of prose, notation, and diagrams as a layout feature, but our results indicate that this combination constitutes the argument itself. A single-modality method would misinterpret transitions from prose to notation, where the text signals a break that the visual layer instead recognizes as continuity. The convergence of textual and visual signals at the page level can thus be read as an empirical confirmation of Peirce's theory of continuous semiosis, where prose, notation, and diagram operate in continuous exchange instead of being separated content layers (both conceptually and materially).

The visual classification relies on a proprietary model whose decisions are not inspectable. The same step could be performed with open-weight models, at an accuracy cost that remains to be measured. Despite these and other limitations, in particular the validation resting on a single manuscript and a single ground-truth diagram, and the reduction of page content to four discrete visual classes, the pipeline is not limited to Peirce and applies to any manuscript collection characterized by rewriting and multimodal composition. Future work could employ multimodal embeddings that combine text and image within a single similarity space, enabling finer-grained detection of cross-modal events. The detected structure can also be exposed as a knowledge graph linked to IIIF facsimiles, supporting scholarly queries and integration with retrieval-augmented systems. More importantly, the workflow can be embedded in digital scholarly editions as a semi-automatic tool: scholars and domain experts could draft, refine, and annotate their own S-diagrams as hermeneutic interpretations of compositional structure, treating annotation as a form of structured, visual argument.

6 Conclusion

This preliminary work presents a computational pipeline for reconstructing the compositional structure of non-linear manuscripts, combining lexical and semantic similarity analysis with visual page classification within a unified corpus. The workflow operationalizes the notion of semiotic continuity, treating prose, notation, and diagrams as co-occurring signs on the same page. Applied to Peirce's MS 431, the analysis confirms semiotic continuity empirically (64% of text-bearing pages combine modalities) and recovers 80% of Zellweger's hand-drawn rewrite annotations through the union of the two similarity measures. The pipeline further extends the manual reconstruction into the upper half of MS 431, accessible only through the microfilm and not annotated in detail by Zellweger, where similar compositional density emerges.

The methodological pattern extends beyond Peirce studies. Any non-linear manuscript collection with multimodal composition can adapt this workflow.

Notes

1Charles Sanders Peirce (1839–1914) was a logician, mathematician, and philosopher, founder of pragmatism and of modern semiotics. He held no permanent academic position after 1884 and published no philosophical book in his lifetime, leaving most of his work in manuscript. The bulk of it reached print only decades after his death, and much of it remains unpublished.

Source

Imported from ACM’s structured HTML source. ACM Reference Format: Carlo Teo Pedretti. 2026. Reconstructing Peirce's Trails of Thought: Multimodal Computational Analysis of Non-Linear Manuscript Composition. In 37th ACM Conference on Hypertext (HT '26), September 14--18, 2026, London, United Kingdom. ACM, New York, NY, USA 4 Pages. https://doi.org/10.1145/3800935.3830878

References

[1] Eva Cetinic. 2021. Towards Generating and Evaluating Iconographic Image Captions of Artworks. Journal of Imaging 7, 8 (2021), 123. https://doi.org/10.3390/jimaging7080123

[2] Cedric De Boom, Steven Van Canneyt, Steven Bohez, Thomas Demeester, and Bart Dhoedt. 2015. Learning Semantic Similarity for Very Short Texts. arxiv:1512.00765

[3] Noa Garcia and George Vogiatzis. 2019. How to read paintings: Semantic art understanding with multi-modal retrieval. In Computer Vision – ECCV 2018 Workshops, Proceedings(Lecture Notes in Computer Science, Vol. 11130), Stefan Roth and Laura Leal-Taixé (Eds.). Springer, Berlin, Heidelberg, 676–691. https://doi.org/10.1007/978-3-030-11012-3_52

[4] Gemini Team. 2025. Gemini 3 Pro Model Card. Retrieved April 27, 2026 from https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf

[5] Yifan Hou, Buse Giledereli, Yilei Tu, and Mrinmaya Sachan. 2024. Do Vision-Language Models Really Understand Visual Language?arxiv:2410.00193 https://doi.org/10.48550/ARXIV.2410.00193

[6] Mary Keeler. 2020. The Hidden Treasure of C. S. Peirce's Manuscripts. Chinese Semiotic Studies 16, 1 (2020), 155–166. https://doi.org/10.1515/css-2020-0008

[7] Mary Keeler. 2020. Pragmatically Improving Access to Peirce's Archive. Chinese Semiotic Studies 16, 1 (2020), 167–187. https://doi.org/10.1515/css-2020-0009

[8] Mary Keeler and Christian Kloesel. 1997. Communication, Semiotic Continuity, and the Margins of the Peircean Text. In Conceptual Structures: Fulfilling Peirce's Dream, Dickson Lukose, Harry Delugach, Mary Keeler, Leroy Searle, and John F. Sowa (Eds.). Springer, Berlin, 275–289. https://doi.org/10.1007/BFb0027879

[9] Priyanka Mandikal and Raymond Mooney. 2024. Sparse Meets Dense: A Hybrid Approach to Enhance Scientific Document Retrieval. In The 4th Workshop on Scientific Document Understanding (SDU) at AAAI 2024. CEUR Workshop Proceedings, Vancouver, BC, Canada, 1–9. arxiv:2401.04055 https://doi.org/10.48550/arXiv.2401.04055

[10] Carlo Teo Pedretti. 2026. Computational Reconstruction of S-Diagrams in Peirce's Manuscripts: Supplementary Materials. https://doi.org/10.5281/zenodo.19848865

[11] Carlo Teo Pedretti, Davide Picca, and Dario Rodighiero. 2025. Moving Pictures of Thought: Extracting Visual Knowledge in Charles S. Peirce's Manuscripts with Vision-Language Models, In Computational Humanities Research 2025 (CHR 2025). Anthology of Computers and the Humanities 3, 1454–1467. https://doi.org/10.63744/fkFGJ6wSzDPV

[12] Charles S. Peirce. 1931. Collected Papers of Charles Sanders Peirce, Vols. 1–6. Harvard University Press, Cambridge, MA. Edited by Charles Hartshorne and Paul Weiss.

[13] Davide Picca, Antonin Schnyder, Eri Kostina, Alessandro Adamou, Dario Rodighiero, and Jeffrey T. Schnapp. 2023. Orchestrating Cultural Heritage: Exploring the Automated Analysis and Organization of Charles S. Peirce's PAP Manuscript. In Proceedings of the 34th ACM Conference on Hypertext and Social Media (Rome, Italy) (HT ’23). Association for Computing Machinery, New York, NY, USA, 1–4. https://doi.org/10.1145/3603163.3609066

[14] Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (Eds.). Association for Computational Linguistics, Hong Kong, China, 3982–3992. https://doi.org/10.18653/v1/D19-1410

[15] Richard S. Robin. 1971. The Peirce Papers: A Supplementary Catalogue. Transactions of the Charles S. Peirce Society 7, 1 (1971), 37–57.

[16] Saurav Sengupta, Nazanin Moradinasab, Jiebei Liu, and Donald E. Brown. 2025. Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features. arxiv:2509.08266

[17] Frederik Stjernfelt. 2019. Dimensions of Peircean diagrammaticality. Semiotica 2019, 228 (2019), 301–331.

[18] Frederik Stjernfelt. 2022. Sheets, Diagrams, and Realism in Peirce. De Gruyter, Berlin, Boston.

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime