Abstract
Recent analysis and critique of commentaries have called for clearer, more structurally robust reinforcement of claims made in commentaries. We introduce a hypertext reading environment for an Ancient Greek source text and an associated commentary, where its claims are supported by transparent, accessible, and replicable data-driven analyses leveraging Jupyter notebook integration. This environment is then evaluated in a user study exploring whether Tompkinsia can reinforce critical perspective of commentaries.
1 Introduction
The commentary tradition has existed for millenia, augmenting reading and translation both on and off the screen. Commentaries are linear lists of annotations on a foreign language source text, and they help readers across expertise levels understand obscure linguistic, historical, and/or cultural context. Commentaries in the Latin and Ancient Greek traditions serve two purposes: first, as a well-informed companion text to help students become acclimated with the language they are learning, and second, as a tool for scholars to further study any topics or lines of inquiry which resonate with the source text [20]. Recent analysis and critique of commentaries has emphasized three specific problem areas: parallels’ failure to account for “bursty" lexical dispersion, vague and misleading claims, and that commentators often assume a more advanced audience, even when commentators nominally produce a commentary that attempts to serve both novice language learners and veterans equally [20, 34]. We propose Tompkinsia, a hypertext reading environment for an Ancient Greek source text and an associated commentary, where its claims are supported by transparent, accessible, and replicable data-driven analyses leveraging Jupyter notebook integration.
Developing Tompkinsia presented an opportunity to explore both how it can change perception of commentaries and commentators, as well as what it can do for critical thinking. To this end, we ran a user study with one early undergraduate level Greek class to better understand Tompkinsia's effect on their perspective of commentary claims, and to evaluate whether Tompkinsia could be used to help reinforce critical thinking skills while participating in a multi-document hypertext translation activity. Students browsed through a version of Tompkinsia featuring lines from the Iliad and a version of the accompanying Basel commentary featuring analysis of 3 claims in the work. Our hypotheses for this user study include:
H1. Participants will form plans that demonstrate understanding of methods reviewed in Tompkinsia.
H2. Exposure to commentary claim analysis will decrease participant acceptance of such claims, with more direct exposure to commentary claim analysis by route of Tompkinsia causing a greater decline in claim acceptance.
2 Background and Motivation
2.1 A History of Claims and Commentaries
Since commentaries are linear lists of annotations on a foreign language source text, we argue this construction approaches the definition of a hypertext. Many commentaries try to provide full clarity on the unfamiliar historical, cultural, and linguistic contexts these texts were written in [5]. Scholars wrote commentaries on texts in many languages for millenia [12]. The first serious editions and commentaries appeared in Alexandria in the 3rd century BCE, and fragments of this work exist on ancient papyri [13]. While a well-developed model for page-based commentaries emerged more than 1,000 years ago, the commentary adopted its modern form by the 18th century [24].
One commentary item type is a claim about the frequency of a phenomenon in the source text. Often the claim is supported by a set of “parallels," or various primary source references for that phenomenon. Commentaries and commentators differ in focus and habit, and the number and quality of parallels were thoroughly surveyed and questioned by commentary critics [14, 18]. Parallels may be a form of or alongside a similar phenomenon we call “coverage creep." When writing a non-initial commentary on a text, convention compels the commentator to echo relevant insights from these commentaries and provide new insights. However, what was once scholars in the 1800s referencing an important manuscript is now the duty of modern commentators to cover centuries of analysis. Plus, a drop in references per item over the last century could be ascribed to lowering the cost of accessing all references supporting a commentator's claim [5]. This decrease in cost helps lower the barrier of entry to engaging with commentaries, as the study of Greco-Roman literature's audience broadened past the upper class. Lastly, in the late 90s/early 2000s, with the advent of computerized literature and hypertext, commentaries’ critical analysts called for varying radical commentary redesigns in a new digital medium [16, 26, 34]. Many hypertext reading environments still represent commentaries where their most radical design change is hyperlinking reference sets, as seen in environments like [17]. Born-digital commentaries show what commentaries could be, yet many digitized commentaries are mostly unaugmented and unrevised.
2.2 Why Validate Commentary Claims?
There are still shortcomings in commentaries that may be clearer to scholars of CL and accessibility, educators, students, and other unconventional and/or underrepresented types of language learners. First, humanist academics write mostly for their peers. [5] found that authors often support claims with parallels without accounting for lexical dispersion and other factors that may invalidate the claim. Next, claims can be vague or misleading. Charles Smith's 1913 commentary on Thucydides’ History of the Peloponnesian War said the verb λωφάω was “used by Thuc. with reference to sicknesses and grave misfortunes" [33]. Searching for all instances of λωφάω in the History, 2 instances reference the Athenian plague, and another refers to other misfortunes Athenians overcame. Given the note's misleading phrasing, readers may think this verb is used more than 3 times in the History. These claim analyses must also be accessibly explained. [5] reports that when academics write commentaries aiming to aid both newer and more advanced language learners, the content often underserves the former. The aid newer language learners need is often absent since commentators assume novices do not read this, leading to their exclusion of not just the novices from understanding the commentary.
Tompkinsia will also aid both hypertext completionists and those uninterested in these notebooks’ analyses. Treebanks have existed for decades, yet while many CL researchers analyzed many treebanks with their own bespoke workflows, such analyses are not often front-facing or novice friendly. This work could lower the barrier of entry into this kind of analysis to those familiar with Jupyter notebooks. Augmenting these commentaries with claim analysis may also keep the cursory reader from being misled by claims with notes summarizing their inconsistency or veracity.
Lastly, this work's scope matters regarding claim validation. Like the analytical work of Dan Tompkins [35], this design's namesake, the framework centers data-driven CL analysis of unambiguous claims (i.e., “Sentences ended with a genitive absolute are rare"). A claim like “Odysseus uses the gnomic aorist when in a rage" is more ambiguous and harder to find with treebanks. Finding aorists in Odysseus’ dialogue may be easy, yet checking results may be time consuming with no notation for what aorists are gnomic or what emotions are there. One could also use a keylist of anger words to further filter the results. Making an accurate keylist with rarer words takes time, but a valid instance could always be excluded. A third method may be training a BERT model to identify Odysseus’ anger, and evaluating its results. Tompkinsia is a presentation medium for checking claims made on a reliable, searchable, and mathematical foundation, and this design centers showing what can be found and analyzed in an unambiguous way.
3 Related Work
Perseus Digital Library [10] digitized about 70 print commentaries, hyperlinking print citations and turning these commentaries into scholarly hypertexts where readers could traverse commentary and source text to see, clearly seeing connections between texts [9]. Other digital platforms include the Dickinson College Commentaries [17], known for extending the traditional commentary model by including media files and linking primary sources and reference works, and more recently New Alexandria [15] built a commentary system where each annotation was peer-reviewed as a separate publication with its own author(s) and metadata. New Alexandria enabled the move beyond the traditional model of single-authored commentaries perpetuating coverage creep [11]. The Ajax Multi-Commentary Project [29] also notably addressed coverage creep [30]. In addition, the Ajax Multi-Commentary Project pioneered work on scalable integration of digitized commentaries and building sustainably minimal computing complex commentary networks [31]. Commentary augmentation is not new, but we have not seen a hypertext reading environment integrating dedicated Jupyter notebooks for analyzing these commentaries’ claims.
[6] is another recent user study identifying different patterns of reading behavior while featuring another type of augmented commentary. Students read the Parodos in Sophocles’ Antigone while accompanied by an augmented version of Richard Jebb's commentary on the text where references were replaced with the relevant excerpt of text referenced. This commentary augmentation design is influenced by findings regarding beginner-friendly ways to integrate augmentations without overloading the reader. The original GLAUx treebanks appeared around 2021 as a result of training multiple classifiers to categorize tokens by Greek morphology, lemmata, part-of-speech, syntactical relation and semantics [23]. It is one of the largest corpora of treebanks spanning many Greek works and is open source, as opposed to the more private TLG [23]. GLAUx also has its own web-based performant treebank query engine, though scholarship on this is currently under development [1, 2]. Web-hosted treebank query engines are not new [22, 25], but they are separate, independent websites. This work seeks to integrate analysis more directly into the experience of reading a foreign language source text with a commentary in a hypertext reading environment.
A few works on digital reading influence this work's perception of potential readers. [32]’s work on content organizers demonstrates that summaries of hypertext partitions can result in more readers being dissuaded from perusing the same partition of hypertext. One can view a commentary claim similarly: readers may view it as an authoritative statement that is factually accurate and does not require verification. [21] and [27] identify hypertext reader groups separated by coverage, or what percent of the total hypertext the user traversed. [21] mentions a “disengaged" group and a more completionist group. This is also resonant with types of commentary readers, especially since references and other language may dissuade language learners from engaging deeper with the material itself [20].
[5] proposed and applied a non-exclusive taxonomy for types of aid provided by commentaries across many commentary types to reveal both differences in aid distribution and references provided, as well as the cost of and other barriers to accessing the full insight of commentaries. This work demonstrates one way to include code and results supporting claims in commentaries. Multiple critiques over the last 20 years have targeted various aspects of commentary writing, reading, and design [19, 20, 34], yet while these various pieces call for born-digital and digitized commentaries to reckon with these issues, few born-digital commentaries answer this call. Hopefully this work will not be the only response to these pieces.
4 Methodology
4.1 Participants
There were n = 12 participants. 10 were undergraduates in a second-semester class reading ancient Greek literature, and 2 were experts who teach CL methods and are familiar with ancient Greek. All were either students or instructors at Tufts University. 1 participant was nonbinary, 5 were women, and 6 were men. Informed consent was obtained for all. Participants were alternatingly assigned to one of 2 experimental groups. Due to technical difficulties, we were unable to record the screen of 1 expert as they browsed.
4.2 Materials
4.2.1 Selecting A Hypertext. We chose lines 37-65 of Book 6 of Homer's Iliad as our source text, with excerpts from the Basel Commentary volume by Magdalene Stoevesandt. To choose between the Basel Commentary and Walter Leaf's 1886 Iliad commentary, a focus group of masters students in digital humanities was convened. Based on discussion with the focus group and the diversity of claims presented in the works, the Basel Commentary was chosen. To be familiar with the text, students were told to read the lines on their own before the user study. This study also features some of the History with two claims from Smith's commentary. The History is a text students may have greater difficulty with translating, but we included the text and claims so that participants did not need to translate any Thucydides, and emphasis was put on clearly explaining how the claims posed a challenge to the reader. That said, a common reason focus group members chose the Basel Commentary was because it was contemporary. Another reason they chose the Basel Commentary was its attention to the source text as a work of poetry. There are notes on the frequency of formulae and epithets, but there is still no universal consensus on frequency's relationship to the Homeric Formula [28]. Participants had up to 10 minutes of browsing time in Tompkinsia: enough time to read and re-read as needed in a time-bounded task setting.
4.2.2 Commentary Augmentation. We needed a portion of the Basel Commentary with some treebank-verifiable claims we could present analyses for while demanding just enough translation and review of the Iliad from participants. [5] already identified claims in this commentary; we chose from those, each roughly 20-30 lines presenting different premises that were easy to analyze and follow at any expertise level. Lines 37-65 did not start or stop abruptly, and corresponding notes included claims regarding a bow epithet, a two-word phrase, and a verse end or VE formula, which we analyzed in Tompkinsia. The first claim proposed that ἀγκύλοv, meaning curved or crooked, only modified the weapon word ἅρμα here and usually modified the bow word τόξov as a bow epithet [7]. The second proposed that σὺ δέ occurred mostly in Homer and Herodotus [7]. The last claimed that ἔπος ηὔδα was a VE formula always preceded by a participle [7]. Claims were translated as statements evaluatable by querying for and verifying instances of a phenomenon in the treebanks. Statements were then evaluated by their respective metrics, and these results were evaluated as a group to determine whether the base claim was valid. This completed base analysis was reviewed and adapted into a Python notebook where the process was explained step-by-step in an accessible way, with code users could run and tweak. These notebooks are available at [3].
We also removed items depending on references or a greater assumed knowledge of ancient Greek. Because our newer language learners had not read the Iliad, we removed items that would not make sense to a reader without exposure to ancient Greek poetry. We also removed notes that mainly pointed to where else the reader could see a list of occurrences or information explaining something the Basel commentary did not. Students often translate with many digital and analog aids, but we wanted to see if Tompkinsia could change reader perception of commentaries without interrupting the reading flow as pursuing references usually does.
4.2.3 Tompkinsia Design.
The Tompkinsia website depicting one view of this hypertext reading environment. The site consists of a white source text pane taking up the left two thirds of the screen, and the white commentary pane taking up the remaining third. These panes are outlined with a thin offwhite border. On the left is a section of the Iliad in black Greek text (6.45-658 specifically). There are line numbers for every multiple of 5 on the left-hand side of the text. In the commentary pane, there are 6 normal commentary notes in a cornflower blue, each one beginning with the relevant line numbers, followed by a bolded substring from the source text, which are in turn followed by notes in English text. The seventh note is a claim in midnight blue text, under which is a summary of what the notebook view finds. This summary is also this midnight blue text, enclosed in a beige rectangle with a rounded black border. Below that are two buttons: a cornflower blue button reading “Show Claim Analysis" and a gold button reading "Download Notebook".
The Tompkinsia website depicting one view of this hypertext reading environment. The site consists of a white source text pane taking up the left two thirds of the screen, and the white commentary pane taking up the remaining third. These panes are outlined with a thin offwhite border and above them both is a thin midnight blue header whose white heading reads “Tompkinsia." On the left is a representation of a Jupyter notebook assessing the claim in the commentary pane. This representation is enclosed in an iframe with a cornflower blue border that users can scroll around in. Inside the iframe, the notebook begins with a heading in large bolded black text. In thinner, body-size black text, the claim is repeated and interpreted. There is material in three slightly smaller black headings that break verification of the claim into steps. 0) contains just a list of requirements for running and tweaking the notebook, 1) investigates the presence of agkulon and arma, and 2) is the heading for an epithet investigation not present in this image. Right under 1), we can see some plain black text and a pale yellow box containing monospaced text meant to signal that it is a code cell. Below that, we have some more clarification of what we're trying to find, and then a larger code chunk for finding all possible instances of agkulon modifying arma. The table of results below consists of a white background, black borders, and black text, showing just one result: the instance the discussed note is attached to. The last thing following that and before 2) is an assessment of whether this verifies the first part of the claim.
In Fig. 1 the Tompkinsia reading environment shows the Iliad and the Basel commentary. The source text is in the left pane, and the Basel commentary is to the right. [6] shows that digital representations of commentaries in a multi-document hypertext space can seem like an overwhelming wall of text. Rather than reprise that design, we made claim comments with notebook analyses a navy blue and comments without notebook analyses a lighter blue. To view a static Jupyter notebook associated with a claim, clicking a blue button in the claim on the right replaces the contents of the left pane with the static notebook view, as Fig. 2 shows. Here, an iframe shows the notebook as Markdown, code, and output sections as common in the Jupyter paradigm, but code sections are static formatted text that do not run in browser. Pressing the blue button reverts the left pane to its prior contents, and the gold button next to it lets the user download the notebook to tweak and run themselves. Between the original comment and the buttons is a summary of the analysis, bordered to imply it is an addendum to the former. Tompkinsia is in action at GitHub [3].
4.3 Procedure
This within-participants study had three phases: onboarding, control, and a test phase. It was also run with physical and Zoom attendance. During onboarding, participants attending physically were first reminded that their screen would be recorded at a certain point in the study, and there was an optional opportunity to have Socratic conversations with one researcher around claim verification recorded. Zoom participants were told that their whole meeting would be recorded. After completing the pre-experiment survey, all participants were given a basic overview of claim verification with Tompkinsia. After onboarding, to prevent unintentional priming, participants would proceed either to the control phase and then the test phase (Group A), or to the test phase and then the control phase (Group B). Participants were assigned to experimental groups alternatingly, and each group had one expert. In the control phase, participants completed a survey that presented a claim, asked participants to note their feelings on the claim, and then build a plan to verify the claim. Lastly in the test phase, participants completed a survey beginning with a different claim, and asked participants to note their feelings about the claim. Participants were told that they would build a verification plan, but first they would browse Tompkinsia for 10 minutes to explore analyses of claims in that same commentary that may use methods helpful for their plan. After browsing, participants completed the last survey and final part of the study, where they would build a verification plan for the claim introduced to them at the start of that phase.
4.4 Surveys
Participants completed 4 surveys hosted on Google Forms. Participant identity was anonymized. The onboarding survey captured important demographic and experience information. The control survey required participants to note how they felt about a commentary claim, then build a plan to verify the claim. The test phase separated the control survey questions into an observation survey before participants’ time with Tompkinsia, and a plan-building survey after that. This subsection reviews this study's surveys, but they can be found at [4].
The pre-study survey captured important demographic and experience information. We asked for participants’ gender and the number of years of experience they had in Ancient Greek. We binned experience by years into no experience, 1-2 years, 3-4 years, 5-6 years, and more than 6 years. Next, there were 2 yes or no questions regarding prior use of 1) a commentary and 2) a hypertext reading environment like Perseus Digital Library. Familiarity with any of these may facilitate participants’ support of our hypotheses, as the germane load of learning the Tompkinsia system may be reduced from its similarity to these translation aids. The survey ended with 4 questions rating familiarity on a 5-point Likert scale with Python, Jupyter notebooks, quantitative textual analysis methods, and working with text in Python to assess the frequency of a textual phenomenon. Unfamiliarity did not disqualify participants, as we wanted to know whether we could reinforce critical thinking in those unfamiliar with such practices.
The control survey showed a claim from Smith's commentary on the History, which participants were not expected to know. The claim reads, “σημανῶ: as 2. 45. 8 and freq. in Thuc.; not common elsewhere in Attic prose" [33]. Participants were asked to guess what the reasoning behind the claim was, and then rated their acceptance of the claim, acceptance of its reasoning, and extent of understanding of the claim's reasoning all on a 5-point Likert scale. We wanted to measure both changes in claim perception and engagement. The rest of the survey guided participants through building a plan to verify the claim with 3 questions. A challenge, recognized here and in [6], is balancing both sufficiently preparing humanist students for engaging in the requisite math and logic, and ensuring that students feel empowered to engage with this content. We wanted this phase to assess student ability to apply what they learned in Tompkinsia, but we also wanted them to feel that overall language learner status or math and logic experience was not an emotional or intellectual blocker for building a plan. Rather than having 2 short answers where we asked what would be verified to assess the claim and what the verification plan was, we chose to include an intermediate question where participants could pick from 11 methods to use in their verification plan, including an “Other" option. Participants then explained how the methods they chose would fit into their plan. This yielded a frame of reference for folks conscious about their level of knowledge or their level of confidence in more mathematical or logical techniques. We still wanted to assess participants’ level of understanding with method choice, so 5 of our 10 defined methods were steps in an original plan to verify the claim and the other steps were nearly correct methods for claim investigation [4]. Lastly, participants could discuss any of these short answers with the researcher, who would co-build the plan with the participant in a Socratic way. The researcher would share knowledge to help the participant understand the goal and any methods they came across so far, but would deflect attempts at being used to devise the plan by asking the participant questions meant to guide them towards building the plan. Such conversations were recorded with informed consent.
The surveys in the Tompkinsia test phase are identical to the control phase, except that the first one shows the claim with perception and engagement questions, and the second one reprises the claim before asking the 3 questions about building a verification plan from the control phase. In this phase, the claim discussed is from the Basel Commentary and is present in Tompkinsia. It reads, “ἄξια: like other adjs. of evaluative content, used almost exclusively in character speech (exception: 23.885); see GRAZIOSI/HAUBOLD ad loc" [7]. The first survey ended with a summary of their task while browsing, and participants were verbally reminded of this task after survey submission. The second survey began with a reminder of the claim. The list of methods question in this plan-building section had only 6 non-other members, as the original plan for verification only had 3 steps [4]. Each claim is in this study is slightly different, but it participants could present a sound verification plan with methods from another analysis.
5 Results
5.1 Demographics and Experience
All participants had translated using a hypertext reading environment. The ratio of participants who had used a physical or digital commentary to those that did not is 3:1. 1 expert had more than 7 years of experience, and the ratio of participants in their first year of ancient Greek to those with 1-2 years of experience was also 3:1. Regarding Python experience, both experts reported peak familiarity (5), 1 participant responded with medium familiarity (3), 1 participant responded with some familiarity (2), and 8 participants reported total unfamiliarity (1). Likewise, regarding Jupyter experience, both experts reported total familiarity (5), 1 participant responded with some familiarity (2), and 9 participants reported total unfamiliarity (1). For both quantitative textual analysis and specific use case experience questions, both experts reported total familiarity (5), 2 participants responded with some familiarity (2), and 8 participants reported total unfamiliarity (1).
5.2 Perception and Engagement
Two histograms depict both average remapped Likert scale scores rating their likelihood to accept the claim, their likelihood to accept the reasoning behind the claim, and their self-reported ability to follow said claim, with the left histogram representing Group A and the right histogram representing Group B. Each histogram consists of three translucent two-bar clusters, with a thinner black 95% confidence interval error bar accompanying each translucent bar representing an average score for the same distribution. In each cluster, the average bar on the left is azure, representing the average for a metric during the control phase, and the average bar on the right is orange, representing the average for a metric during the Tompkinsia-full phase. A legend reports these colors on the right side of the figure. In each histogram, the left cluster depicts Likert scale score distribution for rating their likelihood to accept the claim, labeled on the x-axis as Acceptance, the middle cluster depicts Likert scale score distribution for rating their likelihood to accept the reasoning behind the claim, labeled on the x-axis as Reasoning, and the right cluster depicts Likert scale score distribution for rating their self-reported ability to follow said claim, labeled on the x-axis as Following. What follows is a list of values reporting these average scores as well as the upper and lower bounds of each 95% confidence interval. Let's first look at the scores for Group A. Both average scores across phases for the acceptance metric are 1.6667, and they have matching 95% confidence ranges of 1 to 3. Looking at the reasoning metric, the control phase reports an average score of 3.2222 with a 95% confidence range of 1.8889 to 4.5556, whereas the Tompkinsia-full phase reports an average score of 3.6667 with a 95% confidence range of 2.0667 to 5. As for the following metric, the control phase reports an average score of 2.6667 with a 95% confidence range of 1.3333 to 4, whereas the Tompkinsia-full phase reports an average score of 3.6667 with a 95% confidence range of 2.3333 to 5. The scores for Group B are as follows. Looking at the acceptance metric, the control phase reports an average score of 3.1111 with a 95% confidence range of 1.6667 to 4.3361, whereas the Tompkinsia-full phase reports an average score of 4.4444 with a 95% confidence range of 3.7778 to 5. Turning to the reasoning metric, the control phase reports an average score of 2.2222 with a 95% confidence range of 1.2222 to 3.3333, whereas the Tompkinsia-full phase reports an average score of 2.5556 with a 95% confidence range of 1.5556 to 3.6667. As for the following metric, the control phase reports an average score of 3.3333 with a 95% confidence range of 2 to 4.6667, whereas the Tompkinsia-full phase reports an average score of 1.6667 with a 95% confidence range of 1 to 2.3333.
For each participant, we took their initial range on all perception and engagement Likert scales and remapped all values onto the original range of 1-5. Fig. 3 shows the averages of these scores split by group, with thin black 95% confidence error bars showing the relevant population. There are not significant insights for rating acceptance of the reasoning behind the claim, but there are some changes in group response otherwise. Most notably, the average score for participant acceptance of a claim dropped for Group B between phases. There is also a notable rise in score for participants’ extent of being able to understand and follow the claim's reasoning in Group B.
5.3 Building A Plan
This thematic analysis was conducted with three specific subtopics in mind. To this end, we found our themes fit within one or more of these subtopics. This figure depicts a tri-circular sky blue Venn diagram on a white background where each circle is described by a subtopic in a sky blue and themes are in a navy blue, inhabiting at least one circle of this diagram. All circles overlap with each other, with one circle positioned in the center of the top of the figure, associated with identifying reasoning, a second circle positioned in the lower left, associated with identifying what to verify, and the remaining circle positioned in the lower right of the figure, associated with building a verification plan. This last circle is the only one with themes in it that do not intersect with other circles, and they are Hit Rate and Selection Size. In the overlap only between that last circle and identifying what to verify, the themes present are Filtering Method, Math Level, Search Type, and Number of Steps. Looking at the overlap only betwixt identifying what to verify and identifying reasoning, the only theme present in this space is Uncertainty. Last but not least, all circles share two themes: Phase and Analysis Method Type.
For all short answer questions across non-onboarding phases, 1 author ran a thematic analysis process [8] on each question. In each set, each answer was treated as a data extract and coded with at least 2 codes, which were grouped and iteratively reworked into coherent themes. During thematic analysis, it was found that codes, subthemes, and entire themes were applicable to more than 1 question. Fig. 4 shows these themes’ overlap. Regarding the last short answer, a theme set was developed to track both the complexity of methods chosen and how many different methods were in a plan.
While extracts were labeled by the individual and group IDs of a participant, one universal theme was the phase the extract was written during, noted with the codes control or test across all questions. Another universal theme was types of analysis methods in extracts. Word sense and spot checking were codes associated more with the Tompkinsia-full claim, with word sense often used to describe the meaning of ἄξια (“I would then narrow my search to whether those other uses of the word were in character speech") and spot checking describing reviewing the context of each instance of a phenomenon (“I would want to then look at all occurrences of the word in this subset."). Frequency recognition represented musings on a phenomenon's frequency (“I could then try to find the frequency of whether the word is almost exclusively in direct speech"), and learned method application coded examples of a participant applying methods seen in the onboarding or test phases. This theme has one more code: comparing frequencies (“Getting relative frequencies in and out of direct speech"). Codes in the theme of uncertainty varied from neutral, with guessy (“I would need to search this word…perhaps filtering by part of speech.") to doubting the self (“I am unsure") to not guessing with the I don't know code (“not enough information"). For the verification plan short answers, themes measuring answer quality included number of steps, search type, filtering method, and level of math used. For number of steps, answers to these questions were coded as a one step or multi step approach. Search type had two codes: simple search, for answers only stating what participants would search, and initial search, for answers where search was the first but not final step. It made sense to filter results by different criteria for each claim, so it follows that many control answers were coded filter by author whereas many test answers were coded filter by work. Level of math recognized advanced methods in verification answers (“This would allow me to set up a non-paramteric t-test") though some participants had less of this background, and those answers were coded as proto-math (“One would likely need to run calculation to assess whether the word itself is rare"). 2 themes were exclusive to one verification plan question: how many methods participants chose, and how many of those methods were correct, or hit rate. Selection size had 3 codes: one match covered those that chose only 1 method, hit and a miss covered those that chose 2 methods but only mentioned 1 in their final plan, and multi match covered those that chose more than 2 methods. Hit rate codes included full match, for participants that mentioned each method, unmentioned method, for participants that mentioned all but 1 chosen method in their answer, and unmentioned methods, for participants whose answer did not mention 2 or more chosen methods.
Members of each theme and code were examined to find how Tompkinsia affected critical thinking skills. In the question about what would be verified, half of Group A's answers were coded with simple search and one step, but most Group B members had their answers coded with multi step. 4 Group B members also had their answers coded with filter by work, as did 2 of 6 Group A members. In the verification plan question, defining codes of Group A's answers include initial search (5 in Group A, 3 in Group B), filtering by work (4 in Group A, 2 in Group B), and word sense (6 in Group A, 3 in Group B). 4 Group B members’ answers were coded with learned method application, and 1 Group A answer has the same code.
Table 1: Method Complexity Scores
Complexity Score | Codes |
1 | one step, initial search, simple search |
2 | multi step, analysis method type codes, filtering codes |
3 | comparing frequencies |
4 | advanced methods |
This analysis can also tell us something about the complexity of methods chosen and how many methods the plan included. For each group, we collected data split by phase on the average number of codes per extract. In all phases Group A members had 8.5 codes on average. Yet in Group B, average code volume per extract began at 6.833 codes in the test phase and rose towards 8.6 codes per extract on average in the control phase. Likewise, we gave methods a complexity score as seen in Table 1 and for each extract we took the average complexity score across codes. We averaged those scores similarly to the prior metric, split by phase and group. Group A saw a rise in average complexity score from 1.717 in the control phase to 1.9317 in the test phase, and Group B's rise began at 1.533 in the test phase to 1.93 in the control phase.
5.4 Interaction in Tompkinsia
2 authors coded screen recordings with a schema recording when Tompkinsia sessions started and ended, and when any Show/Hide Claim Analysis button was clicked. Note that 2 Group A members and 1 Group B member did not click on any of those buttons. On average, Group A members spent 1m 34s “scene setting" before clicking on a claim, and in Group B that time was 2m 30s. Group A members spent 1m 58s on average inside each claim, and in Group B this average time was 1 minute. Group B was also the only group where members would return to read the same claim, with the first claim having the highest average of clicks of claims at 1.16 clicks per claim, though the second claim was also revisited.
6 Discussion
H1 proposed that participants will form plans showing understanding of methods from Tompkinsia. We found that 1) 4 Group B members and 1 Group A member applied methods from Tompkinsia to the verification plan, and 2) Group B members reread to claims. Because of time constraints we could not test this with more participants, but one interpretation of this data is that this rereading contributes to better retention and application of methods from Tompkinsia, validating H1. H2 proposed that participants’ acceptance rating of claims would drop, with Group B doing this to a greater degree. Figure 3 shows ratings like acceptance split by phase and group, and though acceptance hardly budges in Group A, retention in Group B leads to a greater drop during the control phase. This partially validates H2. One may think across groups the first phase always primes participants in some way, but a counterpoint to this is what we found with certain Group A outliers, like participant G2A4. Their codes for control phase answers resembled other Group A peers, but after the test phase, their answers and codes fit more with the multi step paradigm defining Group B, and in their verification plan for the test phase, they reference the verifications they viewed. G2A4 was not the only outlier in Group A that may have gained a better understanding of these claims as these results show, but their development in this study was most visible. Given these results, Tompkinsia may effectively be a tool enabling language learners to empower themselves to verify such claims.
7 Conclusion and Future Work
This user study investigated reader perception and reasoning of claims in commentaries, and assessed whether Tompkinsia, a hypertext reading environment for an Ancient Greek source text and a corresponding augmented commentary with data-driven analyses in integrated Jupyter notebooks, could reinforce critical thinking skills regarding such claims. When asked to build a plan for verifying a claim with minimal mathematical background, participants that actively read and re-read the claims in Tompkinsia were more likely to apply methods they learned to this plan. There was also a drop in acceptance of these claims between control and test phases for Group B. There is interest in running this as a larger user study with a multi-institutional cohort of classes at different levels, and we also want to work with commentators to make commentaries that continue this conversation with reliable and transparent analysis.
Acknowledgments
We would like to thank the Schmidt Sciences for funding this work through grant HAVI-2025-2.
Source
Imported from ACM’s structured HTML source. ACM Reference Format: Sarah Abowitz, Sasha Spala, and Gregory Crane. 2026. Tompkinsia: Data-Driven Commentary Augmentation for Critical Thinking. In 37th ACM Conference on Hypertext (HT '26), September 14--18, 2026, London, United Kingdom. ACM, New York, NY, USA 7 Pages. https://doi.org/10.1145/3800935.3830877
References
[1] 2024. GLAUx: How?Retrieved February 10, 2026 from https://glaux.be/how.php
[2] 2024. W3.CSS Template. Retrieved February 10, 2026 from https://glaux.be/search.php
[3] 2026. lepidopterane-atsmith/tompkinsia: A data-driven commentary augmentation design. Retrieved February 10, 2026 from https://github.com/lepidopterane-atsmith/tompkinsia/tree/main
[4] 2026. tompkinsia/surveys.md at main · lepidopterane-atsmith/tompkinsia. Retrieved April 27, 2026 from https://github.com/lepidopterane-atsmith/tompkinsia/blob/main/surveys.md
[5] Sarah Abowitz, Alison Babeu, and Gregory Crane. 2024. Bridging the Understanding Gap: Helping Readers Engage Directly with Foreign-Language Sources More Easily. In Proceedings of the 24th ACM/IEEE Joint Conference on Digital Libraries. Association for Computing Machinery, New York, NY, USA. https://doi.org/10.1145/3677389.3702539
[6] Sarah Abowitz and Gregory Crane. 2025. Student Use of Commentaries with Inline Reference Resolution. In Proceedings of the 36th ACM Conference on Hypertext and Social Media(HT ’25). Association for Computing Machinery, New York, NY, USA, 146–155. https://doi.org/10.1145/3720553.3746681
[7] Anton Bierl, Joachim Latacz, and Magdalene Stoevesandt (Eds.). 2016. Homer's Iliad: the Basel commentary (1st ed.). Walter De Gruyter Inc., Walter De Gruyter Inc. ; Boston, MA.
[8] Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3, 2 (2006), 77–101.
[9] Gregory Crane. 1996. Building a digital library: The Perseus Project as a case study in the humanities. In Proceedings of the first ACM international conference on Digital libraries. 3–10. https://dl.acm.org/doi/pdf/10.1145/226931.226932
[10] Gregory Crane. 2004-. Perseus Digital Library Project. Retrieved February 10, 2026 from http://www.perseus.tufts.edu/hopper/
[11] Gregory Crane, Alison Babeu, Lisa M Cerrato, Amelia Parrish, Carolina Penagos, Faroosh Shamsian, James Tauber, and Jake Wegner. 2023. Beyond translation: engaging with foreign languages in a digital library. International Journal on Digital Libraries 24, 3 (2023), 163–176.
[12] R. Cribiore. 1996. Writing, Teachers, and Students in Graeco-Roman Egypt. Scholars Press. https://books.google.com/books?id=NQwcAQAAIAAJ
[13] Eleanor Dickey. 2007. Ancient Greek scholarship: a guide to finding, reading, and understanding scholia, commentaries, lexica, and grammatical treatises, from their beginnings to the Byzantine period. Oxford University Press, Oxford.
[14] Elaine Fantham. 2002. Commenting on Commentaries: A Pragmatic Postscript. In The Classical Commentary: Histories, Practices, Theory, Roy Gibson and Christina Kraus (Eds.). Brill, 403–421. https://doi.org/10.1163/9789047400943_017
[15] Center for Hellenic Studies. 2023. New Alexandria Open Commentary Platform. Retrieved August 1, 2024 from https://oc.newalexandria.info/
[16] Don Fowler. 1999. Criticism as commentary and commentary as criticism in the age of electronic media. In Commentaries/Kommentare, Glenn Most (Ed.). Vanderhoeck & Ruprecht, 426–442.
[17] Christopher Francese. 2011-. Dickinson College Commentaries. Retrieved August 2, 2024 from https://dcc.dickinson.edu/about-dcc
[18] R.K. Gibson. 2002. Cf. e.g.: a Typology of Parallels and the Role of Commentaries on Latin Poetry. In The Classical Commentary: Histories, Practices, Theory, Roy Gibson and Christina Kraus (Eds.). Brill, 331–358.
[19] Amanda Goodman and Suzanne Conklin Akbari (Eds.). 2023. Practices of Commentary: Medieval Traditions and Transmissions. Arc Humanities Press. https://doi.org/10.2307/jj.6253300
[20] Barbara Graziosi. 2009. Commentaries. In The Oxford Handbook of Hellenic Studies. Oxford University Press. arXiv:https://academic.oup.com/book/0/chapter/334629323/chapter-ag-pdf/44445706/book_38587_section_334629323.ag.pdf https://doi.org/10.1093/oxfordhb/9780199286140.013.0068
[21] Carolin Hahnel, Dara Ramalingam, Ulf Kroehne, and Frank Goldhammer. 2023. Patterns of reading behaviour in digital hypertext environments. Journal of Computer Assisted Learning 39, 3 (2023), 737–750. arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/jcal.12709 https://doi.org/10.1111/jcal.12709
[22] Nick Kallen. 2013. nkallen/pseudw: Language learning software. Retrieved February 10, 2026 from https://github.com/nkallen/pseudw
[23] Alek Keersmaekers. 2021. The GLAUx corpus: methodological issues in designing a long-term, diverse, multi-layered corpus of Ancient Greek. In Proceedings of the 2nd International Workshop on Computational Approaches to Historical Language Change 2021, Nina Tahmasebi, Adam Jatowt, Yang Xu, Simon Hengchen, Syrielle Montariol, and Haim Dubossarsky (Eds.). Association for Computational Linguistics, Online, 39–50. https://doi.org/10.18653/v1/2021.lchange-1.6
[24] Christina S Kraus and Christopher A. Stray. 2015. Form and Content. In Classical Commentaries: Explorations in a Scholarly Genre, Christina S. Kraus and Christopher Stray (Eds.). Oxford University Press, 1–18. https://doi.org/10.1093/acprof:oso/9780199688982.003.0001
[25] Scott Martens. 2012. TüNDRA: A Web Application for Treebank Search and Visualization. In Proceedings of The Twelfth Workshop on Treebanks and Linguistic Theories (TLT12)(TLT ’12). The Institute of Information and Communication Technologies, Bulgarian Academy of Sciences, Sofia, Bulgaria, 133—144. https://www.academia.edu/download/87484106/TLT12Proceedings.pdf#page=139
[26] Willard McCarty. 2002. A Network with a Thousand Entrances: Commentary in an Electronic Age. In The Classical Commentary: Histories, Practices, Theory. Brill, 359–402. https://doi.org/10.1163/9789047400943_016
[27] Tiziano Piccardi, Miriam Redi, Giovanni Colavizza, and Robert West. 2020. Quantifying Engagement with Citations on Wikipedia. In Proceedings of The Web Conference 2020 (Taipei, Taiwan) (WWW ’20). Association for Computing Machinery, New York, NY, USA, 2365–2376. https://doi.org/10.1145/3366423.3380300
[28] M. A. Rodda. 2021. A corpus study of formulaic variation and linguistic productivity in early Greek epic. http://purl.org/dc/dcmitype/Text. University of Oxford. https://ora.ox.ac.uk/objects/uuid:1e682001-b916-4322-adc3-52857d93b92b
[29] Matteo Romanello. 2020-. Ajax Multi-Commentary Project. https://mromanello.github.io/ajax-multi-commentary/, lastaccessed =
[30] Matteo Romanello and Sven Najem-Meyer. 2024. A Named Entity Annotated Corpus of 19th Century Classical Commentaries. Journal of Open Humanities Data 10, 1 (Jan. 2024). https://doi.org/10.5334/johd.150
[31] Matteo Romanello, Sven Najem-Meyer, and Bruce Robertson. 2021. Optical Character Recognition of 19th Century Classical Commentaries: the Current State of Affairs. In The 6th International Workshop on Historical Document Imaging and Processing(HIP ’21). Association for Computing Machinery, Lausanne, Switzerland, 1–6. https://doi.org/10.1145/3476887.3476911
[32] M. Sanchiz, F. Amadieu, J. Lemarié, and A. Tricot. 2022. Do graphic and textual interactive content organizers have the same impact on hypertext processing and learning outcome?Journal of Computing in Higher Education 35 (June 2022). https://doi.org/10.1007/s12528-022-09328-z
[33] Charles Forster Smith (Ed.). 1913. Thucydides. Book VI (1st ed.). Ginn and Company Boston, Boston, MA.
[34] Susan Stephens. 2002. Commenting on Fragments. In The Classical Commentary: Histories, Practices, Theory, Roy K. Gibson and Christina Shuttleworth Kraus (Eds.). Brill, 67–88. https://doi.org/10.1163/9789047400943_005
[35] Daniel P. Tompkins and Adam Parry. 1972. Stylistic characterization in Thucydides: Nicias and Alcibiades. Cambridge University Press, 181–214.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime