Docuverse Despatch: Information Farming For The Collective

Southampton, Hants

ABSTRACT

Since the 1993 paper on Information Farming[5], hypertext has grown in scale and in the degree of its collective editing and use. This paper reflects on what these changes in scale and volume mean for the task of the information farmer and asks if we understand the skills and tools needed for the task of sustaining the docusphere.

CCS CONCEPTS


    Human-centered computingHypertext / hypermedia;

Collaborative content creation; Computer supported cooperative work;

    Information systems → Web searching and information discov-

ery; • Applied computing → Hypertext / hypermedia creation; • Computer systems organizationEmbedded systems; Re- dundancy; Robotics; • Networks → Network reliability;

KEYWORDS

hypertext; docusphere; knowledge management; transclusion; metadata; collaborative working; links; trails; linkbases; information farming; wikis; wikipedia

ACM Reference Format: Mark W. R. Anderson. 2020. Docuverse Despatch: Information Farming For The Collective. In Proceedings of the 31st ACM Conference on Hypertext and Social Media (HT ’20), July 13–15, 2020, Virtual Event, USA. ACM, New York, NY, USA, 4 pages. https://doi.org/10.1145/3372923.3404774

1 INTRODUCTION

In 1993, at the time of ‘Enactment in Information Farming’ [5], the hypertext field was small but burgeoning. Collaborative work tended to be localised and not at the scale now seen. The World Wide Web had arrived, running atop the Internet, but was yet to reach scale or significant visibility to the general public. Since then, though not by intent, the Web’s subsequent dominance—and early technical limitations—have skewed hypertext’s[1] evolution and some good early ideas in the field are yet to be fully realised.

‘Information Farming’[5, p.242] was originally described thus:

1I use the term hypertext in its portmanteau meaning, i.e. to reflect hypermedia as a whole, even though textual use predominates.

Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. © 2020 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-7098-1/20/07...$15.00 https://doi.org/10.1145/3372923.3404774

“Information farming (or gardening) views the cultivation of information as a continuing, collaborative activity performed by groups of people working together to achieve changing individual and common goals. Where the mine and factory serve the organization, the information farm is a computational space where colleagues and employees may work together on shared tasks and also pursue individual goals. The focus is neither on extraction (as in the mine) nor on stockpiling (as in the information factory), but on continuous cultivation and community.”

This notion sets the construction and maintenance of hypertext aside from the then more established field of information retrieval and its more factory-like needs. It recognises both the liberal intertwingularity of Nelson’s ‘docuverse’ [25, p.2/53] and Engelbart’s notion of bootstrapping and ongoing augmentation of human capabilities [11]. The co-creation of human and machine is also acknowledged.

The original description of information farming reflected the emergent exploration of non-hierarchical informational relationships as found in spatial hypertext [20], moving outside the straightjacket of the outline. Also recognised were problems of incompatible existing metaphors [18].

Those might be reflected more generally in McLuhan’s observation of ‘rear-view thinking’ [26, p.98]. Pertinently, McLuhan pointed out [21, p.9] that it is important to attend to the medium—here, hypertext—to properly understand (its message). The information farmer does just that.

2 INFORMATION FARMING TODAY

What, then, has changed in the intervening years? Whereas the erstwhile information farmers were tending smaller, more artisanal, fields and trails, a subsequent development has been scale. The old isolated hypertextual farmsteads are mostly now part of larger, more organised collectives.

These changes are not born from dogma but have evolved from the demands and constraints of large organisations, which necessarily have larger overall bodies of information. I believe the hypertextual challenges of these larger systems pose new challenges and it is the impact of these on which I propose to concentrate in this paper.

These big hypertexts represent the collectivised knowledge of public sector bodies (e.g. government ministries, hospitals, courts), private sector companies, as well as museums, charities and academia. In addition there are public, editable, resources on the Web of which Wikipedia is the pre-eminent example.

Another important contextual shift in recent years has been the move away from paper to digital storage as the primary means of

record. Whereas Bolter talked of a ‘late-print age’ in 1990 [7, p.210], we are moving to a post-print age. A particular outcome of this change is that information and knowledge of note may now be entirely digital, only held within or accessed via a hypertext.

2.1 Hypertext, the docuverse and the ‘cloud’

The concept of the docuverse was defined by Nelson in Literary Machines[2] to describe the larger set of all linked documents within the (Xanadu system) hypertext. In part this overlaps the more contemporary notion of ‘the cloud’[3].

The Internet, primarily via the World Wide Web, has some of the form of a docuverse—or the potential to be so. Of course, not all Internet connections are persistent, and much data resides only a limited (often single) locations. We should be mindful of this shift in medium and McLuhan’s warning about ‘rear-mirror thinking’ [26, p.98]; today’s hypertexts are not simply our grandparents’ bookcase, digitised.

2.2 We’ll need a bigger barn

Regardless of the portmanteau term chosen, it is within this abstracted environment that large hypertexts often now live, their structure not necessarily clear to those working upon them. As storage space can be flexed more easily now, care is necessary to avoid accretion of moribund data though insufficient archiving and pruning of surplus material.

An example of this disintermediation of old media are documents used in Google’s online office applications. Google ‘docs’, ‘sheets’, etc., are actually not truly discrete documents: they exist only on Google’s servers. Additionally, there is no published ‘google docs format’, Google stating “There is no ‘google docs format’.” [13]. Such pieces of data only take document form, e.g. for robust offline use, by exporting them into some other form (PDF, MS Office). This new reality does not detract from the convenience of such a system but it does not necessarily reflect the lay person’s understanding of a ‘document’.

To remove a Google document from the hypertext requires exporting it out translated into a defined file-based format. Within the system, the document has no final state - simply its content at time of last edit. The latter is an example of the shift from digital documents which are a facsimile form of paper, stored as a discrete computer file (e.g. a Word file), to a docuverse where an online document has no published format or discrete data file.

In turn this dis-intermediated form of document mirrors large public hypertexts like wikis (with Wikipedia the most obvious example) where there are not article ‘documents’ files as such.

2All of storage near and far must therefore become a united whole—what is now called a ‘distributed data base’. Actual locations become essentially invisible to the user; or, in that traditional phrase, “You don’t care where it’s stored”. The documents and their links unite into what is essentially a swirling complex of equi-accessible unity, a single great universal text and data grid, or, as we call it, the docuverse. [25, p.2/23]. 3The NIST definition of cloud computing: Cloud computing is a model for enabling ubiquitous, convenient, on-demand network access to a shared pool of configurable com- puting resources (e.g., networks, servers, storage, applications, and services) that can be rapidly provisioned and released with minimal management effort or service provider interaction. [22, p2]

2.3 Who speaks for the hypertext?

However named, organisational knowledge is often now contained in a cloud-based hypertext which is essentially that organisation’s docusphere. Therein lies a problem. Within a cloud-type large hypertext service, the OS substrate and software needed to use the hypertext are made available and supported by an IT department: domain experts contribute their expertise: knowledge managers account for the corporate value and risks of the hypertext (content); librarians archive information no longer in live use; security controls access whilst lawyers and data controllers decide to whom access my be granted.

But, who actually maintains the on-going health of the hypertext, and speaks for it at a managerial level so as to ensure resource for its needs? The purview of these existing skills still leaves an unnourished lacuna within the core of the hypertext.

We lack an enhanced and seriously-used hypertextual expertise. In other words, we lack explicit hypertext structure experts both as part of the teams producing new materials and those maintaining existing material (this is discussed further in Section 5).

2.4 Why can’t somebody else do it?

This expertise gap within the hypertext might be assumed resolved by ‘Linus’ Law’ [27, p.30], i.e. the idea that with enough eyes on the task no problem lies hidden. Though popular in Open Source software (with a narrower contributing skill-base) it does not necessarily work so well for more complex multi-disciplinary and multi-jurisdictional domains like an organisational hypertext.

An example of hidden technical debt can be seen in Wikipedia which has a high level of redirects implemented as a quick-fix for resolving article naming [3, p.120 Table 2]. Hypertext tasks that ‘anybody’ might do are happily left to the presumption of ‘somebody else’ doing them, and that may turn out to be...‘nobody’.

3 DÉJÀ VU?

A touchstone for hypertext’s progress as a medium is Halasz’s ‘Seven Issues’. First set out in 1988 [15], he reprised them in 1991[16]

and 2001 [17]. Many original issues are now resolved and progress was last summarised by Millard & Ross’ ‘Web 2.0: Hypertext by Any Other Name?’ [24], but a few issues (from the 2001 list) merit further note here:

    Search. Effort appears more likely to be invested in new

search code (‘the next version will be better’) than into improving structure of the hypertext structure to aid both search and general traversal. There are some exceptions, where search analysis is used to inform the hypertext’s (re)design [33].

    Composites. Transclusion and re-use occurs less than its util-

ity would suggest. Despite Berners-Lee’s vision for the Semantic Web [4], the amount of integration of semantic systems to general hypertext is low. This in an area that ought to be a natural rendezvous for human and computer derived link networks.

    Versioning. Many current systems record state often, allow-

ing restoration to a particular state or time. However, without additional edit comments (i.e. metadata), the edits alone do not preserve context as to the reason for the edits [2]. Editors

contributing domain expertise may lack the training or encouragement to record meaningful comments, whilst IT staff may not know what extra metadata for which to provision. For hypertexts generating significant information requiring archive, it is important to be able to find such material in the general mass of data. This is especially true for records of special value: historical, legal or governmental.

    Collaboration. Multiple concurrent editing is not necessar-

ily real collaboration, as opposed to work in concert. Large systems can be prone to territorial behaviours. Assertions of ownership—or lack thereof—within a system do not necessarily reduce innate territorial behaviour and so may need policing [29]. It is not clear that collaborative systems are nurturing a bootstrapping environment such as might improve the efficiency of the whole.

    UI for large information spaces. For long-lived hypertexts,

there is a need for better visual tools for complex tasks such as re-arranging content between articles (nodes) and the attendant link triage. As well as seeing what is connected, it is useful to know what might or should be connected, but is not.

4 BEYOND THE ‘SEVEN ISSUES’

New issues have arisen. With a wider user base and external connections, privacy, trust and control of access have all become relevant to the role of the information farmer, or ‘hypertextualist’. Such strictures imply a privileged role, as the information farmer needs to ‘see’ all parts of the hypertext—more so than most other users.

The view across the docusphere also gives them an orthogonal perspective to that of the majority of those using the hypertext, most of whom will be focussed maily on their areas of expertise or authority. This will necessitate care to avoid friction when both are working on the same information but possibly with misaligned purpose.

4.1 Tending the trails

Bush [8, Sect.8, para2] created the idea of trails and trail-blazers for his Memex. Yet, despite constant interest in the idea of predefined paths through hypertext, a robust and defined method for implementing and sharing such trails (paths through linked nodes) has yet to emerge [28, 30, 32, 33].

Planned trails through the hypertext can help users of the resource to navigate the hypertext without resorting to the guesswork of search. It is hard to (keyword) search for what you do not yet know.

Intentional use of trails implies the hypertext author must be sensitive to hypertextual traversal of the docusphere as much as sections of linear text. In turn this implies a consideration of node size for ‘documents’ that can be appropriately sub-divided.

Having effective trails also means giving attention to link rot [23] both within and at the boundary of the hypertext. It also requires careful use of node size and transclusion to enable suitably granular elements to connect with trails.

4.2 Links or just lines?

In the Web-centric view of hypertext with ‘clickable links’ it is easy to envisage (network) lines simply joining documents. Looking further back, links had a richer place in hypertext, signifying more than navigation. Links were used for argumentation (e.g. gIBIS [9]), and to describe intent (Trigg [31]).

More importantly, just before the Web eclipsed other forms of hypertexts, links were stored externally from documents as seen in systems like Hyper-G [19] and Microcosm [12]. The latter went as far as to offer re-configurable stacks of ‘linkbases’ allowing for (the potential) of contextually different sets of links—and similarly, trails.

A dearth, in an hypertext editing context, of external linkbases or link visualisation tools makes link triage harder than it needs to be. A return of richer understanding of linking also offers prospects for better interfacing of human understanding with the machine reasoning of the Semantic Web [4]. This is a societal benefit yet to be properly realised.

4.3 Narrative and structure

Another overlooked area of hypertextual expertise is narrative in non-linear media. Hypertext study used to be more interested in link structures [6] as relating to narrative. Happily, this conference has seen the return of this strand with E-Lit exhibitions and traversals in 2016 [1][4] and 2019 [14]. Narrative is relative to the design and use of trails and creative informative metaphor(s) at hypertext boundaries.

4.4 Trust, privacy and provenance

An emergent challenge still being explored are the trust structures around access to and re-use of knowledge. This also implies having a sense of the provenance of the docuverse.

Exactly how a hypertext specialist will work with these issues is as yet unclear but they will certainly need to understand their general principles.

5 CONCLUSIONS & RECOMMENDATIONS

So, where are the ‘hypertextualists’? For the ‘Silver Stands’ in which the public might use his Xanadu system, Nelson proposed a ‘Hypercorps’ [25, p3/17]. This cadre were to be the experts on the docusphere:

“The Xanadu Hypercorps is expected to be an unusual and elite group. They will circulate amongst Xanadu stations transmitting skills and outlook. They will not be people who can program or repair a computer; ...Like good librarians the Hypercorps will have an understanding of what materials are available, but they will know how to deal with an avalanche, rather than a trickle, of ideas and information. Like good teachers they will have a sense of how to convey ideas. Like good woodsman they will have a sense of the trails and byways of the territory to be explored.” Written in the 1980s, this neatly foreshadows the interface of skills already described above (Section 2.3). Meanwhile, information farming has certainly become more complex as hypertexts have

4ACM Digital Library has no discrete entry for this event.

grown and it has essentially now become a necessary hypertextual role, albeit one yet to have a formal name or definition: we have no Hypercorps.

One possible elucidation of the role is to take an example from the games industry, perhaps the roles of ordinary users and domain experts relate as a Narrative Designer [10] does to the general story writers, guiding the course and purpose (narrative arc) of the hypertext over time. In the hypertext context, this would be the hypertext specialist guiding the work of the domain specialist or general editor.

Our potential hypertextualists would also benefit from:

    A lessened sense of territoriality

    Being as comfortable looking at the big picture as much as

the minutiae

    A sense for the balance of node size and trail opportunities

    An understanding of link networks

    A good grasp of the role of metadata

    An understanding of co-creation of content within social

machines, by humans and computers

Above all, within their organization, they need to be able to give a voice to the hypertext and nurture the knowledge therein. The first two items of the above list may seem atypical normal human behaviours, but are necessary here.

A better understanding of the role should also help elicit a good name for this new type of specialist; without a name their role is otherwise hard to define and thus resource.

This new role also implies a need to think about the tools required to support such work, though this should perhaps wait to reflect a better understanding of the role they will support. Thus, people first; code later.

The challenge is to now find and enable the people with these skills, and to understand how to nurture the same in others. But, are we even looking?

ACKNOWLEDGMENTS

The views here are the author’s own. They draw upon conversations with colleagues in Southampton’s Web Science and WAIS Groups. Further useful background came from a secondment to the UK Cabinet Office in 2019 allowing experience of a large, bounded, digital docusphere. This article also draws on 15 years supporting the Tinderbox and Storyspace user community and its discussions of hypertext techniques. Special thanks go to Rosemary Simpson for her reviews and suggestions.

REFERENCES

[1] 2016. HT ’16: Proceedings of the 27th ACM Conference on Hypertext and Social

Media (Halifax, Nova Scotia, Canada). Association for Computing Machinery, New York, NY, USA.

[2] Mark W. R. Anderson. 2019. Sustainable Knowledge in Hypertext. Ph.D. Disserta tion. University of Southampton, Southampton, Hants. Pub. pending.

[3] Mark W. R. Anderson, Les A. Carr, and David E. Millard. 2017. There and Here:

Patterns of Content Transclusion in Wikipedia, Vol. Proceedings of the 28th ACM Conference on Hypertext and Social Media. ACM, Prague, Czech Republic, 115–124. https://doi.org/10.1145/3078714.3078726

[4] Tim Berners-Lee, James A. Hendler, and Ora Lassila. 2001. The Semantic Web.

Scientific American 284, 5 (2001), 34–43. http://www.jstor.org/stable/26059207

[5] Mark Bernstein. 1993. Enactment in Information Farming, Vol. Proceedings of

the Fifth ACM Conference on Hypertext. ACM, 242–249. https://doi.org/10. 1145/168750.168837

[6] Mark Bernstein. 1998. Patterns of Hypertext, Vol. Proceedings of the Ninth ACM

Conference on Hypertext and Hypermedia : Links, Objects, Time and Space— Structure in Hypermedia Systems. ACM Press, New York, New York, USA, 21–29. https://doi.org/10.1145/276627.276630

[7] Jay David Bolter. 2001. Writing Space. Vol. 2nd ed. Routledge. 246 pages.

[8] Vannevar Bush. 1945. As We May Think. The Atlantic Monthly 176, 1 (1945),

[9] Jeff Conklin and Michael L. Begeman. 1989. gIBIS: A Tool for All Reasons. Journal

of the American Society for Information Science 40, 3 (1989), 200.

[11] Douglas Carl Engelbart. 1962. Augmenting Human Intellect: A Conceptual

Framework. 1 (1962), 2007.

[12] Andrew M. Fountain, Wendy Hall, Ian Heath, and Hugh C. Davis. 1990. MICRO COSM: An Open Model for Hypermedia with Dynamic Linking., Vol. Proceedings of the First European Conference on Hypertext. Cambridge University Press, The Pitt Building, Trumpington Street, Cambridge, 298–311.

[13] Google. 2020. What is Google Docs format? https://support.cloudhq.net/what-is-

[14] Dene Grigar. 2019. Tear Down the Walls: An Exhibition of Hypertext & Partici patory Narrative, Vol. Proceedings of the 30th ACM Conference on Hypertext and Social Media. Association for Computing Machinery, New York, NY, USA Hof, Germany, 1. https://doi.org/10.1145/3342220.3345459

[15] Frank G. Halasz. 1988. Reflections on NoteCards: Seven Issues for the Next

Generation of Hypermedia Systems. Communications of the ACM (CACM) 31, 7 (1988), 836–852. https://doi.org/10.1145/48511.48514

[16] Frank G. Halasz. 1991. “Seven Issues”: Revisited, Vol. Proceedings of ACM

Hypertext’91 Conference. ACM, San Antonio, Texas, 1–18. Author’s papers. Unavailable from ACM Digital Library.

[17] Frank G. Halasz. 2001. Reflections on “Seven Issues”: Hypertext in the Era of

the Web. ACM Journal of Computer Documentation (JCD) 25, 3 (2001), 109–114. https://doi.org/10.1145/507317.507328

[18] Frank G. Halasz and Thomas P. Moran. 1982. Analogy Considered Harmful,

Vol. Proceedings of the 1982 Conference on Human Factors in Computing Systems. ACM, Gaithersburg, Maryland, USA New York, NY, USA, 383–386. https://doi. org/10.1145/800049.801816

[19] Frank Kappe. 1993. Hyper-G: A distributed hypermedia system, Vol. Proceedings

of INET ’93: International Networking Conference. Citeseer, San Francisco, USA.

[20] Catherine C. Marshall and Russell A. Rogers. 1992. Two Years before the Mist: Ex periences with Aquanet, Vol. Proceedings of the ACM Conference on Hypertext. ACM, 53–62. https://doi.org/10.1145/168466.168490

[21] Marshall McLuhan. 1994. Understanding Media: The Extensions of Man. MIT

Press. 390 pages.

[22] Peter Mell and Tim Grance. 2011. The NIST definition of cloud computing. (2011),

[23] Luis Meneses, Sampath Jayarathna, Richard Furuta, and Frank M. Shipman, III.

    Analyzing the Perceptions of Change in a Distributed Collection of Web

Documents, Vol. Proceedings of the 27th ACM Conference on Hypertext and Social Media. ACM, Halifax, Nova Scotia, Canada New York, NY, USA, 273–278. https://doi.org/10.1145/2914586.2914628

[24] David E. Millard and Martin Ross. 2006. Web 2.0: Hypertext by Any Other Name?,

Vol. Proceedings of the Seventeenth Conference on Hypertext and Hypermedia. ACM, 27–30. https://doi.org/10.1145/1149941.1149947

[25] Theodor Holm Nelson. 1987. Literary Machines. Vol. Edition 87.1. Theodor Holm

Nelson. 280 pages.

[26] Neil Postman. 2005. Amusing Ourselves to Death. Penguin. 208 pages.

[27] Eric S. Raymond. 2001. The Cathedral & the Bazaar. O’Reilly Media. 258 pages.

[28] Siegfried Reich, Les A. Carr, David C. De Roure, and Wendy Hall. 1999. Where

Have You Been from Here? Trials in Hypertext Systems. ACM Computing Surveys (CSUR) 31, 4es (1999). https://doi.org/10.1145/345966.345994

[29] Jennifer Thom-Santelli, Dan R. Cosley, and Geri Gay. 2009. What’s Mine is

Mine: Territoriality in Collaborative Authoring, Vol. Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1481–1484. https://doi.org/10.1145/1518701.1518925

[30] trailblazer.io. 2014. Trailblazer. http://www.trailblazer.io

[31] Randall H. Trigg. 1983. A Network-Based Approach to Text Handling for the Online

Scientific Community. Ph.D. Dissertation. University of Maryland at College Park, College Park, MD, USA. (Only Chapter 4 available online.).

[32] Ryen W. White and Jeff Huang. 2010. Assessing the Scenic Route: Measuring the

Value of Search Trails in Web Logs, Vol. Proceedings of the 33rd International ACM SIGIR Conference on Research and Development in Information Retrieval. Association for Computing Machinery, New York, NY, USA, 587–594. https: //doi.org/10.1145/1835449.1835548

[33] Xiaojun Yuan and Ryen White. 2012. Building the Trail Best Traveled: Effects of

Domain Knowledge on Web Search Trailblazing, Vol. Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. ACM, New York, NY, USA, 1795–1804. https://doi.org/10.1145/2207676.2208312

---

Converted from the ACM version of record under supplied ACM authorization.

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime