I just finished David Weinberger’s book “Everything is Miscellaneous” and heartily recommend it. Appropriately enough, Amazon suggested it to me when I was looking for Clay Shirky’s book “Here Comes Everybody” which pretty much everyone in my online world has recommended and for Cass Sunstein’s book “Infotopia” which Michael Nielsen suggested to me in a response to a blog comment. Each of the three books is broadly about changes enabled by the Web, but they each focus on different aspects. To summarize full-length books on a variety of interesting topics in a couple words: Shirky focuses on the effects on human social groups, Sunstein on markets, and Weinberger on information itself. Of the three, I liked Weinberger’s the best – probably because it resonates most strongly with my own recent studies and interests. The topics cited in the book were uncannily related to my own personal reading/dreaming/experimenting list –Aristotle, Shirky, tagging, ontology, semantic Web, intertwingularity, blogging, Nature Publishing, Vanevar Bush, Dewey, it even mentions LSIDs of all things!
Weinberger writes in an imminently quotable style. Practically every paragraph contains a pithy sentence worthy of repetition. To illustrate, here are just a few that caught my attention (page numbers from paperback):
“...the solution to the overabundance of information is more information” p. 13
“...Making complex, meaningful phenomena explicit can leave us rudderless, force us to oversimplify, and result in statements that are incomplete and misleading.” p. 156
“The meaning of a particular thing is enabled by the web of implicit meanings we call the world” p. 170
“A Semantic Web that loosely stitches together imperfect, smushy, local efforts is not only more likely, it is to be preferred” p. 195
“Paper drives thought into our heads. The Web releases thoughts before they’re ready so we can work on them together.” p. 203
“That was Aristotle’s startling discovery: a thing, standing on its own, is what it is because of its connection to other things like it and other things not like it” p. 219
Obviously, I liked the book and tended to agree with most of its arguments (funny how liking and agreeing often go together). My only complaints about it related to its treatment of social tagging and of the semantic Web; these complaints likely arise from the book’s successful attempt to appeal to a very general audience and the coincidence that I might know a tiny bit more about these areas than the intended audience – hence I would like more detailed treatment of the consequences of more specific aspects of the technologies than is provided. For example, it annoys me when people talk about social tagging as if it is a permanent, stable technology, thus assuming that the ambiguity of tags (as simple Strings unlinked to concept definitions) is a fundamental aspect of all such systems rather than an optional weakness of the design of particular instantiations. It also annoys me when people talk about RDF as an “ontology language” – especially without at least making some attempt to explain its relationship to things that are most definitely ontology languages like OWL..
Overall a good quick read and a useful reference for anyone interested in the continuing evolution of the Web and, through it, our species.
Showing posts with label information science. Show all posts
Showing posts with label information science. Show all posts
Thursday, February 12, 2009
Notes from Everything is Miscellaneous
Posted by Benjamin Good at 1:36 PM View Comments
Labels: book, information science, miscellaneous, quotes, weinberger
Wednesday, October 22, 2008
Heading to ASIST
This Saturday I will be heading down to Columbus, Ohio to attend the annual meeting of the American Society for Information Technology. If you are coming, hope to see you there!
I will be presenting a poster about an empirical approach to the study of (human) indexing systems based on the comparative analysis of the syntax of the terms that are used within the different systems. The work emerged from investigations of the relationship between social tagging (e.g. Connotea) and professional indexing (e.g. MEDLINE). As has been typical in my career thus far, when I got started on the project, I found that I didn't have the tools I felt that I needed to answer the question satisfactorily so I spent most of my time trying to make them.
One of the more interesting outcomes of the work are visualizations of the differences between the structures of the terms used within ontologies, thesauri, and folksonomies. They are surprisingly distinct from one another. The abstract for a full paper (just accepted today) describing the work as well as the programs written and the data processed are available here.
Posted by Benjamin Good at 7:55 PM 0 comments
Labels: asist, conference, information science, library science
Sunday, August 31, 2008
Peter (Google) and Christine (the librarians)
I was very lucky to have my new wife with me at SciFoo for many reasons, not the least of which is that she is much better at socializing than I am and thus managed to introduce me to many people I would never normally have met. One of those people was Christine Borgman, Professor & Presidential Chair in Information Studies at UCLA. While I was struggling to explain my work on the Entity Describer project to her, she noticed Peter Norvig walk by and dragged him over to join our conversation. I guess she must have known him from somewhere but I'm not sure where. Anyway, I didn't realize this at the time, but Peter is head of research at Google. Ahem... did I mention that SciFoo interactions could be somewhat intimidating? So there I am, standing between two giants of modern information science trying to explain what it was I was doing there but mostly trying to get some insight into their respective thoughts on the organization of the world's information. It wasn't a long discussion, but here are the basics.
The main question that I posed to them was whether or not and how semantic tagging (a la ED) is or might be useful. On the surface, the answer from Peter was no and the answer from Christine was yes. However, the truth of the matter is that they were really talking about supporting different functions for the end user - though this fundamental difference became a little lost during the conversation. Google is principally focused on providing the best possible results, to the most people, given the least amount of information in the query - that is, keyword based search of the entire Web. Library-science is typically much more concerned with providing the capacity for people to make very specific requests using much more sophisticated queries that operate over much smaller collections of information (e.g. the library of congress). The fact that there is some overlap in the information needs of the users of these different kinds of systems often brings up the desire for combative comparison, but I think that, in reality, there is clearly no need for combat because they are simply too different in the functions that they intend to provide.
My interpretation is that Google isn't really concerned with intentionally provided meta-data in the name of end-user, full-Web search because the scale that they operate on seems to render any such indexing by one or even a number of parties almost laughably shallow in its characterization of both the nature of any particular item and its expected relevance to a query. When you have literally millions of people passively voting and indexing every item of the Web through their decisions to link to it or not, you have very sophisticated algorithms for understanding the text in the pages generating and receiving those links and to top it off you record and process millions of people's behavior when faced with your search results, why should you care what some person or institution says the item is about? The fact that they (among other search engines) beat out the directory-based approach to finding information on the Web is a clear demonstration that automatic indexing and link based relevance ranking do a better job than meta-data based classification - for the problem of Web scale search. Google clearly doesn't need human semantic indexers to succeed, though, as Peter said, they certainly use all of the information that exists. If there happen to be good indexes online (as Connotea turned out to be be for a fairly brief window), then their algorithms will certainly find them and use them - if not, no worries, the algorithms will take advantage of the 'normal' data on the Web and do just fine thank you.
From the library-science professional perspective, this attitude is clearly annoying. If human indexing isn't really necessary to find things, what is the point of the field that has devoted itself to the creation of effective ways for people to categorize things for retrieval? For example, there is a lot of annoyance that the Google Books initiative seems to ignore most, if not all of the meta-data already associated with the books that they are scanning and indexing. This means that meta-data, even as basic as volume numbers, is inaccessible for searching. For the library professional that is trained to both search through and construct careful and precise classification structures,the inability to even search for a specific volume of a book is infuriating - particularly knowing that it is well within Google's power to incorporate such abilities into their system.
So on the one side we have the perspective that there is still value in the careful, intentional use of meta-data in the search and retrieval process while on the other side we are quite happy to let the intersection of algorithm and massive passive indexing do the work. I guess, as is the usual answer, I'd suggest that both sides provide useful functions that are both worth keeping and advancing. A detailed classification system, either constructed intentionally through the work of professional labor or semi-intentionally through the work of social taggers, provides functionality that is clearly different than what can be achieved by automatic indexing; however, it may not provide any help whatsoever in improving a full-Web-scale keyword-based search. The essence of the power of intentional classification is the precision of the queries that it enables. For example, if I want only version 3 of "The Devil's Rights and the Redemption" and thats it or I want only those items that have been tagged as with bioinformatics and to_read by Jaa, there is really no way ( AFAIK) to accomplish this without the intentional recording and utilization of meta-data about those resources.
So, though Peter and Google may have little direct use for ED and its semantic meta-data generating and consuming brethren emanating from the library and information sciences, there are still clearly meaningful applications of such work. It just happens that providing effective search over the contents of the entire Web based on a string like 'Britney Spears' isn't really one of them.
I'm ok with that.
Posted by Benjamin Good at 5:24 PM 0 comments
Labels: ED, google, indexing, information science, library science, scifoo, scifoo2008, search, semantic tagging, social semantic tagging
Subscribe to:
Posts (Atom)