Showing posts with label conceptual hypermedia. Show all posts
Showing posts with label conceptual hypermedia. Show all posts

Wednesday, October 17, 2007

Where is the API?

Yes.. I am procrastinating. I should be sleeping, working on the OntoLoki automatic ontology evaluation system, or preparing for our meeting with the SWAN team tomorrow morning; but instead, I am perusing Project Prospect and thinking about what a journal should look like. This is largely because of my disappointment in reading this nascent blog post in which Ian Mulvaney (a person who I think I respect and leader of a project I obviously find fascinating) suggests that enforcing the application of naming standards for chemical entities at the time of publication would a) be too hard for authors, b) not provide much benefit, c) that it would be better to let this be a voluntary step - all of which I absolutely disagree with.

This, and comments on the post, lead me to Project Prospect - which seems to be the first real publisher to take the idea of semantic enhancements of online manuscripts seriously.

Project Prospect provides semantic annotation (e.g. labeling GO terms etc. in manuscripts) and uses this to provide some enhanced navigation patterns and some additional information (e.g. definitions) for any of the annotations. Doing a pretty nice job at this was apparently enough to win them the 2007 ALPSP/Charlesworth Award for Publishing Innovation. While this is certainly a nice addition and a step in what I think is the right direction, it is 1) overwhelmingly similar to the much older and much more flexible, Conceptual Open Hypermedia ServicE (COHSE) from the University of Manchester and 2) does not seem to provide any capacity for semantic integration of the manuscripts in the collection.

Is this really the best we can do?

What I would like to see is a journal with an API. An API that would let me ask it questions like "what genes are present in articles published in this journal that contained both go:0005576, or any of its children , the word 'vaccine', and are described in the article as being upregulated". Right now, we can approach this sort of question with text-mining, but, with extensions to work like that done to enable the hypermedia browsing defined above (which fundamentally depends on solid entity identification and annotation within the document), this question (which spans multiple manuscripts) could be answered with a relatively straightforward query.

Its time for journals to step up an stop wasting talented researchers time writing text mining algorithms. Lets build a journal with a proper API, one with standards compliant methods for both writing content to it and querying the content inside it programmatically. Such a journal would not only improve human navigation and understanding of its independent textual documents, but would also enable entirely new modes of interaction with the integrated knowledge spanning all of its semantic content.