While we (Andrew Su and I) like to talk about the successes of the Gene Wiki - articles like the one for Reelin that represent arguably the best consolidated body of text associated with the gene - there remain some rather glaring holes in its content. A couple months ago I had a look for under-developed articles linked to genes with extensive numbers of publications. With a small bit of hacking I uncovered a list of 2,553 genes that were linked directly to more than 20 PubMed citations (using NCBI's gene2pubmed) but had less than 100 words of text in their Gene Wiki article. (Up to the previous period, this post contained 105 words.) From this list I found 151 genes with more than 100 PubMed citations and less than 100 words of wiki text.
An example is the PIN1 gene. When the analysis was run, this gene was linked to 154 citations in PubMed yet had only 2 sentences in the Gene Wiki. So... how do we fill in these gaps? This is, of course, the fundamental question associated with wikis or any other attempt to harness community intelligence and there is no easy answer. One model that we are very interested in was pioneered by Alex Bateman and colleagues at the journal of RNA Biology. When hopeful authors submit an article about a new RNA family to the journal, it is a condition of publication that they contribute an article to Wikipedia about that family. Aside from being a generally good thing to do as far as sharing knowledge with the world, these articles are subsequently used to manage the annotations for RNA families in the Rfam database (e.g. snoZ107_R87). After a few years of operation, the Rfam team published an article that, among others things, celebrated the success of the Wikipedia connection. So, how might we expand upon this model to tackle the challenges facing the Gene Wiki?
The beauty of this approach is that it does not rely on any changes to the incentive system currently operational in science. Scientists need to publish in peer-reviewed journals. Rather than complaining about the inefficiency of this outdated process and suggesting social changes with no obvious way to achieve them, lets see what we can do to make the system work for us as it stands. Lets create a way for scientists to obtain real publications in real journals and have Gene Wiki article content generated as a natural part of the process. Here is one idea.
Friday, May 6, 2011
Integrating the Gene Wiki with traditional publishing?
Posted by Benjamin Good at 4:42 PM 4 comments
Labels: academic publishing, gene wiki, journals
Wednesday, October 17, 2007
Where is the API?
Yes.. I am procrastinating. I should be sleeping, working on the OntoLoki automatic ontology evaluation system, or preparing for our meeting with the SWAN team tomorrow morning; but instead, I am perusing Project Prospect and thinking about what a journal should look like. This is largely because of my disappointment in reading this nascent blog post in which Ian Mulvaney (a person who I think I respect and leader of a project I obviously find fascinating) suggests that enforcing the application of naming standards for chemical entities at the time of publication would a) be too hard for authors, b) not provide much benefit, c) that it would be better to let this be a voluntary step - all of which I absolutely disagree with.
This, and comments on the post, lead me to Project Prospect - which seems to be the first real publisher to take the idea of semantic enhancements of online manuscripts seriously.
Project Prospect provides semantic annotation (e.g. labeling GO terms etc. in manuscripts) and uses this to provide some enhanced navigation patterns and some additional information (e.g. definitions) for any of the annotations. Doing a pretty nice job at this was apparently enough to win them the 2007 ALPSP/Charlesworth Award for Publishing Innovation. While this is certainly a nice addition and a step in what I think is the right direction, it is 1) overwhelmingly similar to the much older and much more flexible, Conceptual Open Hypermedia ServicE (COHSE) from the University of Manchester and 2) does not seem to provide any capacity for semantic integration of the manuscripts in the collection.
Is this really the best we can do?
What I would like to see is a journal with an API. An API that would let me ask it questions like "what genes are present in articles published in this journal that contained both go:0005576, or any of its children , the word 'vaccine', and are described in the article as being upregulated". Right now, we can approach this sort of question with text-mining, but, with extensions to work like that done to enable the hypermedia browsing defined above (which fundamentally depends on solid entity identification and annotation within the document), this question (which spans multiple manuscripts) could be answered with a relatively straightforward query.
Its time for journals to step up an stop wasting talented researchers time writing text mining algorithms. Lets build a journal with a proper API, one with standards compliant methods for both writing content to it and querying the content inside it programmatically. Such a journal would not only improve human navigation and understanding of its independent textual documents, but would also enable entirely new modes of interaction with the integrated knowledge spanning all of its semantic content.
Posted by Benjamin Good at 12:15 AM 6 comments
Labels: academic publishing, AP, conceptual hypermedia, journals, nature, semantic web