Showing posts with label publication. Show all posts
Showing posts with label publication. Show all posts

Monday, January 2, 2012

Scientific Games in Genome Biology

Happy New year everyone! In case you are looking for some inspirational reading to start off 2012, Andrew and I wrote a Research Highlight for Genome Biology on Games With a Scientific Purpose.  Kudos to Andrew for convincing them to make it open access so you can actually read it...

Research highlight      Games with a scientific purpose Good BM, Su AI
Genome Biology 2011, 12:135 (28 December 2011)
[Abstract] [Full text] [PDF]

Thursday, December 15, 2011

Mining the Gene Wiki

Our article about mining ontology-based gene annotations from the text of the Gene Wiki just came out at BMC Genomics.  Yay!

In the article, we discuss the results of what I think might be the simplest text-mining strategy that could possibly work.  Based on the premise that each Gene Wiki article is fundamentally about one particular gene, we make the simplifying assumption that all of the concepts detectable in the article are descriptors of what that gene does.  With those assumptions in place, we use the NCBO annotator to detect concepts from the Gene Ontology (GO) and the Human Disease Ontology (DO) in the text of articles about genes.  Each detected occurrence thus produces a candidate annotation for the gene.  From the article:

For example, we identified the GO term ‘embryonic development (GO:0009790)’ in the text of the article on the DAX1 gene: “DAX1 controls the activity of certain genes in the cells that form these tissues during embryonic development”.  From this occurrence, our system proposed the structured annotation ‘DAX1 participates in the biological process of embryonic development’.  Following the same pattern, we found a potential annotation to the DO term ‘Congenital Adrenal Hypoplasia’ (DOID:10492) in the sentence: “Mutations in this gene result in both X-linked congenital adrenal hypoplasia and hypogonadotropic
hypogonadism”.
We found that, in terms of precision, this simple approach worked pretty well on detecting gene-disease  annotations (90-93%) but not nearly as well at detecting gene-function (GO) annotations (48-64%).  As you might expect, the recall equation worked in the opposite direction with many more potential GO annotations discovered (11,022) then DO annotations (2,983).  Though there was some overlap, the majority of the predicted annotations did not have any match in existing annotation databases, showing that the Gene Wiki contains some knowledge that centralized resources like the Gene Ontology Annotation database do not yet represent and that basic text mining provides a way to access that knowledge computationally.

But, you say, that precision for the GO is really low, what use is this really?  For applications that require 100% accuracy, like a curated database, well you would need to curate the predicted results and that might be quite a lot faster than searching through PubMed to find them all from scratch.  As it turns out, there are also other kinds of applications that can take advantage of data like this that has noise in it.  As long as there is a strong signal within the noise, probabilistic techniques, like enrichment analysis, can work.  This is possible because, although many of the individual annotations might turn out to be incorrect, as a group they are far far from random.

For more details, read the paper ;).

Friday, November 11, 2011

Gene Wiki article out today at NAR

The articles for the annual database issue are starting to appear in the NAR collection.  My favorite one this year is, immodestly perhaps, ours about the Gene Wiki!  The simple message here is that the Gene Wiki is continuing to grow and that the content remains very high quality overall.  For more information, the abstract is below, and of course the paper is freely accessible online:

"The Gene Wiki is an open-access and openly editable collection of Wikipedia articles about human genes. Initiated in 2008, it has grown to include articles about more than 10 000 genes that, collectively, contain more than 1.4 million words of gene-centric text with extensive citations back to the primary scientific literature. This growing body of useful, gene-centric content is the result of the work of thousands of individuals throughout the scientific community. Here, we describe recent improvements to the automated system that keeps the structured data presented on Gene Wiki articles in sync with the data from trusted primary databases. We also describe the expanding contents, editors and users of the Gene Wiki. Finally, we introduce a new automated system, called WikiTrust, which can effectively compute the quality of Wikipedia articles, including Gene Wiki articles, at the word level. All articles in the Gene Wiki can be freely accessed and edited at Wikipedia, and additional links and information can be found at the project's Wikipedia portal page:http://en.wikipedia.org/wiki/Portal:Gene_Wiki."

Monday, February 11, 2008

another Scientific American article on the Semantic Web

 

A follow-up to the much cited Scientific American article on the Semantic Web has recently come out.  Would love to say more about it, but I can't find a free copy at the moment..  

Links related to the new article now:
If anyone has a copy, perhaps you could bravely post it somewhere and comment a link to it?

Its both funny and I think instructive that the number one hit in Google Scholar when searching for "semantic web" is to the first (2001) article in this now two-part series, but that, out of the 36 versions of that article that are listed, none of them link to Scientific American in any way!  Seeing that the article has been cited more then 5,000 times, it seems that S.A. could probably have made a quite a bit of money one way or another if they had decided to host a free public version from the outset and remained the main provider of that information.



Wednesday, February 14, 2007

iCAPTURER1 paper accepted

First knowledge gardening paper accepted in September of 2005

Good B,Tranfield EM,Tan PC, Shehata M, Singhera GK, Gosselink J, Okon EB, Wilkinson
Fast, Cheap, and Out of Control: A Zero Curation Model for Ontology Development
at the Pacific Symposium on Biocomputing 2006.

See here for PDF