Happy New year everyone! In case you are looking for some inspirational reading to start off 2012, Andrew and I wrote a Research Highlight for Genome Biology on Games With a Scientific Purpose. Kudos to Andrew for convincing them to make it open access so you can actually read it...
Research highlight Games with a scientific purpose Good BM, Su AI
Genome Biology 2011, 12:135 (28 December 2011)
[Abstract] [Full text] [PDF]
Monday, January 2, 2012
Scientific Games in Genome Biology
Posted by Benjamin Good at 7:00 AM 0 comments
Labels: foldit, games, genome biology, gwaps, publication, sulab
Thursday, December 15, 2011
Mining the Gene Wiki
Our article about mining ontology-based gene annotations from the text of the Gene Wiki just came out at BMC Genomics. Yay!
In the article, we discuss the results of what I think might be the simplest text-mining strategy that could possibly work. Based on the premise that each Gene Wiki article is fundamentally about one particular gene, we make the simplifying assumption that all of the concepts detectable in the article are descriptors of what that gene does. With those assumptions in place, we use the NCBO annotator to detect concepts from the Gene Ontology (GO) and the Human Disease Ontology (DO) in the text of articles about genes. Each detected occurrence thus produces a candidate annotation for the gene. From the article:
For example, we identified the GO term ‘embryonic development (GO:0009790)’ in the text of the article on the DAX1 gene: “DAX1 controls the activity of certain genes in the cells that form these tissues during embryonic development”. From this occurrence, our system proposed the structured annotation ‘DAX1 participates in the biological process of embryonic development’. Following the same pattern, we found a potential annotation to the DO term ‘Congenital Adrenal Hypoplasia’ (DOID:10492) in the sentence: “Mutations in this gene result in both X-linked congenital adrenal hypoplasia and hypogonadotropicWe found that, in terms of precision, this simple approach worked pretty well on detecting gene-disease annotations (90-93%) but not nearly as well at detecting gene-function (GO) annotations (48-64%). As you might expect, the recall equation worked in the opposite direction with many more potential GO annotations discovered (11,022) then DO annotations (2,983). Though there was some overlap, the majority of the predicted annotations did not have any match in existing annotation databases, showing that the Gene Wiki contains some knowledge that centralized resources like the Gene Ontology Annotation database do not yet represent and that basic text mining provides a way to access that knowledge computationally.
hypogonadism”.
But, you say, that precision for the GO is really low, what use is this really? For applications that require 100% accuracy, like a curated database, well you would need to curate the predicted results and that might be quite a lot faster than searching through PubMed to find them all from scratch. As it turns out, there are also other kinds of applications that can take advantage of data like this that has noise in it. As long as there is a strong signal within the noise, probabilistic techniques, like enrichment analysis, can work. This is possible because, although many of the individual annotations might turn out to be incorrect, as a group they are far far from random.
For more details, read the paper ;).
Posted by Benjamin Good at 10:20 AM 0 comments
Labels: disease ontology, gene ontology, gene wiki, genomics, publication, sulab, text-mining
Friday, November 11, 2011
Gene Wiki article out today at NAR
The articles for the annual database issue are starting to appear in the NAR collection. My favorite one this year is, immodestly perhaps, ours about the Gene Wiki! The simple message here is that the Gene Wiki is continuing to grow and that the content remains very high quality overall. For more information, the abstract is below, and of course the paper is freely accessible online:
"The Gene Wiki is an open-access and openly editable collection of Wikipedia articles about human genes. Initiated in 2008, it has grown to include articles about more than 10 000 genes that, collectively, contain more than 1.4 million words of gene-centric text with extensive citations back to the primary scientific literature. This growing body of useful, gene-centric content is the result of the work of thousands of individuals throughout the scientific community. Here, we describe recent improvements to the automated system that keeps the structured data presented on Gene Wiki articles in sync with the data from trusted primary databases. We also describe the expanding contents, editors and users of the Gene Wiki. Finally, we introduce a new automated system, called WikiTrust, which can effectively compute the quality of Wikipedia articles, including Gene Wiki articles, at the word level. All articles in the Gene Wiki can be freely accessed and edited at Wikipedia, and additional links and information can be found at the project's Wikipedia portal page:http://en.wikipedia.org/wiki/Portal:Gene_Wiki."
Posted by Benjamin Good at 11:36 AM 0 comments
Labels: database, gene wiki, NAR, publication, sulab
Monday, February 11, 2008
another Scientific American article on the Semantic Web
Posted by Benjamin Good at 9:32 AM 0 comments
Labels: "scientific american", academic publishing, publication, semantic web
Wednesday, February 14, 2007
iCAPTURER1 paper accepted
First knowledge gardening paper accepted in September of 2005
Good B,Tranfield EM,Tan PC, Shehata M, Singhera GK, Gosselink J, Okon EB, Wilkinson
Fast, Cheap, and Out of Control: A Zero Curation Model for Ontology Development at the Pacific Symposium on Biocomputing 2006.
See here for PDF
Posted by Benjamin Good at 1:25 PM 0 comments
Labels: icapturer, mass collaboration, mechanical turk, ontology, publication, semantic web, volunteer