Showing posts with label genomics. Show all posts
Showing posts with label genomics. Show all posts

Thursday, December 15, 2011

Mining the Gene Wiki

Our article about mining ontology-based gene annotations from the text of the Gene Wiki just came out at BMC Genomics.  Yay!

In the article, we discuss the results of what I think might be the simplest text-mining strategy that could possibly work.  Based on the premise that each Gene Wiki article is fundamentally about one particular gene, we make the simplifying assumption that all of the concepts detectable in the article are descriptors of what that gene does.  With those assumptions in place, we use the NCBO annotator to detect concepts from the Gene Ontology (GO) and the Human Disease Ontology (DO) in the text of articles about genes.  Each detected occurrence thus produces a candidate annotation for the gene.  From the article:

For example, we identified the GO term ‘embryonic development (GO:0009790)’ in the text of the article on the DAX1 gene: “DAX1 controls the activity of certain genes in the cells that form these tissues during embryonic development”.  From this occurrence, our system proposed the structured annotation ‘DAX1 participates in the biological process of embryonic development’.  Following the same pattern, we found a potential annotation to the DO term ‘Congenital Adrenal Hypoplasia’ (DOID:10492) in the sentence: “Mutations in this gene result in both X-linked congenital adrenal hypoplasia and hypogonadotropic
hypogonadism”.
We found that, in terms of precision, this simple approach worked pretty well on detecting gene-disease  annotations (90-93%) but not nearly as well at detecting gene-function (GO) annotations (48-64%).  As you might expect, the recall equation worked in the opposite direction with many more potential GO annotations discovered (11,022) then DO annotations (2,983).  Though there was some overlap, the majority of the predicted annotations did not have any match in existing annotation databases, showing that the Gene Wiki contains some knowledge that centralized resources like the Gene Ontology Annotation database do not yet represent and that basic text mining provides a way to access that knowledge computationally.

But, you say, that precision for the GO is really low, what use is this really?  For applications that require 100% accuracy, like a curated database, well you would need to curate the predicted results and that might be quite a lot faster than searching through PubMed to find them all from scratch.  As it turns out, there are also other kinds of applications that can take advantage of data like this that has noise in it.  As long as there is a strong signal within the noise, probabilistic techniques, like enrichment analysis, can work.  This is possible because, although many of the individual annotations might turn out to be incorrect, as a group they are far far from random.

For more details, read the paper ;).

Tuesday, December 21, 2010

About the Code Delusion


After randomly stumbling across a link to the article “Getting over the Code Delusion” in a link posted by Deepak Sing a couple months ago I’ve just spent the last hour reading it instead of getting my own work done and now I’m spending more time writing about it.  Sorry boss, but I couldn't help myself... 
The central point that the article makes with surprisingly elegant prose is that it is impossible to understand how life works entirely through the lens of a linear DNA sequence.  Put very simply, life is much more complicated than we had hoped – read his piece to get a better understanding of some of the mechanisms of that complexity. 
While I generally found the piece informative, enlightening and a pleasure to read, one element that I found unnecessary was a thread of what seem like bizarrely placed calls to some higher power.  For example,
 “When you encounter the meaningful, directed, and well-shaped movements of a dance, it’s hard to ignore the active principle — some would say the agency or being — coordinating the movements.”
 And close after,
 “Seemingly in the grip of the encircling DNA with its relatively fixed and stable structure, yet responsive to the varying flow of life around it, the nucleosome holds the balance between gene and context — a task ­requiring flexibility, a ‘sense’ of appropriate rhythm, and perhaps we could even say ‘grace.’”
These, perhaps unscientific, references seem to arise in the text because the complexity of the totality of the factors that affect how cells and organisms function appears too great to hope to understand in a reductive way.  Along the same lines he suggests: 
"The search for precise explanatory mechanisms and codes leads us along a path of least resistance toward the reduction of understanding."
Now what does that mean for a scientist?  How can we increase understanding without searching for precise explanations? 
It puts me in mind of the unlimited complexity produced by simple underlying rules in Wolfram's "A new kind of science" - would love to hear what it puts in your mind.

Saturday, March 29, 2008

not so personal genomics

So my idea of requesting sponsorship for getting my 23andme results processed and sharing the results was already carried out successfully by some one named Andrew Meyer. Follow his progress on his blog .  


Mike Arrington also has a nice post about his decision to give it a try and his fears about what it might tell him (and other people) about himself.  I think one of the reasons he's been such a successful blogger is how incredibly quotable his writing is.  To sum up why, despite the list of fears he discusses, he went ahead and sent in his spit, he says:

"The future is coming, and I want to know"

My sentiments exactly.

Friday, March 28, 2008

I'll show you mine...


23andme is giving a whole new meaning to the profile page of the social network.  Now you can share your genetic information as well as your relationship status and your interests in underwater basket weaving.  If I had a spare $999.00 I would love to try this out.  Anyone want to sponsor it?  

It brings us closer to the time when my (much mocked) relatedness party game becomes a reality.  The game would be to guess the person at the party that was either the most distant or the closest genetically to you - and then use the (I assume) forthcoming Facebook/23andme application to find out who won.  You could also do celebrity variations and so forth.

So, if I show you mine, will you ???


Tuesday, June 26, 2007

(Very) personalized medicine (everything)

You have to love their tagline at the top of their website,

don't panic, we're here to help

So, do you trust Google/23andMe to find and hold all of your genetic secrets? I'd trust them a lot more than Joe Shmoe in I.T... Seriously, if you are afraid Google is setting up some secret clone army bent on world domination, your genetic information really isn't going to help them much anyway. The truth is, they don't really have any incentive to do anything malicious with your personal information. Their famous moniker "don't be evil" isn't just could p.r., its good business - and thats why I actually do trust them. Not too mention that if anyone could actually keep your information safe from those that would do evil with it (perhaps certain governments..) they are about the only group in the world that might stand a chance.

I absolutely can not wait until I can have my whole genome sequenced and can start playing with it. I want to see specifically what differences exist between me and my sister, my cat, my dog, and my plant. I want to know what pills I should take when, if I should be getting tested sooner for prostate cancer, and whether I am more closely related to Socrates or to Gandhi. I want to go to a party and play the 4th cousin game.. who in the room is my closest relative? I want a t-shirt with the damn thing on it!

Never fear knowledge.