If there was any previous doubt about whether 2010 will be known as the year that the Semantic Web came of age, this acquisition seals the deal.
I wonder if my life would be different if I had listened to my wife's advice and walked up to Danny Ayers at sci foo in 2008 and asked for a job... His three word intro "just, an, engineer" still inspires me.
Friday, July 16, 2010
Google buys Metaweb
Tuesday, April 6, 2010
Jerzy Lewak at SDSW Meetup
Image via Wikipedia
- Hmm, why don't I set up some nice hierarchies of concepts to place my emails into so it will be easier to find them later?
- Darn.. that really isn't working out very well. The more data I get, the harder it is to organize and I keep running into the problem that almost every single item in my collection might be placed under more than one category. Perhaps a physical filing cabinet is actually a terrible thing to base a completely virtual information storage and retrieval system on... (Though he didn't bring it up, he was describing exactly the same thing that Clay Shirky got so excited about in the 'Ontology is Overrated' essay in 2005 that - in turn - got me all excited in 2006, except of course Jerzy was thinking in 1991. And, as pointed out to me by my LIS friend Joe, this basic problem and the following conclusion were pretty well fleshed out by Ranganathan in the 1930's...).
- After experimenting with plain content search (a la Google) he arrived at faceted classification as the most powerful and flexible way to describe and access data.
- Choose some database field like 'last name' and type in a name like 'Smith'
- Show two things - one, the number of results went down and two, the number of possible values for the other fields (e.g. 'first name', 'height', 'date', etc.) would immediately be constrained to only show values for objects linked to Smiths in the database.
The system works on both structured and unstructured data (he showed a quick example of newspaper articles) but requires fairly heavy manual labor to get it started. Overall it looked like it would be useful and fun to use and I can imagine many potential directions they could take it.
My only complaint for the talk was that there was absolutely zero Web in it - no mention of any native ability to consume or produce RDF, no mention of OWL, no discussion of scaling possibilities, and the proverbial elephant in the room of large-scale data access was left more or less untouched. I guess that might be one sign of a good talk - I got interested and it left me thirsting for more..
Posted by Benjamin Good at 9:50 PM 1 comments
Labels: Clay Shirky, google, Jerzy Lewak, Knowledge Management, Knowledge Representation, meetup, Ontologies, san diego, semantic web, User interface
Sunday, August 31, 2008
Peter (Google) and Christine (the librarians)
Posted by Benjamin Good at 5:24 PM 0 comments
Labels: ED, google, indexing, information science, library science, scifoo, scifoo2008, search, semantic tagging, social semantic tagging
Friday, June 13, 2008
IndentationError
Today was a good day. I made absolutely no progress on my research and did not otherwise move myself any closer to graduation. But, today, that is ok , because, today, I was not actually trying to do either. Today, frustration at delays from a collaborator, burning curiosity, and the need to devise a Father's Day present, practically forced me to take the day and play with the Google App Engine. Apologies to my supervisor..
- Python is really much easier to learn from scratch and by example than javascript (thankfully)
- In fact, Python is actually pretty cool, despite having the rather surprising "feature" of interpreting the amount of indentation in the code as having meaning (that was a surprise, but easily adapted to)
- I was already pretty convinced from the videos of the Google developer's meeting, but now its really just bleedingly obvious that this is how the next generation of web application is going to be born
Posted by Benjamin Good at 12:04 AM 0 comments
Labels: appengine, google, luddite, procrastination, python
Wednesday, January 23, 2008
centralized content and decentralized control
No time for depth of thought here, just wanted to quickly jot down some current observations and recent links to ideas related to the future of the semantic Web.
- Early January 2008, the annual NAR database issue comes out listing more than 1000 distinct databases in molecular biology.
- Shortly after, Duncan Hull complains that essentially none of the "dark data" residing in these databases will ever be used because of the near complete lack of both syntactic and semantic interoperability between these isolated, decentralized silos.
- I participate in writing up a paper about the Banff Manifesto, which, among other things seeks to improve cross-database interoperability through the introduction of a single, open, centralized, resource for defining public namespaces for use in the construction of unique identifiers.
- I am forwarded a link to a Wired blog post about Google base - Google providing a centralized repository for scientific data.
- the Freebase dev blog quotes an article about Wine tasting websites that indicates that Freebase is an example of the semantic Web (which they also refer to as Web3.0) and that everyone that creates a community-directed database should be using Freebase to do so - because of the dramatic advantages of centralization.
Not as far as I can tell right now.
Posted by Benjamin Good at 12:39 PM 0 comments
Labels: centralization, control, duncan, freebase, google, semantic web, wine
Thursday, January 17, 2008
Extreme Writing for ISMB
"... a grander vision of integrative bioinformatics. In this vision, researchers or their computational agents not only discover and access all the databases that they require, but also clearly understand how each entity in each database relates to the entities in all the others. The information required to realize this vision can be conceptualized as a map that, rather than describing the interconnectivity of points in physical space, describes the interconnectivity of biological entities in the context of the Web..."
"...Right now, this meta-database makes it possible to answer queries about the connectivity of the various data sources but does not yet enable queries of the connectivity of their individual components. Before it is possible to 'zoom in' in on the lowest level entities in the global map of bioinformatics data, it is first necessary to establish a consistent strategy for their identification. The bioinformatics meta-database presented here provides a centralized, community-governed repository of public namespaces that we propose might serve as the foundation of a global unique identifier system for bioinformatics on the Web. This identifier system is based on principles outlined in the Banff Manifesto (BM) instigated at the 2007 World Wide Web conference and presented here for the first time..."
In the end, the paper really could have been much a better with a some more time to put it together. That being said, it does have some interesting ideas in it and so might just squeak into the conference. If it does, Francois and co. will certainly produce a very exciting presentation by the time the conference finally comes to pass. If not, we'll try try again.
Posted by Benjamin Good at 11:39 AM 0 comments
Labels: Banff Manifesto, collaboration, Extreme Writing, freebase, google, ISMB, map, public namespaces, social locking, Web mapping, writing
Friday, November 9, 2007
Google Gadget
Testing labmate Byron Kuo's Google Gadget (iPubCloud)
Type in something likely to show up in a PubMed query like: Good BM
Tuesday, June 26, 2007
(Very) personalized medicine (everything)
You have to love their tagline at the top of their website,
don't panic, we're here to help
So, do you trust Google/23andMe to find and hold all of your genetic secrets? I'd trust them a lot more than Joe Shmoe in I.T... Seriously, if you are afraid Google is setting up some secret clone army bent on world domination, your genetic information really isn't going to help them much anyway. The truth is, they don't really have any incentive to do anything malicious with your personal information. Their famous moniker "don't be evil" isn't just could p.r., its good business - and thats why I actually do trust them. Not too mention that if anyone could actually keep your information safe from those that would do evil with it (perhaps certain governments..) they are about the only group in the world that might stand a chance.
I absolutely can not wait until I can have my whole genome sequenced and can start playing with it. I want to see specifically what differences exist between me and my sister, my cat, my dog, and my plant. I want to know what pills I should take when, if I should be getting tested sooner for prostate cancer, and whether I am more closely related to Socrates or to Gandhi. I want to go to a party and play the 4th cousin game.. who in the room is my closest relative? I want a t-shirt with the damn thing on it!
Never fear knowledge.
Posted by Benjamin Good at 10:15 AM 2 comments
Labels: 23andme, fear, genetics, genome, genomics, google, personal information, personlized medicine, politics, trust