Showing posts with label google. Show all posts
Showing posts with label google. Show all posts

Friday, July 16, 2010

Google buys Metaweb

If there was any previous doubt about whether 2010 will be known as the year that the Semantic Web came of age, this acquisition seals the deal.

I wonder if my life would be different if I had listened to my wife's advice and walked up to Danny Ayers at sci foo in 2008 and asked for a job... His three word intro "just, an, engineer" still inspires me.

Well done Danny and friends!

Tuesday, April 6, 2010

Jerzy Lewak at SDSW Meetup

Two tall metal file cabinets for work or home useImage via Wikipedia

I just got back from another interesting San Diego Semantic Web Meetup, here are my notes before I forget.

The presenter this evening was Jerzy Lewak, a former theoretical physicist, a professor emeritus at UCSD, and cofounder of (at least) Nisus Software and SpeedTrack Inc.. (Once again, I continue to be impressed at the level of the speakers that are recruited for these events!). Jerzy presented the history and current applications of a novel human interface for databases that he calls GIA (Guided Information Access) that is made possible by the underlying TIE (Technology for Information Engineering) framework.

Jerzy began his talk as a true computer scientist by defining 'The Problem' that originally inspired this work (back in 1991). At that time email was just starting to pick up steam and already the task of finding old emails was becoming unmanageable. So, he set out to find a better way. (Always nice to be working on a solving a problem for yourself.) His search for a solution took a decidedly familiar path:
  1. Hmm, why don't I set up some nice hierarchies of concepts to place my emails into so it will be easier to find them later?
  2. Darn.. that really isn't working out very well. The more data I get, the harder it is to organize and I keep running into the problem that almost every single item in my collection might be placed under more than one category. Perhaps a physical filing cabinet is actually a terrible thing to base a completely virtual information storage and retrieval system on... (Though he didn't bring it up, he was describing exactly the same thing that Clay Shirky got so excited about in the 'Ontology is Overrated' essay in 2005 that - in turn - got me all excited in 2006, except of course Jerzy was thinking in 1991. And, as pointed out to me by my LIS friend Joe, this basic problem and the following conclusion were pretty well fleshed out by Ranganathan in the 1930's...).
  3. After experimenting with plain content search (a la Google) he arrived at faceted classification as the most powerful and flexible way to describe and access data.
So far so good, I am paying attention.

Now the problem arises that many combinations of facets actually produce zero results. As he noted, just 200 facets (he uses the word 'selectors') is enough to uniquely describe every particle in the universe. This became the real problem. The breakthrough that got him going and has led to SpeedTrack and everything else he presented was the idea of dynamically limiting the potential facets based on those that are already selected. In the interfaces he demonstrated, he would:
  1. Choose some database field like 'last name' and type in a name like 'Smith'
  2. Show two things - one, the number of results went down and two, the number of possible values for the other fields (e.g. 'first name', 'height', 'date', etc.) would immediately be constrained to only show values for objects linked to Smiths in the database.
His interfaces might be described as an advanced multi-parameter type-ahead. They work by guiding the user (GIA) in the creation of potentially very complex queries that are guaranteed to return results. This is achieved by dynamically exposing the underlying indexes that drive boolean queries.

The system works on both structured and unstructured data (he showed a quick example of newspaper articles) but requires fairly heavy manual labor to get it started. Overall it looked like it would be useful and fun to use and I can imagine many potential directions they could take it.

My only complaint for the talk was that there was absolutely zero Web in it - no mention of any native ability to consume or produce RDF, no mention of OWL, no discussion of scaling possibilities, and the proverbial elephant in the room of large-scale data access was left more or less untouched. I guess that might be one sign of a good talk - I got interested and it left me thirsting for more..
Reblog this post [with Zemanta]

Sunday, August 31, 2008

Peter (Google) and Christine (the librarians)

I was very lucky to have my new wife with me at SciFoo for many reasons, not the least of which is that she is much better at socializing than I am and thus managed to introduce me to many people I would never normally have met.  One of those people was Christine Borgman, Professor & Presidential Chair in Information Studies at UCLA.  While I was struggling to explain my work on the Entity Describer project to her, she noticed Peter Norvig walk by and dragged him over to join our conversation.  I guess she must have known him from somewhere but I'm not sure where.  Anyway, I didn't realize this at the time, but Peter is head of research at Google.  Ahem... did I mention that SciFoo interactions could be somewhat intimidating?  So there I am, standing between two giants of modern information science trying to explain what it was I was doing there but mostly trying to get some insight into their respective thoughts on the organization of the world's information. It wasn't a long discussion, but here are the basics.
The main question that I posed to them was whether or not and how semantic tagging (a la ED) is or might be useful.  On the surface, the answer from Peter was no and the answer from Christine was yes.  However, the truth of the matter is that they were really talking about supporting different functions for the end user - though this fundamental difference became a little lost during the conversation.  Google is principally focused on providing the best possible results, to the most people, given the least amount of information in the query - that is, keyword based search of the entire Web.  Library-science is typically much more concerned with providing the capacity for people to make very specific requests using much more sophisticated queries that operate over much smaller collections of information (e.g. the library of congress). The fact that there is some overlap in the information needs of the users of these different kinds of systems often brings up the desire for combative  comparison, but I think that, in reality, there is clearly no need for combat because they are simply too different in the functions that they intend to provide.
My interpretation is that Google isn't really concerned with intentionally provided meta-data in the name of end-user, full-Web search because the scale that they operate on seems to render any such indexing by one or even a number of parties almost laughably shallow in its characterization of both the nature of any particular item and its expected relevance to a query.  When you have literally millions of people passively voting and indexing every item of the Web through their decisions to link to it or not, you have very sophisticated algorithms for understanding the text in the pages generating and receiving those links and to top it off you record and process millions of people's behavior when faced with your search results, why should you care what some person or institution says the item is about?  The fact that they (among other search engines) beat out the directory-based approach to finding information on the Web is a clear demonstration that automatic indexing and link based relevance ranking do a better job than meta-data based classification - for the problem of Web scale search.  Google clearly doesn't need human semantic indexers to succeed, though, as Peter said, they certainly use all of the information that exists.  If there happen to be good indexes online (as Connotea turned out to be be for a fairly brief window), then their algorithms will certainly find them and use them - if not, no worries, the algorithms will take advantage of the 'normal' data on the Web and do just fine thank you. 
From the library-science professional perspective, this attitude is clearly annoying.  If human indexing isn't really necessary to find things, what is the point of the field that has devoted itself to the creation of effective ways for people to categorize things for retrieval?  For example, there is a lot of annoyance that the Google Books initiative seems to ignore most, if not all of the meta-data already associated with the books that they are scanning and indexing.  This means that meta-data, even as basic as volume numbers, is inaccessible for searching.  For the library professional that is trained to both search through and construct careful and precise classification structures,the inability to even search for a specific volume of a book is infuriating - particularly knowing that it is well within Google's power to incorporate such abilities into their system.
So on the one side we have the perspective that there is still value in the careful, intentional use of meta-data in the search and retrieval process while on the other side we are quite happy to let the intersection of algorithm and massive passive indexing do the work.  I guess, as is the usual answer, I'd suggest that both sides provide useful functions that are both worth keeping and advancing.  A detailed classification system, either constructed intentionally through the work of professional labor or semi-intentionally through the work of social taggers, provides functionality that is clearly different than what can be achieved by automatic indexing; however, it may not provide any help whatsoever in improving a full-Web-scale keyword-based search.  The essence of the power of intentional classification is the precision of the queries that it enables.  For example, if I want only version 3 of "The Devil's Rights and the Redemption" and thats it or I want only those items that have been tagged as with bioinformatics and to_read by Jaa, there is really no way ( AFAIK) to accomplish this without the intentional recording and utilization of meta-data about those resources.
So, though Peter and Google may have little direct use for ED and its semantic meta-data generating and consuming brethren emanating from the library and information sciences, there are still clearly meaningful applications of such work.  It just happens that providing effective search over the contents of the entire Web based on a string like 'Britney Spears' isn't really one of them.
I'm ok with that.

Friday, June 13, 2008

IndentationError

Today was a good day. I made absolutely no progress on my research and did not otherwise move myself any closer to graduation. But, today, that is ok , because, today, I was not actually trying to do either. Today, frustration at delays from a collaborator, burning curiosity, and the need to devise a Father's Day present, practically forced me to take the day and play with the Google App Engine. Apologies to my supervisor..


What I learned from my explorations today:
  • Python is really much easier to learn from scratch and by example than javascript (thankfully)
  • In fact, Python is actually pretty cool, despite having the rather surprising "feature" of interpreting the amount of indentation in the code as having meaning (that was a surprise, but easily adapted to)
  • I was already pretty convinced from the videos of the Google developer's meeting, but now its really just bleedingly obvious that this is how the next generation of web application is going to be born
I'd show you what I did... but Google currently has added the requirement of the possession of an SMS receiving phone to authenticate appengine accounts for deployments and, surprise!, I don't actually have one. I bet the combination of interest in developing with the appengine and lack of a mobile phone puts me in a very small group of people.. (I'm not a luddite, but mobile rates are ridiculous in Canada - and we don't even have legal iPhones yet).

Wednesday, January 23, 2008

centralized content and decentralized control

No time for depth of thought here, just wanted to quickly jot down some current observations and recent links to ideas related to the future of the semantic Web.

In chronological order
  1. Early January 2008, the annual NAR database issue comes out listing more than 1000 distinct databases in molecular biology.
  2. Shortly after, Duncan Hull complains that essentially none of the "dark data" residing in these databases will ever be used because of the near complete lack of both syntactic and semantic interoperability between these isolated, decentralized silos.
  3. I participate in writing up a paper about the Banff Manifesto, which, among other things seeks to improve cross-database interoperability through the introduction of a single, open, centralized, resource for defining public namespaces for use in the construction of unique identifiers.
  4. I am forwarded a link to a Wired blog post about Google base - Google providing a centralized repository for scientific data.
  5. the Freebase dev blog quotes an article about Wine tasting websites that indicates that Freebase is an example of the semantic Web (which they also refer to as Web3.0) and that everyone that creates a community-directed database should be using Freebase to do so - because of the dramatic advantages of centralization.
Notice anything interesting here?

It seems that many people think that success in achieving the goals of the semantic Web is really all about centralization of content. This of course, seems to smack against the decentralized roots of the first Web and is, in that way, rather disappointing. On the other hand, it also seems that success on the semantic Web is really all about decentralization of content control - which was, IMHO, the fundamental characteristic of the first Web that made it such a success (providing it with the coveted capital letter).

Is it possible to achieve a useful giant global graph without centralizing content?

Not as far as I can tell right now.

In fact, the current Web would hardly be very useful without massive efforts to centralize its content. How useful would the Web be if Google and others stopped downloading it and indexing its content on their servers, thus depriving us all of our god given right to find out anything about anything in milliseconds??









Thursday, January 17, 2008

Extreme Writing for ISMB

Over the last 48 hours or so I participated in a Google-Doc enabled Extreme Writing session as the lead (English-speaking) author of an article submitted to ISMB. The article describes a new area of the Freebase knowledge space meant, in addition to quite a number of other things, to capture knowledge about the interconnectivity of different biological databases. (See the temporary development version in the Sandbox or the eventually more permanent URI). The idea for this project, which has evolved a bit from the description in the Wiki available today, and the actual work of building this resource came almost entirely from Francois Belleau of the Bio2RDF project - he invited me to help with the writing and to add my two cents about knowledge gardening on the Web, a subject I am supposed to be becoming an expert in somehow.
Here are a couple snippets from the paper:
"... a grander vision of integrative bioinformatics. In this vision, researchers or their computational agents not only discover and access all the databases that they require, but also clearly understand how each entity in each database relates to the entities in all the others. The information required to realize this vision can be conceptualized as a map that, rather than describing the interconnectivity of points in physical space, describes the interconnectivity of biological entities in the context of the Web..."
"...Right now, this meta-database makes it possible to answer queries about the connectivity of the various data sources but does not yet enable queries of the connectivity of their individual components. Before it is possible to 'zoom in' in on the lowest level entities in the global map of bioinformatics data, it is first necessary to establish a consistent strategy for their identification. The bioinformatics meta-database presented here provides a centralized, community-governed repository of public namespaces that we propose might serve as the foundation of a global unique identifier system for bioinformatics on the Web. This identifier system is based on principles outlined in the Banff Manifesto (BM) instigated at the 2007 World Wide Web conference and presented here for the first time..."
It was an exciting, intriguing, sometimes frustrating experience simultaneously co-authoring this document with 5 other people (Francois, Michel Dumontier, Mark Wilkinson, Marc-Alexandre, Jamie Taylor). Overall, Google docs handled the experience surprisingly well. A quick look indicates that it captured 5925 revisions over the past three days. My only major complaint with it was that the screen tended to jump all over the place when lots of people working at the same time. I guess this may be one mechanism they use to keep people from editing in the same text area before its had a chance to sync everyone up for that area (otherwise it wouldn't know what to display). To avoid this problem to some extent and to make the work a bit more efficient we evolved strategies for social locking of different sections of the document. It would be great if support for this could be added to G-Doc - for example, a user could select a segment of text and request a lock on it that was enforced by the Google system.

In the end, the paper really could have been much a better with a some more time to put it together. That being said, it does have some interesting ideas in it and so might just squeak into the conference. If it does, Francois and co. will certainly produce a very exciting presentation by the time the conference finally comes to pass. If not, we'll try try again.









Friday, November 9, 2007

Google Gadget

Testing labmate Byron Kuo's Google Gadget (iPubCloud)

Type in something likely to show up in a PubMed query like: Good BM

Tuesday, June 26, 2007

(Very) personalized medicine (everything)

You have to love their tagline at the top of their website,

don't panic, we're here to help

So, do you trust Google/23andMe to find and hold all of your genetic secrets? I'd trust them a lot more than Joe Shmoe in I.T... Seriously, if you are afraid Google is setting up some secret clone army bent on world domination, your genetic information really isn't going to help them much anyway. The truth is, they don't really have any incentive to do anything malicious with your personal information. Their famous moniker "don't be evil" isn't just could p.r., its good business - and thats why I actually do trust them. Not too mention that if anyone could actually keep your information safe from those that would do evil with it (perhaps certain governments..) they are about the only group in the world that might stand a chance.

I absolutely can not wait until I can have my whole genome sequenced and can start playing with it. I want to see specifically what differences exist between me and my sister, my cat, my dog, and my plant. I want to know what pills I should take when, if I should be getting tested sooner for prostate cancer, and whether I am more closely related to Socrates or to Gandhi. I want to go to a party and play the 4th cousin game.. who in the room is my closest relative? I want a t-shirt with the damn thing on it!

Never fear knowledge.