Showing posts with label scifoo. Show all posts
Showing posts with label scifoo. Show all posts

Saturday, September 13, 2008

one more sci foo tag cloud

Alright, I know this is getting a bit old now, but I read this summative post about the meeting by Frank Wilczek (I got the first signed copy of his new book ;) and thought it would make a much more descriptive cloud than the ones I put together during the meeting about the attendees and their ideas for sessions.  So, here is yet another wordle-generated scifoo cloud based on Frank's post.  (It is better if you click through to the larger image).



Sunday, August 31, 2008

Peter (Google) and Christine (the librarians)

I was very lucky to have my new wife with me at SciFoo for many reasons, not the least of which is that she is much better at socializing than I am and thus managed to introduce me to many people I would never normally have met.  One of those people was Christine Borgman, Professor & Presidential Chair in Information Studies at UCLA.  While I was struggling to explain my work on the Entity Describer project to her, she noticed Peter Norvig walk by and dragged him over to join our conversation.  I guess she must have known him from somewhere but I'm not sure where.  Anyway, I didn't realize this at the time, but Peter is head of research at Google.  Ahem... did I mention that SciFoo interactions could be somewhat intimidating?  So there I am, standing between two giants of modern information science trying to explain what it was I was doing there but mostly trying to get some insight into their respective thoughts on the organization of the world's information. It wasn't a long discussion, but here are the basics.
The main question that I posed to them was whether or not and how semantic tagging (a la ED) is or might be useful.  On the surface, the answer from Peter was no and the answer from Christine was yes.  However, the truth of the matter is that they were really talking about supporting different functions for the end user - though this fundamental difference became a little lost during the conversation.  Google is principally focused on providing the best possible results, to the most people, given the least amount of information in the query - that is, keyword based search of the entire Web.  Library-science is typically much more concerned with providing the capacity for people to make very specific requests using much more sophisticated queries that operate over much smaller collections of information (e.g. the library of congress). The fact that there is some overlap in the information needs of the users of these different kinds of systems often brings up the desire for combative  comparison, but I think that, in reality, there is clearly no need for combat because they are simply too different in the functions that they intend to provide.
My interpretation is that Google isn't really concerned with intentionally provided meta-data in the name of end-user, full-Web search because the scale that they operate on seems to render any such indexing by one or even a number of parties almost laughably shallow in its characterization of both the nature of any particular item and its expected relevance to a query.  When you have literally millions of people passively voting and indexing every item of the Web through their decisions to link to it or not, you have very sophisticated algorithms for understanding the text in the pages generating and receiving those links and to top it off you record and process millions of people's behavior when faced with your search results, why should you care what some person or institution says the item is about?  The fact that they (among other search engines) beat out the directory-based approach to finding information on the Web is a clear demonstration that automatic indexing and link based relevance ranking do a better job than meta-data based classification - for the problem of Web scale search.  Google clearly doesn't need human semantic indexers to succeed, though, as Peter said, they certainly use all of the information that exists.  If there happen to be good indexes online (as Connotea turned out to be be for a fairly brief window), then their algorithms will certainly find them and use them - if not, no worries, the algorithms will take advantage of the 'normal' data on the Web and do just fine thank you. 
From the library-science professional perspective, this attitude is clearly annoying.  If human indexing isn't really necessary to find things, what is the point of the field that has devoted itself to the creation of effective ways for people to categorize things for retrieval?  For example, there is a lot of annoyance that the Google Books initiative seems to ignore most, if not all of the meta-data already associated with the books that they are scanning and indexing.  This means that meta-data, even as basic as volume numbers, is inaccessible for searching.  For the library professional that is trained to both search through and construct careful and precise classification structures,the inability to even search for a specific volume of a book is infuriating - particularly knowing that it is well within Google's power to incorporate such abilities into their system.
So on the one side we have the perspective that there is still value in the careful, intentional use of meta-data in the search and retrieval process while on the other side we are quite happy to let the intersection of algorithm and massive passive indexing do the work.  I guess, as is the usual answer, I'd suggest that both sides provide useful functions that are both worth keeping and advancing.  A detailed classification system, either constructed intentionally through the work of professional labor or semi-intentionally through the work of social taggers, provides functionality that is clearly different than what can be achieved by automatic indexing; however, it may not provide any help whatsoever in improving a full-Web-scale keyword-based search.  The essence of the power of intentional classification is the precision of the queries that it enables.  For example, if I want only version 3 of "The Devil's Rights and the Redemption" and thats it or I want only those items that have been tagged as with bioinformatics and to_read by Jaa, there is really no way ( AFAIK) to accomplish this without the intentional recording and utilization of meta-data about those resources.
So, though Peter and Google may have little direct use for ED and its semantic meta-data generating and consuming brethren emanating from the library and information sciences, there are still clearly meaningful applications of such work.  It just happens that providing effective search over the contents of the entire Web based on a string like 'Britney Spears' isn't really one of them.
I'm ok with that.

Sunday, August 24, 2008

things I should write

This always happens.  I go to a meeting of some kind, get loaded up with new ideas and experiences, get excited about writing about them, and then reality comes crashing back down and I never end up finding the time to do the writing properly, if at all.  Before the memories of scifoo slip further away, I just want to note some ideas I had for posts about the experience.  If any of them strike your interest, let me know and I will do my best to do a complete post on the topic.

  1. People I met at scifoo
  2. Things I collected at scifoo
  3. Google's role in the future of bioinformatics
  4. Micro-credits.  How to assign credit and blame in the smaller, newer bits that compose pieces of open science projects or other useful collaborative endeavors
  5. My brief talk with Geoffrey Carr, the science editor for The Economist
  6. Google could care less about meta-data, my talk with Peter Norvig and Christine Borgman.  
  7. What I would do differently if I ever go to another foo camp
  8. Why I should have been twittering
  9. cell phones versus one laptop per child
  10. Impressions of the Google campus
I'll leave it at that, let me know if anything strikes your fancy!

Monday, August 18, 2008

biodiversity researchers at scifoo

As mentioned in a previous scifoo post, I met Brad Zlotnick, Dan Janzen, and Winnie Hallwachs before even getting on the bus to head to Google. As it turns out, they are involved in a DNA-bar coding effort that is being applied to Dan and Winnie's field research in Costa Rica as part of a global effort to identify unique DNA signatures of all of the world's species. Another person involved in this effort (its big..) that I had several good conversations with was Vince Smith, a 'cybertaxonomist' at London’s Natural History Museum. These conversations about biodiversity, taxonomy, tagging, and semantic web-based knowledge representation really excited me because it felt like I might be able to make a real contribution to their efforts. Who knows, maybe biodiversity research will have an important part in the next phase of my working life - which would be great because it might mean I get additional offers to occasionally go on sample collection trips ;).

The biodiversity folks were also exciting because they are actually in the process of constructing a portable organism identification system (built to detect the DNA bar codes) that sounds like it will actually become a reality in the not-too-distant future. I've wanted one for backpacking trips for many years, but everyone always said it was impossible. Guess not :). I think the only time I wrote about this was just last year though its something I've annoyed nearly anyone I've ever taken a walk in the woods with.

Sunday, August 10, 2008

bye bye scifoo

Well, scifoo is over now. As I reluctantly take off my conference badge I can't help but feel sad to leave all of the people I've met here behind. The only recompense I suppose is that I do get to take a tiny bit of their ideas and their inspirational energy along with me. My main regret is that I think I really took away a lot more than I gave - though I think that is probably a pretty universal feeling for the attendees.  


I'll dribble out posts about specific parts of the meeting over the next few weeks, so that I have a bit of time to write them and the brain dump is not too overwhelming to you, the ethereal readers (both of you!).




Saturday, August 9, 2008

scifoo 2008 friday night

A quick recap, hopefully I will follow these eventually with something more coherent but I need to sleep.

Here are a few of the people that I spent the most time talking with today - listed in the order I met them (with the text describing them copied from their entries in the wiki - if anyone wants their entry changed or removed from here please let me know):

  • Pablo Gleiser: Physicist at the Statistical and Interdisciplinary Physics Group in Centro Atómico Bariloche, Argentina. I am currently working on the relation between complex networks and synchronization phenomena.
  • Brad Zlotnick emergency physician; furthering the common good; struck out Barry Bonds. Collaborative support of biodiversity development and innovation (DNA barcoding, index and application; carbon sequestration/forest sustained and regenerated).
  • Dan Janzen - Professor of Conservation Biology, University of Pennsylvania, Philadelphia; I am, along with Winnie Hallwachs (see below) and many others, inventorying all of the 10,000+ species of caterpillars, their food plants and their parasites, of Area de Conservacion Guanacaste in northwestern Costa Rica http://janzen.sas.upenn.edu; this 30-year and ongoing biodiversity project is attempting to manage massive amounts of biological collateral (text, tables and images), have it all available to anyone, and also have it linked to the DNA barcodes (a 650 base pair species-level-unique tag) for these species, as a guinea pig project for DNA barcoding the entire world by iBOL and CBOL, which in turn should lead to anyone anywhere at any time being able to identify any specimen or specimen fragment to species, and link that to the rest of what the world knows; the world can be barcoded for $3 b in 20 years.
  • Winnie Hallwachs: tropical ecologist. Tropical forests with their great species richness and boggling interactions - somehow biologists haven't managed to convey the grand scale, importance and fascination of what is out there. Looking for ways to understand (caterpillar website and barcoding - see Dan Janzen entry) and video their enormous diversity and interactions; to preserve some of the expert knowledge of biologists who have seen what is now gone; to make that data available in unconventional products that are intricately searchable, easily absorbed, and compelling; to keep it alive.
  • Barend Mons Originally amolecular biologist, I am presently focusing on computational biology and I am one of the co-founders of WikiProfessional. My main initerest is in speeding up scientific discovery by literature based knowledge discovery and notably the prediction of connections between concepts hitherto only implicitly associated in scientific data and literature. I am also co-founder of the company Knewco I will demo our concept web linker approach to anyone interested and will be looking for partners interested in building out the concept of Intellectual Networking.
  • Christine Borgman, Professor of Information Studies at UCLA. I conduct research on how data are made, asking fundamental questions about what are data, with a fabulous team of students, as part of the Center for Embedded Networked Sensing, an NSF Science and Technology Center. We apply research methods from anthropology, sociology, social studies of science, bibliometrics, and computer science to study scientific processes and to build better data management tools to support them. What is “information studies”?, you ask? We’re still trying to define it ourselves. My PhD students have brought their expertise in biology, physics, engineering, computer science, humanities, and studio art to bear in studying the very notion of data and its relationship to notions of information. My book Scholarship in the Digital Age: Information, Infrastructure, and the Internet (MIT Press, 2007) was reviewed in Nature and in Science this year. I also chaired the NSF Task Force on Cyberlearning. The Cyberlearning report, complete except for some minor editoral corrections, is posted at this URL until August 14, with the following disclaimer: Any opinions, findings, conclusions and recommendations expressed in this report are those of the Task Force and do not necessarily reflect or represent the views of the National Science Foundation. By August 14, we expect the final report to be posted on the NSF site. I’m eager to talk with other Sci Foo campers about scientific data, the changing nature of scientific scholarship, and how to use our technologies to improve learning in science.
  • TimoHannay: Publishing Director, Nature.com; ex-neurophysiologist; SciFoo co-organiser.
  • Matt Brown: I'm Editor of Nature Network, working from London with Timo Hannay. My role is to create a friendly and useful place online for scientists to exchange ideas - through blogs, forums and other tools.
  • Alexander Griekspoor: Independent software developer with a scientific background located in Cambridge UK who writes innovative software for scientists on the Mac. I'm perhaps most well-known as "Mek" from the duo "Mekentosj". Together with my friend Tom "Tosj" Groothuis I developed a number of scientific Mac applications, two of which won Apple Design Awards. What started in my spare time has become my passion and now also my work. I aim to create novel and revolutionary platforms for scientists, of which my most recent program Papers is a prime example.
  • Hilary Spencer
  • Joel Selanikio: pediatrician / social entrepreneur / photographer. As co-founder of DataDyne.org, I find ways to use open-source and lowest-common-denominator tech like SMS to bring the power of the ICT revolution to bear against problems of health and development in poor countries. I also practice pediatrics at Georgetown University, and spend a lot of time on planes, on my bike, and talking. And taking pictures.
  • Jon Kuniholm: Grad student in biomedical engineering and founder of The Open Prosthetics Project.
More to come tomorrow.. hopefully I will synthesize some of this in my dreams tonight.

Friday, August 8, 2008

scifoo session suggestion wiki tag cloud

Here is a tag cloud from the session suggestions page. Again, the length of the page makes it difficult to get a grasp of everything all at once, so the tag cloud gives a nice, quick feeling for what is happening.

scifoo attendee tag cloud

To get a feeling for the way scifooers describe their work and themselves, I used Wordle to create this tag cloud from the attendee page of the scifoo wiki.



(click for larger view)

scifoo arrival

Well its now just past 12:00am on August 8 and I'm in my official scifoo hotel room already, browsing through the scifoo wiki.  Its exciting, inspirational, and, of course, somewhat intimidating to see the array of brilliance they have assembled here.  Here are just a few session suggestions that caught my bleary eye in a rapid scan (the list is very long already)

"Make your own future" – Can scientists come out and play? I'm launching two massively multiplayer forecasting games this fall, and the goal is to get a LOT of scientists to play. I'll demo the platforms and talk about why I think scientists should spend more time playing games and imagining the future. (Jane McGonigal)"
"Why Whales are Weird: Comparative Anatomy and Evolution" by — Joy Reidenberg.
Her self-bio is characteristic of many scifooers - "(My research is in comparative anatomy particularly of marine animals, I am an artist, and I also teach Human Anatomy, Histology, and Anatomic Radiology at the Mount Sinai School of Medicine in New York".
Another interesting session proposal on "Collaborative Action Networks" by Paul Biondich.

So much to share! but I am jetlagged and should be sleeping up for tomorrow.. hope to write often this weekend.

Wednesday, May 21, 2008

Scratch that

So in the end, she convinced to try to go to scifoo after all.. She is awesome! It turned out that Timo extended the invitation to her as well!  It is definitely going to be one of the most unique honeymoon experiences ever.


If you are coming, see you at SciFoo !  
If not, look forward to some good blogging here in August this year.

oh, the pain!

I was lucky enough to be invited to this year's scifoo camp (see email included below for posterity). It is a huge honor to be invited and excruciatingly painful not to be able to go. Ack!

Of the few reasons that one might have for missing an opportunity to meet some of the smartest, most creative people in the world, and to visit the Mecca of geek society, mine is perhaps the best. I'll will be stranded on a tropical island with the woman I love.

Damn I'm a lucky man. (but oh, the pain of missing scifoo!)

Ben,

We'd like to invite you to join us on the weekend of August 8-10 for
Science Foo Camp (or "Sci Foo"), a unique, invitation-only gathering
organized by Nature, O'Reilly Media, and Google, and hosted at the
Googleplex in Mountain View, CA.

Now in its third year, SciFoo is already achieving cult status among those
with a passion for science and technology. The Economist said that it
"capture[s] the essence of innovation"; in a photo essay for Edge
(http://tinyurl.com/3o9sam), George Dyson wrote of the "the impossible
choice" when deciding which sessions to attend; another attendee described
it simply as "The best gathering ever. Period."

As before, we will be inviting about 200 people from around the world who
are doing groundbreaking work in diverse areas of science and technology.
Participants will include not only researchers, but also writers, artists,
investors, and other thought-leaders.

The format is highly informal: all delegates are also presenters and
demonstrators; the schedule is determined collaboratively on the first
evening; and sessions continue to be organized and re-organized throughout
the weekend. This creates a unique opportunity to explore topics that
transcend traditional boundaries, and discussions are of a kind that
happens at the best conferences during breaks and late into the night.
Of course, there will also be time to have fun and relax at Google's
legendary campus.

SciFoo 2008 will run from about 6pm on Friday, August 8 until after lunch
on Sunday, August 10. Campers need to make their own way to and from the
event, but Google will provide accommodation and meals, and there is no
registration fee. For those who don't have cars, there will also be free
shuttle buses between the hotel and the Googleplex.

Please RSVP by replying to this email. We do have space restrictions, so
if you'd like to attend please be sure to reply as soon as possible, and
in any case by May 16.

We hope to see you at the Googleplex in August!

Tim O'Reilly, O'Reilly Media
Chris DiBona, Google
Timo Hannay, Nature