Showing posts with label conference. Show all posts
Showing posts with label conference. Show all posts

Wednesday, October 1, 2014

Conference proceedings are citable, stop double-dipping

Over the last several years there has been a trend among conferences such as the Bio-Ontologies Special Interest Group meeting at ISMB and the Semantic Web applications and tools for life sciences (SWAT4LS) to invite article submissions for people that want to present at the conference and then to subsequently invite presenters to expand their article in an "official" publication in an associated journal.  Since the PDFs of the articles submitted to the conference usually end up online, as they should, this results in a situation where first reviewers and later readers are often confronted with two versions of essentially the same article - typically with the same title, author list, and often the same abstract.  This causes problems for reviewers as this kind of overlap with prior work (even from the same authors) would typically be grounds for rejection - yet because of this bizarre arrangement with the conference, reviewers are supposed to treat the original article as if it were a pre-print of the first, despite the fact that it is a citable entity on the Web and never referred to as a preprint anywhere.

Conference organizers, please stop this madness.  Here are three models that would be better.

  1. Following the International Biocuration Conference model, invite submissions directly to the partner journal first and then choose presenters from the successful submissions and independently submitted abstracts.  No confusion.  One good, citable paper.  Probably higher quality conference submissions.
  2. Put the articles submitted to the conference in a pre-print server such as arXiv or BioRxiv and continue with the concept of an expanded article in a journal.
  3. Do what the computer science community does and recognize contributions to conference proceedings as citable articles and do away with the attempt to get an 'official' journal publication in addition to the conference citation.

Monday, April 21, 2014

The Cure at Salk Cancer Day Symposium

Karthik G. and I will be presenting a poster tomorrow at the Salk Institute's Cancer Day Symposium.  We will be presenting data from a year with the scientific discovery game The Cure.  You can read more about those results on the arXiv.

If you are coming, please stop by for a chat!  We would especially love the chance to discuss the new, collaborative decision tree-building interface that Karthik has created.  Who knows if the conference wifi will work, so please try it now!


Thursday, April 10, 2014

Microtask crowdsourcing and biocuration

Thanks to the hard work of my coauthors @x0xMaximus and @andrewsu , I was able to nab the award for the best presentation at The Seventh International Biocuration Conference from the International Society of Biocuration.   The slides for the presentation and the poster are available from slideshare.

Yay team!

I think the presentation garnered the interest it did because many of the people in the audience had heard the term "crowdsourcing" before, but had never seen a real example of a specific application - let alone one in science.  I was surprised by the number of people that I spoke to that had no idea what the Amazon Mechanical Turk was - nevermind that it might be applicable to some of the problems they were working on.  We had a decent result to talk about, but much more importantly, we taught the audience about a powerful new tool that they might be able to use in their own work.

For those that do want to try scientific applications of microtask crowdsourcing I'd like to emphasize that its probably not going to be an easy process.  The result we presented was from the third iteration of our system and represents several months of developer time.  While resources are emerging that should make this process much faster to get started (e.g. [1-4]), expect to engage in an iterative cycle to get your system dialed in!

If you do want to give crowdsourcing a try for biocuration or other scientific objectives, (1) we would love to hear about it! and (2) it might be worth a quick look at our review of the domain [5].  Microtask systems such as the one we worked with here are just one of many ways that scientific challenges can be opened up to much broader communities.

References
  1. Our code: mark2cure 
  2. Soltilab mention tagger for crowdflower
  3. GATE crowdsourcing plugin 
  4. Crowd Watson from IBM
  5. Good, Benjamin M., and Andrew I. Su. "Crowdsourcing for bioinformatics" Bioinformatics 29.16 (2013): 1925-1933.

Tuesday, November 6, 2012

Return to Moscone

I'm sitting in the main hall at the enormous Moscone conference center in San Francisco awaiting the first plenary at ASHG2012 and remembering the last time I was here.  Back in spring of 2007, I saw Jeff Bezos and several other Web luminaries speak here about the Web2.0 phenomenon - what it was and how they were planning to make money on it.  The buzz throughout that conference was Twitter, though I admit I hadn't really noticed it before then, did not understand it, and was very skeptical that it would amount to anything.  That was the meeting that really inspired me to start writing here.  5 and a half years and 214 posts later its clear that, because of that, it was probably one of the most significant meetings in my professional career.  Who knows?  Perhaps this genetics business will prove even more inspirational.

Friday, November 2, 2012

Gene-disease annotation with Mobianga

Over the summer, an enterprising high school student named Nishant Mandapaty approached our research group about doing a project with us.  He found us through the "Crowdsourcing Biology" group that we created for the Google Summer of Code program.  To make a long story short, he has been doing good work for us ever since, quickly learning what he needs to on his own.

His primary contribution so far is a nascent game for collecting gene-disease connections called Mobianga!.  He is planning to submit the results of an experiment using this game to the Intel Science Talent Search, a prestigious national science fair.  But, he needs help if he is going to succeed!  If you know anything about genes and their relationship to disease or are capable of using resources like OMIM, PubMed, and Google find such information he needs you to play a few games!  Even better, invite several of your friends to play a few games.

Help a 17 year old computational biologist reach his dreams, play Mobianga! today!

Mobianga contest at the American Society for Human Genetics annual Meeting

Next week, I will be attending ASHG at the Moscone Center in San Francisco.  I will be presenting a poster (number 3584W) on Wed., Nov. 7 from 3:15-4:15 about game applications in biology.  If you come by and can prove that you have played a game of Mobianga by identifying your name on the leader board I will give you a prize!  I will also be giving out a more substantial prize at the end of the meeting to the top-scoring player as of the morning of Saturday November 10.  Details on the ASHG Mobianga contest will be distributed at my poster.

Technical details

Mobianga! makes use of the human disease ontology to provide the opportunity to easily annotate genes at varying levels of granularity.  For each gene challenge, you start at the top of the hierarchy (e.g. choose between 'disease of cellular proliferation' and 'disease of mental health') and you work your way down to specific diseases.  At each step you earn points based on an algorithm that assesses the precision of the annotation and degree of consensus among prior players in a manner similar to the Herdit game for music tagging recently published in PNAS.  

The game is implemented as a Python-powered Web Application that runs in the Google App Engine.  The code is open source and he would welcome collaborators.  The game is intended to eventually run smoothly on phone-sized browsers (the name Mobianga came from 'the mobile annotation game'), but this optimization has not yet been achieved.  Anyone that wants to help, please get in touch.

Friday, July 20, 2012

Social machines at ISMB 2012


This year's 20th annual conference on Intelligent Systems for Molecular Biology (ISMB) was a busy one for me and the rest of the Su Lab.  As a group of 5, we were responsible for 4 oral presentations, three posters, and the administration of one special session.  Keeping that all together while catching up with old and new friends was a fun, though exhausting experience.  And nevermind trying to follow the ISMB twitter stream!

Very briefly, we presented as follows:

You may notice a fairly consistent theme here (with the exception of Erik's award-winning contribution).  The idea of 'community intelligence' was also explored in the four invited presentations in the special session that I helped to organize.  (Thanks to Alex Garcia for the original idea and rabble rousing that lead to this session's conception nearly two years ago).  At that session we heard from Alex Bateman on the integration of Wikipedia with RFAM, PFAM, and traditional publishing, Alex Pico about lessons learned in the WikiPathways project, Andrew Su about the Gene Wiki initiative, and finally Firas Khatib on the protein folding game Foldit.

Of all the very impressive and interesting presentations, Alex Pico's stood out for me.  To keep this short, I'll leave the specific recap on the other projects to your Googling and finish this with a couple thoughts that percolated from Dr. Pico's perceptive presentation.

A Pico Lesson

WikiPathways is doing great right now.  They have a very rapidly growing user base and are on their way to becoming the de facto standard resource for pathway information.  Of particular interest is that more than 20% of its registered users have made edits to pathways.  That is an astoundingly high rate of user-to-editor conversion.  We claim great success with the Gene Wiki and I am fairly certain that our ratio is less than 1% (though its difficult to tell exactly as the definition of 'user' is fuzzier).

So, how are they making it work?  One of the things that Alex emphasized in his presentation was that WikiPathways was created to solve their own problems in collaboratively editing and sharing pathways (as part of the GenMapp project).  It would have been useful and used (by them) even if no one outside of their research group ever got involved.  The fact that it has been taken up by a broader community is a very valuable, but secondary effect.  This basic idea was echoed in the twitter echoes of Carole Goble's talk (I missed hearing it directly) and resonates yet again with the inescapable Del.icio.us lesson.  Personal value precedes network value.  

As we forge ahead into the realm of Community Intelligence, we need to keep that lesson foremost in our minds.  When we are thinking about games, that means that the game actually has to be fun, really fun!  Time will tell if we can cross that threshold...

Tuesday, March 6, 2012

Sepublica 2012 deadline extended


Deadline extended: now March 18

http://sepublica.mywikipaper.org/ – the future of scholarly communication and scientific publishing

SePublica2012 an ESWC2012 Workshop.
May 27 or 28 (exact day to be announced), Heraklion, Greece.

At Sepublica we want to explore the FUTURE OF SCHOLARLY COMMUNICATION and SCIENTIFIC PUBLISHING. As we are going through a transition between print media and Web media, Sepublica aims to provide researchers with a venue in which this future can be shaped. Consider research publications: Data sets and code are essential elements of data intensive research, but these are absent when the research is recorded and preserved by way of a scholarly journal article. Or consider news reports: Governments increasingly make public sector information available on the Web, and reporters use it, but news reports very rarely contain fine-grained links to such data sources.  At Sepublica we will discuss and present new ways of publishing, sharing, linking, and analyzing such scientific resources as well as reasoning over the data to discover new links  and scientific insights.


View Larger Map

Tuesday, January 10, 2012

Semantic Publishing workshop at ESWC in Greece


I'm helping to organize this exciting event, please consider submitting a manuscript and or attending.  From the official call for papers:
http://sepublica.mywikipaper.org/SePublica2012 an ESWC2012 Workshop.  May 27-31, Heraklion, Greece.
At Sepublica we want to explore the future of scholarly communication and scientific publishing. As we are going through a transition between print media and Web media, Sepublica aims to provide researchers with a venue in which this future can be shaped. Consider research publications: Data sets and code are essential elements of data intensive research, but these are absent when the research is recorded and preserved by way of a scholarly journal article. Or consider news reports: Governments increasingly make public sector information available on the Web, and reporters use it, but news reports very rarely contain fine-grained links to such data sources.  At Sepublica we will discuss and present new ways of publishing, sharing, linking, and analyzing such scientific resources as well as reasoning over the data to discover new links  and scientific insights. 
Workshop Format 
We are planning to have a full day workshop with two main sessions. During the first part of the workshop accepted papers will be presented; the second part of the workshop will address by means of focus groups two main questions, namely “what do we want the future of scholarly communication to be?”  and “how could data be preserved and delivered in an interactive manner over scholarly communications?”. These focus groups will be followed by a panel discussion. As an outcome of these activities we will have a communique that will be the editorial for the workshop proceedings,

Dates 
* workshop papers submission deadline: Feb 29
* workshop papers acceptance notification: April 1
* workshop papers camera ready: April 15 
Submission
 https://www.easychair.org/conferences/?conf=sepublica2012

Issues to be addressed
  • Representation:
    • Formal representations of scientific data; ontologies for scientific information
    • What ontologies do we need for representing structural elements in a document?
    • How can we capture the semantics of rhetorical structures in scholarly communication, and of  hypotheses and scientific evidence?
    • Integration of quantitative and qualitative scientific information
    • How could RDF(a) and ontologies be used to represent the knowledge encoded in scientific documents and in general-interest media publications?
    • Connecting scientific publications with underlying research data sets
  • Technological Foundations:
    • Ontology-based visualization of scientific data
    • Provenance, quality, privacy and trust of scientific information
    • Linked Data for dissemination and archiving of research results, for collaboration and research networks, and for research assessment
    • How could we realize a paper with an API?  How could we have a paper as a database, as a knowledge base?
    • How is the paper an interface, gateway, to the web of data? How could such and interface be delivered in a contextual manner?
Applications and Use Cases:
  • Case studies on linked science, i.e., astronomy, biology, environmental and socio-economic impacts of global warming, statistics, environmental monitoring, cultural heritage, etc.
  • Barriers to the acceptance of linked science solutions and strategies to address these
  • Legal, ethical and economic aspects of Linked Data in science

Tuesday, April 5, 2011

Collaborative Innovation in Philadelphia

Today I presented a high-level view of the Gene Wiki project to a group of people interested in expanding the application of "collaborative innovation" in the biomedical domain.  My presentation (and my lack of a suit) definitely felt like an outlier in this group.  Many of the other talks brought up the apparently looming demise of the pharmaceutical industry due to a lack of innovation and made suggestions about how companies and consortia could make changes that would allow them to tap into much broader, more diverse intellectual and physical resources using a variety of mechanisms.  We also had presentations about how the same kinds of cost-saving, creativity-enhancing approaches might be used to tackle neglected diseases (aka diseases that are not likely to make the pharma companies huge amounts of money by curing as opposed to say, cancer).  The special sauce here is this "collaborative innovation" stuff.  This seems to some extent to reiterate the age-old solution to complexity - 'divide and conquer'.  The problems of developing novel drugs today are diverse - any successful solution is the result of the execution of many many different tasks.  The key challenges for people in this industry (and many others I imagine) seem to be knowing which of the tasks they are very good at completing, which tasks they are not, who the people are that can complete these tasks, and how to legally transact mutually beneficial collaborations with those other people.

For more information see the conference website for abstracts and links to the talks.

Random notes to self:

  1. I love talking about the gene wiki!  Its just a great story to tell.  I'm really looking forward to the time that my part of this story is half as interesting as the parts that have already transpired.
  2. Why do I spend so much time worrying about and getting ready for presentations that are largely gone the moment they are completed and so little time on these blog posts that tend to stick around for years?
  3. This evening I spent far too much time preaching that games were the future of this domain and that Jane McDonigal had the keys to saving the world in her book 'Reality is Broken'.  I couldn't help it, I got excited!
  4. I should have spent more time listening to what people like Hassan had to say - especially as he seems to have been studying game-like systems for large-scale collaborative innovation for a long time...

Thursday, March 10, 2011

TBI-AMIA 2011 Brain Dump

Here is a very rapid dump of a few of the tidbits that stuck in my brain after spending the last few days at the Translational Bioinformatics (TBI) conference for 2011.

The main recurring themes of interest for the year were:

Social networking for scientists: dirrecttoexperts, vivo, open social, patient recruitment (e.g. army of women).  All efforts are heavily into the linked open data concept.

Temporal reasoning for clinical data: the chronology of the patient's symptoms is an important and underrepresented aspect to mine/model in patient records

Naive Bayes : came up at least twice in important places - once in Lincoln Stein's keynote (used in the automatic expansion of the Reactome database) and once in the nominated best paper by Wei Wei "The Application of Naive Bayes Model Averaging to Predict Alzheimer’s Disease from Genome-Wide Data".  Its an effective method to integrate many many different sources of evidence into a single predictor - a relatively simple machine learning algorithm that scales up well with a lot of data.

Genomic complexity : keynotes about cancer and microbiome highlighted the incredible diversity of genomes.  One example described the difference between a tumor cell and a normal cell one inch away on the order of 50,000 SNP differences - add on top of that genomic rearrangements and epigenetics and well... its complicated.  Complete genome sequencing renders all diseases orphans from a drug development perspective - major changes needed for pharma...  Major opportunities as well.  

Politics : NCBO vs. OBO vs. UMLS
Each of these groups is or will be competing for the same pot of money to solve very similar problems.  While members do collaborate, e.g. key UMLS players advise NCBO, they have serious disagreements about how things should be done and I noticed some pretty emotional discussions - mainly related to fears caused by expected funding cuts.  
 name      goal  conflict with OBO  conflict with UMLS
 NCBO  empower researchers with access to ontologies and related tools   NCBO is very open about including ontologies.  OBO is not - they believe strongly in their foundry principles (and their distinctive file format) and it annoys them that anyone can get their ontology into NCBO.  See goals.. and think $$

NCBO has a different approach to concept detection that annoys the MetaMap people.  MetaMap feels like NCBO is throwing away years of work on NLP that should be used.
 OBO  provide "high quality" collection of biomedical ontologies Some OBO members seem to have a longstanding dislike of the way that the UMLS is modeled.  OBO believes they are right and the UMLS is wrong and have effectively ignored its presence in the development of their ontologies.
 UMLS      empower researchers with access to ontologies and related tools

Unsurprising : lots of people building ontologies to use for structuring data and lots of people applying text mining approaches to mine unstructured data

There was a whole lot more than that, for some more info, have a look at the conference papers, and at Russ Altman's year in review.  



Sunday, March 6, 2011

with a flower in my hair

Assuming the fog clears enough to land, I'm heading to San Francisco today to attend the 2011 AMIA Summit on Translational Bioinformatics.   I'll be presenting Tuesday morning about mining structured gene annotations from the text of the Gene Wiki.  Supporters and hecklers would be welcome!

Tuesday, February 1, 2011

cash for semantic publishing research

Following from my last post, here is an update from the SePublica workshop


SUBMISSION DEADLINE February 28
ELSEVIER BEST SEMANTIC PAPER AWARD
The Best Paper Award is presented to the author(s) deemed to have written the paper covering the most innovative and feasible proposal concerning semantic publishing in the workshop. All submissions to the SePuBlica workshop will be considered, and a panel of experts will rate the papers according to originality of the idea, feasibility and  presentation. The Best Paper award is sponsored by Elsevier as an incentive for researchers working on defining the next generation of scientific publishing concepts.  The Best Paper Award will be handed out at the end of the SePuBlica workshop.
As a cash prize, the Best Paper Award will receive: US$ 750
The runner-up will be awarded a prize of US$ 250.

Wednesday, January 12, 2011

CfP Semantic Publishing

Minoan rhyton from Crete!
I'm on the program committee for this conference workshop so clearly you should submit something (and its in Crete!).  See the call for papers below:

-----------------------------------------------------------------------------------------------------------------

1st International Workshop on Semantic Publication (SePublica 2011)
http://sepublica.mywikipaper.org
at the 8th Extended Semantic Web Conference (ESWC 2011)
http://www.eswc2011.org
May 29th or 30th, Hersonissos, Crete, Greece
Keynote by Steve Pettifer, Manchester University, UK.
“Utopia Documents and The Semantic Biochemical Journal experiment”

SUBMISSION DEADLINE February 28

The MISSION of the SePublica workshop is to bring together researchers
and practitioners dealing with different aspects of Semantic
Technologies in the Publishing Industry. How is the Semantic Web
impacting the publishing industry? How is our experience of
publications changing because of Semantic Web technologies being
applied to the publishing industry?

The CHALLENGE of the Semantic Web is to allow the Web to move from a
dissemination platform to an interactive platform for networked
information. The Semantic Web promises to “fundamentally change our
experience of the Web”.

In spite of improvements in the distribution, accessibility and
retrieval of information, little has changed in the publishing
industry so far. The Web has succeeded as a dissemination platform for
scientific and non-scientific papers, news, and communication in
general; however, most of that information remains locked up in
discrete documents, which are poorly interconnected to one another and
to the Web.

The connectivity tissues provided by RDF technology and the Social Web
have barely made an impact on scientific communication nor on ebook
publishing, neither on the format of publications, nor on repositories
and digital libraries. The worst problem is in accessing and reusing
the computable data which the literature represents and describes.

• Consider research publications: Data sets and code are essential
elements of data intensive research, but these are absent when the
research is recorded and preserved in perpetuity by way of a scholarly
journal article.
• Or consider news reports: Governments increasingly make public
sector information available on the Web, and reporters use it, but
news reports very rarely contain fine-grained links to such data
sources.

QUESTIONS AND TOPICS OF INTEREST

• What does a network of truly interconnected papers look like?
How could interoperability across documents be enabled?
• How could concept-centric social networks emerge?
• Are blogs and wikis new means for scholarly communication?
• What lessons can be learned from humanities and social science publishers
(i.e. going beyond scientific publishing towards scholarly publishing)?
• How could we move beyond the PDF?
How can we embed and link semantics in EPUB and other e-book formats?
• How are digital libraries related to semantic e-science?
What is the relationship between a paper and its digital library?
• How could we realize a paper with an API?
How could we have a paper as a database, as a knowledge base?
• How is the paper an interface, gateway, to the web of data?
How could such and interface be delivered in a contextual manner?
• How could RDF(a) and ontologies be used to represent the knowledge encoded
in scientific documents and in general-interest media publications?
• What ontologies do we need for representing structural elements in a
document?
• How can we capture the semantics of rhetorical structures in
scholarly communication, and of  hypotheses and scientific evidence?

AUDIENCE

• researchers from diverse backgrounds such as argumentative
structures, scholarly communication, multi-modality in publications,
digital libraries, semantics in publications, and ontology
engineers.
• practitioners active in the publishing industry, repositories of
experimental information and document standards.

IMPORTANT DATES

Paper/Demo Submission Deadline: February 28, 23:59 Hawaii Time
Acceptance Notification: April 1
Camera Ready Version: April 15
SePublica Workshop: May 29 or May 30 (to be announced)

SUBMISSION AND PROCEEDINGS

Research papers are limited to 12 pages and position papers to 5
pages. For system descriptions, a 5 page paper should be
submitted. All papers and system descriptions should be formatted
according to the LNCS format

http://www.springer.com/computer/lncs?SGWID=0-164-6-793341-0

We encourage the submission of semantic documents. LaTeX documents in
the LNCS format can, e.g., be annotated using SALT
(http://salt.semanticauthoring.org) or sTeX
(http://trac.kwarc.info/sTeX/). We also invite submissions in
XHTML+RDFa or in the format or YOUR semantic publishing tool.
However, to ensure a fair review procedure, authors must additionally
export them to PDF.  For submissions that are not in the LNCS PDF
format, 400 words count as one page. Submissions that exceed the page
limit will be rejected without review.

Depending on the number and quality of submissions, authors might
be invited to present their papers during a poster session.

Please submit your paper via EasyChair at
http://www.easychair.org/conferences/?conf=sepublica2011

The author list does not need to be anonymized, as we do not have a
double-blind review process in place.

Submissions will be peer reviewed by three independent
reviewers. Accepted papers have to be presented at the workshop
(requires registering for the ESWC conference and the workshop) and
will be included in the workshop proceedings that are published online
at CEUR-WS.

PROGRAM COMMITTEE

• Robert Stevens, Manchester University, UK
• Benjamin Good, GNF, USA
• Michael Kohlhase, Jacobs University, Germany
• Oscar Corcho, Politecnica de Madrid, Spain
• Steve Pettifer, Manchester University, UK
• Jodi Schneider, DERI, NUI Galway, Ireland
• Sebastian Kruk, knowledgehives.com, Poland
• Henrik Eriksson,  Linköping University, Sweden
• Dagobert Soergel, University of Maryland, USA
• Tim Clark, Harvard Medical School, USA
• Paolo Ciccarese, Harvard Medical School, USA

ORGANIZING COMMITTEE

• Alexander García Castro, University of Bremen, Germany
• Christoph Lange, Jacobs University Bremen, Germany
• Anita de Waard, Elsevier, USA/Netherlands
• Evan Sandhaus, New York Times, USA

QUESTIONS? → sepublica@googlegroups.com

Thursday, December 3, 2009

Heading to China

ChinaImage via Wikipedia
Tomorrow I am departing to attend the Asian Semantic Web Conference in Shanghai, China.  I'll be manning the demo for my former labmate's project CardioSHARE (which I played no real part in building).

Looking forward to getting caught up on the latest from the semantic web research community.  Out of the accepted papers, I am most looking forward to hearing "Merging and Ranking answers in the Semantic Web: The Wisdom of Crowds".

Hope to see you on the other side of the Great Firewall.
Reblog this post [with Zemanta]

Wednesday, October 22, 2008

Heading to ASIST

This Saturday I will be heading down to Columbus, Ohio to attend the annual meeting of the American Society for Information Technology.  If you are coming, hope to see you there!


I will be presenting a poster about an empirical approach to the study of (human) indexing systems based on the comparative analysis of the syntax of the terms that are used within the different systems.  The work emerged from investigations of the relationship between social tagging (e.g. Connotea) and professional indexing (e.g. MEDLINE).  As has been typical in my career thus far, when I got started on the project, I found that I didn't have the tools I felt that I needed to answer the question satisfactorily so I spent most of my time trying to make them. 

One of the more interesting outcomes of the work are visualizations of the differences between the structures of the terms used within ontologies, thesauri, and folksonomies.  They are surprisingly distinct from one another. The abstract for a full paper (just accepted today) describing the work as well as the programs written and the data processed are available here.


Saturday, November 3, 2007

KCAP 2007


Aside from getting ready for that committee meeting, I had the chance to go the 4th international conference on knowledge capture (KCAP) up in Whistler, British Columbia. From the new technology perspective, the most impressive results I saw came from the Turing Center down in Seattle. Though they didn't say so directly (in fact the opposite), it seems to me the dream of Strong A.I. is alive and well down there.. In a massively reduced nutshell, they are capitalizing on the scale of the Web to start translating its textual content into machine processable units of knowledge automatically. They are working on the reading machine. See TextRunner and Alice for examples.

I met and reconnected with quite a few people at the conference, here are a few of the main ones I spoke with:

  • Reconnect -> Marja Koivunen, founder of Annotea
  • Reconnect -> Mikele Pasin, creator of PhiloSURFical
  • Met -> Anna Tordai, part of the MultimediaN E-Culture project
  • Reconnect -> Derek Sleeman, among many things, one of the founders of the KCAP conference.
  • Reconnect -> Dragan Gasevic, professor at SFU and Athabasca
  • Met -> Andriy Nikolov, my roommate for the conference.
  • Reconnect -> Michele Dumontier, professor at Carleton university and one of the few bioinformatics representatives at the conference
  • Met!!!! -> Luis von Ahn, professor at Carnegie Mellon, inventor of the CAPTCHA, and coiner of one of my favorite terms 'Human Computation' for describing Games With a Purpose (GWAP). It was really exciting to meet him in person after reading and thinking so much about is work. He was very friendly and, in contrast to other famous super-geeks I've met in the past, didn't make me feel like a complete idiot when I spoke with him.

Though the most exciting new and working thing at the conference was the Web-scale NLP, the Next Big Thing clearly seems to be the integration of this work with Luis' human computation. Have to say, I'm happy to have made a small stab in that general direction already and am excited to do a better job of it in the coming months...

Wednesday, May 9, 2007

WWW2007, workshop on the collaborative construction of structured knowledge

I just returned from the World Wide Web conference in Banff, Alberta, Canada and now I'm frantically trying to assemble some notes before it all fades into the memory fog.

The first day of the conference I attended a workshop on the collaborative construction of structured knowledge. The purpose of the workshop was to explore requirements, opportunities, and current tools for enabling multiple people to contribute to the same knowledge resource. The resources ranged from tags at the bottom, through wikis, to the creation of description logic ontologies. Below I offer a few highlights of the day from my perspective, starting with my own little brush with fame.

Prior to the workshop, there was a competition involving 6 new collaborative KR tools. In this competition, the users of the tools competed with one another based on how many bits of knowledge that they constructed with the tools and on the quality of the comments that they provided to the tool builders. The contest incentivized people to actually try the (mostly pre-alpha..) tools thoroughly, thus providing the tool builders with good feedback and stimulating discussion at the workshop. At the end of the workshop, Natasha Noy, one of the organizers and a pre-eminent figure in the world of ontology based knowledge representation, preceeded her presentation of the results of the competition by acknowledging that a paper by "Mark Wilkinson" provided some of the inspiration for the idea of the competition. Ahem.

As the first author of the paper that she was referring to, I was 99% thrilled to have it recognized by such a prestigious figure but (at least) 1% distressed that it was referred to as Mark's paper without any mention of me or any of the other author's (though the others had a very limited role). I spoke up and she immediately apologized - and did so again in person after the end of the session. I have no hard feelings about it, and like I said, am delighted that people like Natasha, who are at the very top of this field actually knew about my work; however, the whole episode brings up a lot of interesting facets of the culture and perhaps the economics of academia (and perhaps all human society in the networked world) that might help to shed some light on the questions about blogging I am still trying to work through. I'll talk about this more in another entry..

A few highlights:
1) Andrew Gibson - Described the need for people of different backgrounds and with different desires to interact with ontologies. This provided some of the motivation for a web2.0 classification system for ontologies. Each ontology would have its own 'profile' consisting of meta-data (kept distinct from the ontology file) that would include information about aspects such as authorship, revision state, history, purpose, deployments, users, and even threaded discussions.

Cool, lets make it happen.

2) Michael Backhaus presented BOWiki - a biology specific semantic media wiki. Interesting to see if the work they did in grounding it in an upper ontology of function will matter. (Only if they can actually get users..)

3) SOBOLEO - My personal favorite entry in the challenge because a) it worked for me b) it was simple and c) I enjoyed the live edit tracking feature.

4) Also from SOBOLEO group: A very clever idea called imagenotion. Basically, it replaces concept term labels with pictures. I think this is an obvious (in retrospect) and potentially extremely powerful thing to enable. Is a picture worth a thousand words? Has to be a paper title for them if it isn't already. The incorporation of multimedia in the ontology browsing/creation experience is going to be exciting because it will really make it much easier (I think) to communicate the concepts with real people (not robotic logicians).

5) Great keynote "Stone Soup", "freebase" on economic lessons for the now web.

More to come on this..

Saturday, May 5, 2007

WWW conference


Along with three trusty geek companions (Dr. Mark, Dr. Cartik and Mr. Byron), I depart on Monday for the 16th Int. World Wide Web conference in beautiful Banff, Alberta, Canada. Aside from taking in the fantastic scenery and learning what I can, I'll be presenting a poster describing experiences with volunteer knowledge engineers.

Following the conference, we will take a quick trip to dinosaur valley in Drumheller Alberta Should be lots of fun! (Then back to ontology evaluation via inductive learning).