Thursday, March 27, 2008

Twine, Applications and Green Fields

Looking at Robert Scoble's interview with Twine's Nova Spivack reminded me of Freebase, Knol and countless others before. Basically the interview goes like this:

1. Sir Tim Berners-Lee* has talked about the importance of the data web, or more recently about the Giant Global Graph (GGG)*.
2. Twine has built technology that can create, maintain and query a data web.
3. Let's sit at our editor, create a little data web, and run a query
4. Look how beautiful, imagine all the wonderful things you could do with this!

* Substitute appropriate visionary here, like Danny Hillis or Won Kim
** Substitute appropriate information infrastructure technology like "semantic database", "object-oriented database", etc.

Now I did enjoy the interview and the editor does look nifty but I was still a bit disappointed.

First because the presentation is very technology centric, cleverly leaving the applications to the imagination. But what's a dataweb (or any information infrastructure) without applications? No, the editor from step 2 doesn't count as an application. Even a query interface or browser barely deserves that qualification.

Compare that to a database like Oracle. The relational database concept developed by Jim Grey et al was a major breakthrough, as was the query language SQL. But there really isn't much use to a relational database until you run business applications on top of it -- think of Payroll or Shopping Cart. Oracle didn't win the database war from IBM, Informix, Tandem, Microsoft and countless others because their database was technologically most advanced, but because they focused on the applications -- lately in the extreme by selling the database as well as its applications, either developed in-house or obtained through myriad acquisitions. Here's a list, and that's only the "strategic" acquisitions!

So I perked up when Robert asked "what are the applications you have in mind", but I slumped back when I heard Nova answer "we think this is something that would be used by work groups".

The second disappointment was what I call the "Green Field Approach". The demo starts with Nova creating a Twine, cutting and pasting some text, and then creating relations with other twines and some existing web pages. A bit further in, Nova shows how you can import information from other sources, but that seems an afterthought: "we're looking at what other sources might be interesting".

But given the terabytes of information that are already out there it seems the last thing we need is human authors creating new pages -- shouldn't the main goal be to navigate, organize, link and clean up existing information? Shouldn't creating the dataweb be as easy as tagging -- which very successfully avoids the creation of original content?

The best of luck to information technology companies -- we definitely think of ourselves as one.

But remember:
No Apps, No Glory!
and
There is Plenty Information Already!

Sunday, March 23, 2008

Wikinomics and Maslow's hierarchy of needs

Reading Wikinomics I wondered what it takes for wikis to work in the enterprise. Jimmy Whales, the head Honcho of Wikipedia, has said that 0.7% of users did 50% of the edits on Wikipedia. If those same numbers held for the enterprise, a 1,000 person company's wiki would hinge on the contributions of 7 people. There has been debate about the Wikipedia numbers, but beyond that debate the numbers don't translate to the enterprise because the incentives are different.

What are the incentives to produce content on Wikipedia? There is the altruistic motive of providing free (as in beer) truth to the world, there is the satisfaction of finding and correcting flaws and, perhaps most important, the value of asserting yourself as a topic expert.

Referring to Maslow's hierarchy of needs, none of these incentives translate into level one (physiological) and two (safety) needs, perhaps "belonging to the Wikipedia community" might qualify at level three.

Contrast that with creating or editing a useful page on your company's wiki. Saving colleagues time not only improves the bottom line but might get you noticed by your manager or perhaps even your manager's manager, both of which translate into food and safety (of employment).

While we couldn't extrapolate our numbers because service networks go far beyond wikis -- especially by automatically providing much-needed structure -- we have been seeing the power of those incentives during implementations of Service Networks for Upgrades. Traditionally, non-IT employees shy away from tasks in planning or testing an Oracle Upgrade, but when they see how their contributions will be visible on the network, they often jump on board. With the help of their published experience, upgrades are completed in half the traditional time and at one quarter the traditional cost.

As a lot of momentum is going into Enterprise 2.0, I hope to see more of this type of quantification. While Wikipedia and Facebook may run on Maslow's belonging, esteem and self-actualization, mainstream enterprises will need to understand the material impact on their business for Enterprise 2.0 to really take off.

Wednesday, January 16, 2008

Knols, Wikipedia and truth in Openwater Service Networks

Last month, Google unveiled knols. In Udi Manber's words,

The key idea behind the Knol project is to highlight authors.

Unlike Wikipedia, which contains a single definition for each topic, authors write competing knols about the same topic.

Udi argues that competition is a good thing. Google will not edit the knols, but use its considerable Search Quality expertise to let the best version float to the top.

Considering Wikipedia and Knol raises interesting philosophical questions about truth. Take Hugo Chávez . Many view him as a socialist liberator, many others view him as an authoritarian demagogue.

Wikipedia editors deal with these versions of the truth in a variety of ways, labeling actions or statements as "controversial", requiring references to facts from reputable sources, flagging text as "opinion" and even introducing a separate topic Criticism of Hugo Chávez . However admirable these efforts, some bias undoubtedly remains. For example, these graphs make him look like a hero to a visual person like me. The text next to it has enough nuance that I couldn't tell. And that's where Wikipedia (or any encyclopedia, for that matter) usually leaves me -- my head full of facts from more or less reputable sources, various conflicting interpretations of these facts but no point of view.

Enter Knol. Knol isn't live yet, but I imagine we'd get at least two knols: Hugo the Good and Hugo the Bad. It will be interesting to see which Hugo will float to the top -- it will be a reflection of the community that views and rates knols.

The word "community" provides a good segueway into what we do at Openwater. Rather than demanding one version of the truth, or letting several truths compete, we build service networks for communities (specifically the installed base of a high-tech product), each of which creates its own Community Encyclopedia.

As in Wikipedia, anyone in the service network can edit a topic and each topic has experts assigned who monitor the quality. But like Knol, the editors don't have to balance every sentence that appears to carry an opinion or bias. And perhaps the biggest difference with both Knol and Wikipedia: there is no need to consider all possible meanings of each word -- just the current meaning within the installed base community is fine.

As a result, topics can be more specific and concise than those in either Knol or Wikipedia, and at the same time they can index more specific (proprietary) information and people in the installed base community.

There is one other major difference I'd like to mention in this blog -- this is what makes us an "implicit web" company. Unlike Wikipedia and Knol, topic pages don't start as an empty page and a blinking cursor.

In an Openwater Service Network, we start by indexing all the structured and unstructured information that's already available to the installed base, such as technical documentation, forums, bug tracking systems, LDAP directories, mailing lists, project plans, etc.

From the index, we then extract candidate topics by applying Natural Language Processing techniques to the unstructured text and by applying queries to the structured information. These candidate topics come with links to relevant documents, people, terms and other information. Users select suitable topics and refine them into the Community Encyclopedia.

The resulting pages look a lot like those in Wikipedia and Knol, but they're a lot more relevant to the community and a lot less effort to produce.

Because of the way it's built, we can also use the encyclopedia to program the network, but that's another story.

Friday, December 21, 2007

Spock attacks the identity crisis

Two days ago I got Spocked! I'm impressed with their approach to solving the identity crisis:

- Index as many external "people" sites (linkedin, plaxo, address books from the largest webmail providers, DBLP bibliography database, etc., etc.)

- Let users invite each other into trust networks

- Let users tag each other with relevant attributes

- Let users vote on tags

- Allow users to tag relations -- i.e., assert a relation between two different people. Other than in Freebase I haven't seen this, but Spock is much more liberal, allowing any type of relation (with the flip side of not establishing inverse relations).

- Use all the resulting metrics (believability of a tag based on votes, authority of a user based on number of correct tags, social metric of people according to relations with other people and their authority) to create an equivalent to pagerank (Spock Power)

- And the kicker; allow users to merge records about people

Obviously the latter one is a huge contributor to solving the identity crisis using collective intelligence. I think Spock recognizes the power of particularly that type of contribution because after I merged a couple of records about people I know, my Spock power went from 161 to 1729!

There has been a lot of discussion on the (lack of) ethics of the Spock team here and obviously they will have major hurdles fighting spam, but the Spock uri (e.g. mine is http://www.spock.com/Jasper-Kamperman-z32I1yG ) might become a pretty usable online identity.

But being in the top 1% of users after only two hours of tagging and merging and the top 1000 leader board having scores as low as 2,500 leads me to believe their number of users is less than a million -- which includes a lot of europeans. So they don't appear to cover a very large part of the web just yet.

Friday, November 16, 2007

Solving the identity crisis

Will the large scale Semantic Web ever happen? As in a significant portion of the web explicitly annotated with classes, taxonomies, properties, even microformats? I'm not holding my breath, and a lot of implicit webbers with me.

Just consider how hard it is to design a sound ontology. How much harder it is to standardize on one ontology. Or to map between ontologies. Then imagine how all web publishers are going to deal with those issues. Hundreds of millions of them, including you and me.

Implicit webbers don't wait for publishers. We like shallow taxonomies and we guess if we need to. And of course we'll accept any help we can get, even if it's labeled "semantic web".

Some of that help may come from OKKAM. These people seem to have found a tractable corner of the Semantic Web. Here's my take on their recipe:
  • Forget about classes, taxonomies, properties and description logics and build a service for just identity -- a taxonomy doesn't get flatter than that.
  • Create a unique ID for Jasper Kamperman. Resist the temptation to classify him as a Person, Male, Musician, Dutchman, Sunnyvale-dweller, Computer Scientist, Openwater Architect. Don't try to maintain his current phone number, address or affiliation. Don't send him email to "update his ecard".
  • But do record that jasper@cwi.nl, jasper.kamperman@cwi.nl, jasper.kamperman@idr.nl, jasper.kamperman@reasoning.com, jasper.kamperman@intel.com, Jasper F. Th. Kamperman, Jasper Kamperman, PhD, http://www.linkedin.com/in/jasperk64 , jasperk64 at yahoo dot com, and jasper dot kamperman at openwaternet dot com all refer to the same entity.
The result is a unique ID for each entity and a large set of clues that helps us make the right guess.

I wonder if or when they'll have an OpenSocial adapter .. imagine the possibilities!

Friday, November 2, 2007

Along came the Implicit Web

A conference about the implicit web inspired me to finally start this blog. The concept isn't very well defined yet, but then again, how much fun would that be?

"Implicit Web" caught my attention because it captures so well what we're doing at Openwater -- building service networks that connect people and information to create better service.

And right there you have an example; how do you know if service is better? That's simple, you conduct a survey. Or do you? Well, I don't know about you, but personally, I detest surveys. And I get really mad at software like WebEx's Meeting Manager that insists on a survey after every single meeting. And rating those movies on Netflix gets old quickly, too. Surveys are annoying exactly because they require you to be explicit.

So how about analyzing service network activity and creating implicit measures of quality?

Take online forums: How long does it take for a question to get an answer? Is it really an answer? How many people read that answer? How many people link to it? Does anyone write "thank you"? Do they come back?

Or take a technique called Search Analytics, explained eloquently here by Gery Angel. How often do users type slight variations of a search query? How many queries return any results at all? Do users click the first, the second or the third search result?

There are endless possibilities in measuring service quality alone, which in itself is only a small but necessary part of what we do. Which brings me back to the reason for starting this blog -- "implicit web" seems to cover a whole swath of subjects we're working on at Openwater. Stay tuned for more.