This reconstruction follows the Future of Web Apps presentation given in San Francisco in September 2006. The interface screenshots, figures and predictions are preserved as they appeared then; the speaker notes have been lightly edited for reading.
At the time, I was working for Yahoo in a small rapid-prototyping and innovation unit in London, mostly on social software.
What I wanted to talk about was how people can interact to make something that’s greater than the sum of its parts.
Now, probably luckily for me, social software and ideas like harnessing collective intelligence are quite big at the moment and lots of sites are doing quite well out of them.
But despite the astronomical growth in the popularity of these services some people have recently called the whole social software scene into question and asked what it is really good for. Now, I think I understand why people feel this way, and I think it’s because of this...
This is MySpace and this is the page for celebrated expert in client technologies and co-creator of the Django web framework, Simon Willison.
Now of course the first thing you notice about MySpace is that user pages are eye-bleedingly horrible, so I can totally see how people - particularly those of us over twenty - can look at this stuff and say to themselves, “Where’s the value? Who has time for this kind of stuff?”
But of course it’s incredibly popular and has eaten a whole generation of people who are basically using it to function socially. And to see that you have to step back one stage and see how people interrelate through the software.
Looking close up is often a mistake in these territories.
The same thing happened with weblogs a few years ago - people are seeing the individual bits of stuff that a specific person makes and they don’t see the processes and systems that expose the great stuff, or that focus all that energy in one direction or another.
As far as I’m concerned, social software is the latest in a long line of technologies of collaboration - things that help people operate more effectively together than alone, and which go all the way back to things like:
The creation of money created an abstracted mechanism for trade that transformed ancient tribal chieftains into trading empires. I was a Classicist before I became a technologist and this period in time is particularly interesting.
Money also made it possible for us to collaborate in a market system that straddles the planet. Its effect has been enormously valuable, but it has also had its costs.
What I mean by structured mediation is that it’s the thing in the middle. It’s the thing that sits between all the people who are talking or participating with one another and it makes that participation possible and it can help steer or structure that participation so that something good comes out of the other end of it.
And when you interrogate these in a little more detail you come down to three major territories:
(1) Individual motives - why are people participating
(2) Social value - how you create an experience for multiple people and what happens when it goes wrong
(3) Business or organisational value - why you’re doing it and what it’s good for.
But before I do that, I want to talk briefly about the two major models I know that work in this way - allowing people to build individual value, social value and organisational value at the same time...
Now the most interesting thing about the Wikipedia model is that it really shouldn’t work. It is a ridiculously difficult project to hold together, and I think the only reason it’s managed to do so is because against all the odds it’s managed to create and focus a strong, large and well-organised community of people, create structures for them to work within and a goal towards which they can all strive diligently.
The reason I say all of this is because I’m not sure it’s particularly replicable.
We’re all familiar with Wikipedia, of course.
And probably a good proportion of people are aware that alongside the ease of going in and editing the pages there is a bit more of a system behind the scenes that helps things tick over. But I wonder how many of you are aware of how large it is and how complex.
The scale of the administrative and organisational work around Wikipedia is quite phenomenal, and I suspect a direct consequence of the fact that it has a strong, altruistic and highly focused objective. I think without that it would fall apart.
Now don’t get me wrong, there are other sites that operate in this territory, like OpenStreetMap and...
MusicBrainz, but they are the exception rather than the rule and they are orders of magnitude smaller than Wikipedia itself.
For the most part, online communities of people - and by this I mean where individuals come together in a common space to communicate or collaborate - tend to have a natural lifecycle and over time decay as the people within the group’s interests change or they move on, or they get stale or overrun by new people or spammers.
It’s my contention here that unless you have a really solid mission that’s honourable and grand and unless you’re prepared to develop the kinds of structures that Wikipedia has, that most attempts at large-scale consensus-driven projects will tend towards failure. This is why, I think, there is only really one Wikipedia.
Which brings me to the polyphonic approach.
Now I strongly believe that this model for the most part operates much more effectively than the consensus model, and can be adapted much more readily to different projects.
My evidence is the sheer number of projects that already operate on the principle that by getting innumerable contributions from your users you can generate some kind of order or value from it.
You’ve got here YouTube, del.icio.us, Plazes, HSX, Amazon and last.fm - all places where you can contribute in an individual way in accordance with things that make sense and are useful to you (almost without consideration to the things that other people are doing around you) and yet something emerges out of this space that’s useful to everyone.
And you can see in sites like Flickr how it manages to avoid the problems of one community gradually falling apart.
The polyphonic model - when structured with social networks - can support as many different communities operating in parallel as it has users.
You can also layer interest communities onto the polyphonic model - and these communities can be born, have a lifecycle and die without the larger project collapsing.
With every contribution - put up for your own purposes, or for your friends or for your tiny hat community, the larger commons is enriched and survives and becomes more valuable.
So if you’re thinking about exploring this space, I recommend looking first for models that operate in this kind of territory.
But none of this matters if nobody contributes. So the next question is one of user motives: why contribute at all?
A while back, someone in the industry told me they understood only three types of motivation: money, power and winning, and that these were fundamental to all human interaction.
That did not match my experience, so I investigated the territory more thoroughly and gathered several other perspectives.
More seriously, writing about virtual communities, message boards and Usenet identified four reasons why people contribute valuable things for free. There are useful ideas here, although they do not quite describe the whole picture.
More usefully for our purposes is Steve Weber’s book on the success of open source. He gave these motivations.
Again there are some useful motivations here, but I don’t think they really get what’s going on with our stuff.
Different projects need different motives. OpenStreetMap, for example, needs definitive data, so altruism and contributing to the common good matter. Other services can draw on saving something for personal use, sharing with friends, joining an interest community, self-expression or simply showing off.
The important thing is to understand what value a contribution creates for the person making it, then how those individual acts accumulate into something larger.
Three motives that are perhaps a little more shaky or clumsy are money, points and competition, and I’d advise you to be wary of these. There are a few reasons for this. First, it is normally extremely hard to correlate simple abstracted scoring mechanisms to the kinds of actions you’re trying to improve. So you end up actually promoting the wrong kinds of behaviour. If you can completely correlate them, of course, then that’s great, and you may find that points work for you.
Using money is a good example of how these things go wrong. Most users of most sites are unlikely to derive a lot of money from their contributions. But if there’s money in the picture then some users will do anything to get it - and so you become a magnet for spammers. So I’d avoid using money as a motivator for anything other than sites built specifically around improving or facilitating actual sales of goods. Money works fine as a motivator for Amazon and eBay, because there are physical things that need to be sold and redistributed. But it would be a terrible motivator for something like Plazes for example...
There’s a famous article by Richard Bartle, the guy who created the world’s first MMORPG, about the various kinds of users that you need to create a successful MUD, and he describes their motivations as follows:
(1) Achievement with the game context (Diamonds)
(2) Exploration of the game (Spades)
(3) Socialising with others (Hearts)
(4) Imposition upon others (Clubs)
The most interesting thing about this article is that it proposes that people move between these motivations, and that an imbalance between the different kinds of participant makes the environment less enjoyable for everyone.
Rewarding only one group of users in pretty much any system will create an imbalanced environment and severely damage your ecosystem. It is often better to rely on interpersonal rewards that a community establishes for itself.
Digg’s a really good example of a site that could fall prey to problems like this and which supports a variety of different motivations. It requires at least two different types of user and probably supports more than that. It requires people who will go out and specifically find new links to feed it, and a larger group of individuals who will go through the site and digg the things they think are of quality. That latter group probably subdivides a number of times as well, with a large proportion of users probably only digging things in the territories that they browse around in every day - so they’re probably a bit of an echo-chamber. The site also requires people who are more interested in digging deeper and reviewing the newly submitted sites and bumping them up to the top.
The motivations in each group will be different and supporting one over all of the others could only result in a worse experience. There’s an old rule of thumb - anything you measure goes up - which applies here. Make it a good thing to submit new links, you’ll get spammed to death. Make it a good thing to digg something and you’ll get a lot of the wrong things dugg. Make it a good thing to be the first to spot things that get big and - to an extent - you’ll push a community towards uniformity. Balancing these needs, and supporting multiple ones is the only way to get this stuff to work.
Digg does have a ‘top users’ section, but it is subdued and kept in the background and most importantly it reports on a number of different ways to evaluate a user’s success on the site. If you’re going to do it, this is normally the way to go about it.
We now have some inkling of why an individual might be sufficiently invested to contribute. The next question is where their peers get value from each contribution.
One obvious answer is to analyse the data people submit and use it to improve the service. Search engines, for example, can learn from the words people search for and the results they choose.
They look at the words that people search for and the results that they click on and weigh their results accordingly. But while this is useful, it’s limited compared to the stuff that you can do if you expose all the content and data directly to users.
The other set of social value is stuff that I’ve already touched on - it’s the value of the social contexts of sharing - the ways that this software enables people to build relationships - and perhaps more importantly maintain existing relationships.
From MySpace, where the social glue and maintenance is the most interesting element through to Flickr where people communicate through photos revealing their ambient presence, to last.fm where they communicate through music and del.icio.us and the weblog world where they communicate and relate to one another through links and writing - the social interpersonal communicative element is one of the most powerful things about this entire scene.
But more interesting still is the way that these motives can be fed back upon the data that people have shared initially to create a resource of even more significant data that’s of value to the whole community, and eventually to you as a business as well.
Let’s look briefly at last.fm.
Here’s my profile page on last.fm. Whenever I play a song using iTunes, a small application on my computer sends last.fm a little blob of information recording that I have done so.
On this page you can see not only my beautiful bearded visage, but also the songs I was listening to while trying to finish this bloody talk.
But why would I want to do this? First, there’s personal value.
I get interesting and exciting nerdy lists of the songs I’ve been listening to most recently, for any week in history or ever - for any music obsessive this is fascinating, by the way.
Or I can download their client and listen to a form of heavily personalised radio that will occasionally play songs that I in particular might find interesting or like. So that’s an interesting service too. Yahoo has a similar service by the way but I figured you’d lynch me if I only talked about Yahoo-related stuff throughout the talk.
But it also has a communicative or social value - you can see my friends on the right over there - at any time I can go and see what they’re listening to recently or have been listening to most often recently.
On this page, I can get a sense of the presence of my friends: what they’re doing, and what they have been listening to. If something looks bizarre, I can message them through the site or over IM.
This is where the first bit of aggregate value starts to emerge. I can explore through their music, find new things that I might like and eventually arrive on the page for an artist.
Here I can see similar artists, worked out from other people’s listening; a biography that anyone can edit and improve; user tags that describe the music; journal entries about the artist; and charts showing how often they have been played.
And here are the top tracks for last week and the last six months, some people who particularly like that artist and places where I can go and talk about those artists with other people.
And this exists for pretty much every single artist you can ever imagine - and it is a profoundly useful corpus of data, created collectively and pretty much available to all. It’s good for music discovery, for collaborative filtering, for understanding what might be good or bad, or popular or not popular, for getting a variety of opinions or perspectives or meeting new people who share your tastes.
That annotated data space is pretty much all you need to make sense of an enormous territory like music in ways that would be almost impossible to manage and scale in a top-down way. In fact it’s so useful I’ve always been completely mystified as to why Apple haven’t bought them. It would seem like a completely natural fit between their player and their store.
To get another sense of how productive this space can be, look at what Flickr and its users created together.
It was already an enormous, ever-growing repository of photographs: with clear ownership and permissions, location data, tags and signals about which images people found interesting. That had obvious implications for businesses such as stock photography.
Because you’re trying to generate as much useful data as humanly possible - data that humans have touched and rubbed themselves against and bent a bit to their will.
Before continuing, I want to mention a few things that can go wrong when you try to derive aggregate data from individual contributions.
The issue is not simply whether something is private or public, but whether people understand the boundary and expect their data to be used in that way. Last.fm probably collected more useful, personally identifiable information about a participant than AOL held about a typical search user. But because people understood the circumstances in which it would be used, they went in with their eyes open.
The Facebook example from the week of this talk is a good illustration. The company had just introduced a feature which displayed your activity around the site. This caused a hell of a lot of anxiety around the place, not because the data that was exposed wasn’t previously ‘in public’ but because previously it was hard to discover. Users entered it with one expectation and found it displayed back to them in another. This is enough to freak people out.
I’d recommend thinking about this in any of those cases where you’re encouraging people to save for personal use in public. Del.icio.us users understand what they’re doing, but it’s a harder thing to explain to naive users.
Not every user needs to participate for a service to generate social value. On Wikipedia, a relatively small group made most edits, while a much wider population added material occasionally or anonymously; the number of readers was larger again by several orders of magnitude.
Data was already an important part of the web’s ecosystem, and my belief was that it would become still more important as the web moved towards a set of interconnected data sources.
This is a diagram I used to indicate how the web worked at the time - a set of interconnected pages each fuelled by the data and content management behind the scenes.
This is what I thought the web was moving towards: data sources connecting in formal, informal, commercial and non-commercial ways, with new sites drawing on services and information from across the ecosystem.
You do not have to accept that whole prediction for the rest of the argument to make sense, but I thought the progression would look something like this.
At the time, large proprietary data providers were doing well because they controlled repositories of information that were hard and expensive to replicate, and could license access to them.
But large, open and distributed projects, whether founded for the common good or to make money, were already beginning to challenge those providers through exactly the mechanisms we have been discussing.
We’ve already seen Wikipedia rise and pretty much defeat Britannica in the online encyclopedia wars. We’ll see how that goes in the future but at the moment I’m pretty confident.
MusicBrainz is a project that seems increasingly to be gaining its feet and could be seriously challenging data providers like CDDB in a few years’ time.
OpenStreetMap - although in its very earliest stages - could potentially come up to challenge data providers like NavTeq, MapQuest or Britain’s Ordnance Survey.
I’d say at the moment that’s a bit of a stretch, but I’m not counting it out quite yet.
And in a way Flickr’s doing the same thing - it’s working with its users to make something truly greater than the sum of its parts and finding new ways to make money.
It’s a new model that doesn’t rely on owning content, but instead on premium accounts, careful advertising and services that help you do things with the photos that you’ve uploaded into the community. And yet every day they become a more valuable and interesting resource on the web.
So where was the money? At the time I saw four main territories. But I also wanted to go out on a limb and propose something more far-reaching.
In a world where data becomes increasingly important, proprietary data can become core infrastructure for the web. But perhaps the game-changing move is to throw that model away.
Perhaps the future is not in owning the data, but in creating an ecosystem filled with data that has bubbled up from the people using it.
A business could then offer new services by becoming the principal facilitator of a more collaborative relationship between company and consumer.
If so, the future of web applications might lie in working with users to create something collectively from which everyone can benefit: personally, socially and organisationally.
And that’s where I’m going to leave things.



































































































