Talks / Native to a Web of Data
A B C D
Title slide: Native to a Web of Data, by Tom Coates.

This write-up approximately follows the revised version of this talk from the Future of Web Apps in February 2006. Its presenter notes have been lightly edited for reading, while the period artwork and examples have been retained.

Hi, My name’s Tom Coates. I’ve spent the last couple of years working for the BBC running a small R&D team in Radio and Music, and I’m currently doing something pretty similar for Yahoo, although I should state straight off that I’ve only been with them a couple of months and this talk is definitely my thoughts and not corporate policy.

When I joined the BBC, I was really very focused on social software as a means to help people collaborate and create together. But over the last couple of years, I’ve become interested in a larger problem: how sites and services connect together in similar ways, collaborating to make something much larger than themselves.

The only other part of my team in London at the moment is esteemed Python nerd Simon Willison, who is in the audience today. Ladies and Gentlemen, Simon Willison. Simon has been incredibly helpful in forcing me to get my head together for this presentation and I’d just like to start off by saying an enormous thank you to him, and to Matt Biddulph and Andy Budd who have been really cool.

Now, a lot of the work that I’m going to talk about today has tended to be lumped together under the flag of Web 2.0, and I think we all know what Web 2.0 is most known for - rounded corners and gradient fills...

Slide reading: Chatsum, Blinksale, Rollyo and Blogger.

But I’m not going to be talking about those things at all. I’m going to be talking about product design at a higher level - about what it means to build a product that fits and works and thrives in its environment. Because the web as an environment is starting to change quite dramatically. It’s starting to become more than the sum of its parts.

It’s that new environment I’m going to be talking about over the next forty-one minutes. In particular I’m going to be talking about three things: what the web is changing into, what we can and should build on it, and the architectural principles that might help us build products native to that environment.

Section slide: A web of data.

Probably the most famous attempt to make sense of all the trends that are happening at the moment is Tim O’Reilly’s piece on “What is Web 2.0” based on a session he hosted at last year’s FOO Camp. Tim and the O’Reilly group also coined the Web 2.0 term. I can really recommend the article - Web 2.0 as a term generates some scepticism from people, but the article itself is way better than most of the commentary that followed.

It’s easy to be cynical about buzzwords like these, but the truth is that we have been seeing a mass of new developments over the last couple of years and they need to be named. That's all it is - an attempt to make order from the chaos, to figure out what the hell is going on in an environment which is seeing tremendous change. I mean think about it, before the bubble, the internet was about communication, then with the arrival of the big companies it became all about sales and publishing, then it started to be about networks. But things have been changing again for a few years and we’re now in a stage when people are trying to make sense of it. Web 2.0 as a phrase is just one of those attempts.

Personally, I think my favorite representation of Web 2.0 is this word cloud below, as made by Markus Angermeier at the end of last year. It’s a kind of mind-mappish attempt to pull together some patterns - the larger things being more important - and I think gestures at how many things seem to be changing and pushing forward concurrently.

Visual from slide 9 of Native to a Web of Data.

I’d argue though, that the term Web 2.0 isn’t strong enough to hold together all these disparate streams. You’ve got a whole lot of activity going on here at different levels - architectural changes, cosmetic changes, environmental changes, technological, social and business changes - and I think it’s over-ambitious to make it all representative of one underlying process.

So for the rest of this talk I'm going to abandon this term and concentrate on a subset where I think the web is really changing, and I use the web here really carefully. Because you only have to look at Google and Yahoo and Flickr and del.icio.us and the various other web start-ups to see that the idea of a web of connected resources has developed pretty quickly over the last few years beyond the simple idea of the navigational hyperlink between published pages and - through APIs, web services and RSS - right down into the heart of the data itself.

We’ve got a new web manifesting here, a web of connected data. So for example, if this is the web that was....

Slide reading: The web as it was.

Then this is the web that seems to be manifesting. A web of data resources, each of which connected to the others around them, able to create more by their combination than they could apart.

Slide reading: Web of the future?.

This is the web I think we’re moving towards and it’s very different from the web we have now. Unlike today, where we have pretty much a web of pages, we’re on our way towards something like a web of data. It's the data sources are becoming more intertwined. Everything is still publishing pages. Pages make things discoverable, interesting, findable. They give people interfaces. But the actual, the data itself is connecting together in web-like ways in a way that it never used to. And I think that's really interesting

By data, I mean the underlying things a service knows about: people, photographs, places, events, products and all the relationships between them. Once those things are exposed, other services can help people explore and manipulate them, and users can connect them together in ways their original creators may never have anticipated.

Visual from slide 12 of Native to a Web of Data.

I think mash-ups are kind of our pilot fish for whatever the web is becoming. In and of themselves they’re fairly interesting, but if you think of them as the beginning of a trend towards interconnected data and reuse they're a much better deal.

Visual from slide 13 of Native to a Web of Data.

Think of them this way - Mash-ups show us that there are things that you can do with two bits of data hybridized together. And that points in a new direction. To what? Well that's less clear. That's why there's a question mark in the diagram.

Let me give you an example: This is a little project that Simon and I knocked up a couple of weeks ago. It's not public, unfortunately, because it's too cool and you couldn't cope with it. It's called Yahoo! Astronewsology. What it does is it takes information from the Yahoo! News site and it takes information from Yahoo! Astrology site. Because what could really be a more solid indicator of the truths of the universe than what the stars tell us? And it splices them together in an interesting way.

Slide reading: Astronewsology.

One way you can navigate your way through it is by - for example - looking at things happened to Leos today! That would be interesting, right? But perhaps more interesting - you can explore a news story and compare what was supposed to happen to someone with what actually did! This is extremely interesting when you're exploring obituaries!

Visual from slide 16 of Native to a Web of Data.

Obviously, in and of itself, this is a pretty trivial example. It was an attempt to show just how crazily you can merge datasets together. But if you think about it more seriously for a second, what we've done here and what most mashups do is they get two disparate data sources and they hybridize them together in ways that make each of those data sources better.

This gives you the ability to navigate one data source in terms of another. And that makes each one of those data sources more useful, more impactful, more valuable.

As you combine these things together, you start to get a network effect. Only here rather than a network effect of connected people, you're getting a network effect of services. And I think that's really important and interesting and foundational to where we're moving.

Or to put it another way - every new service that you create can potentially build on top of every other services that are in existence to create something new, that itself can be a building block for other people. That, in itself, is pretty amazing. And it goes further. Every single service that adds data into that ecosystem enhances all the ones that are already out there.

Visual from slide 17 of Native to a Web of Data.

Now obviously this is exciting from a technical perspective. But I think it's also exciting from a kind of utopian liberal hippie perspective, which—bluntly—is my perspective.

But I think it's also significant in other ways as well. By hybrising things together, I think you get accelerating innovation. Because no one has to go and build the same service twice. Someone can just go out there and build one thing. The next person can build on top of it, connect in with it. And it also then results in more competition, which creates opportunities for businesses and for the creation of more componentized services, and specialized services. Basically, I think it's good for capitalists too.

Visual from slide 18 of Native to a Web of Data.

Which makes me even more certain that things are going to move in this direction. There's money to be made in services! And that's directly charging for them, or by using APIs to drive people to your stuff or as a beachhead for other new services.

Visual from slide 19 of Native to a Web of Data.
Visual from slide 20 of Native to a Web of Data.

Now, when I was working at the BBC, I worked on a project called Programme Information Pages. And this was about representing every programme the BBC produces with a unique addressable web page and building a database of programme information behind the scenes, which people could access and explore. And that project is still ongoing. You can see it at the moment in Radio 3 and Radio 4's websites. Now, if the BBC opens that up as APIs, people will build new ways of navigating and exploring around that kind of stuff. Which means, in the end, whatever anyone builds brings people back to the BBC's programmes. If the BBC wants to get those programmes out in public, there's no better way of doing it for a minimum effort for them. Amazon is the prime example of this.

Slide reading: $$$.

In the end I think everyone ends up playing in the same ecosystem, in the same space. Because if you’re not participating, if you're not part of that ecosystem, then you're not benefitting from the accelerating change, network effects and added value of being part of the web of data.

So my argument is that if you’re part of this ecosystem you will be pulled along and caught up in a web of accelerating value and reuse, whereas if you are not you’ll be stuck in a disconnected backwater.

But what kinds of products work well in this space? How do you decide what to build?

Section slide: Choosing what to build.
Slide reading: What can I build that will make the whole Web better?.
Slide reading: How can I add value to the Aggregate Web?.

So the first question is - can you find way to add data to the aggregate web? Can you create or open up a database of information that already exists, become a definitive home for a particular kind of data on the web? Can you own a kind of data that people want or are prepared to pay for? Can you work with the wider web to help your users create data, to publish or annotate or enhance some things that are already there? Or help them organise a part of their lives, help them turn their own information into data, and share and use it in more powerful ways?

Visual from slide 25 of Native to a Web of Data.

To be more cynical and businesslike, one shift in this ecosystem is going to be towards people trying to control and own certain key types of data, or to become synonymous with it. Tim O’Reilly says it best in this particularly blunt couple of quotes from the What is Web 2.0 piece...

Visual from slide 26 of Native to a Web of Data.

The next way you can add value is by making a service that helps people explore, use or manipulate data in some way. The arrival of weblogs and the popularisation of RSS amounted to pretty much the first improvement to structured data publishing on the popular internet for a long time and loads of people have built on top of the data they’ve created. Similarly Amazon web services and the Flickr APIs have created an incredibly fertile space both for individuals to play creatively and - increasingly - to build businesses on top of other people’s data stores.

Visual from slide 27 of Native to a Web of Data.

At it’s most basic - can you move from people or organisations from capturing and organising their data into making it a more embedded part of this data ecosystem. Can you help them syndicate, help them cross publish - ideally without any extra work - and show them the benefits of being able to connect one sort of data with another? Can you yourselves just join up two bits of data that haven’t been integrated before and make something better out of them?

Visual from slide 28 of Native to a Web of Data.
Section slide: Architectural principles.
Visual from slide 30 of Native to a Web of Data.
Visual from slide 31 of Native to a Web of Data.

Mission: Try and build something that adds value to the aggregate web - improving a data source, finding a new way to connect disparate data sources, build a new interface for manipulating data.

Slide reading: 1 Look to add value to the Aggregate Web of data.

User interface changes in a web of data, because you have at least three types of user - normal humans who are looking to explore or use the information on your site, developers who are looking for the hooks that they can use to build upon your service, and software that's been trained to look for common features and standards directly at the data level. To be part of a web of data you need to build for all of them.

Slide reading: 2 Build for normal users, developers and machines.

Always think about what you're making in terms of data/information and not pages. This sounds like it would scare off real people, but actually the opposite is the case. Having a clear understanding of the information that a page represents is a good thing for normal human users. In both jobs, what you're trying to do is turn information into navigable, explorable, reusable, connectable units.

You're looking for a best of both world's scenario, where you have data that is rich and consistent enough for machines to work with reliably organised in ways that make it explorable and comprehendable to humans.

The process of product design starts with designing the data, and the data structures and relationships. If you don't capture information and relationships that will make it easy to navigate through your application or service then it will never work.

Slide reading: 3 Start by designing explorable data, not pages.

Every page on your site will end up being an addressable view of your data. When I say addressable, I mean that you can point to it, or link to it, or send it to your friends or use it as a marker to stand in for the full contents of the page. So your first job is to understand what the core concepts are that you're going to be working with - whether it be people, addresses, events, photographs, television programmes or whatever - and to give each of these a unique and well structured URL.

These will be your ‘destination’ pages.

Slide reading: 4 Identify your first order objects and make them addressable.
Slide reading: Search, Technorati, Digg and Delicious.

Even if you do no more than that, you're already playing well in the aggregate web of data - search engines and aggregators like Search, Technorati, Digg and del.icio.us can already start usefully aggregating information about each addressable component of your site, based on how people link to and reference them.

But URLs aren't the only kind of addressability that you might need, because not all concepts are created natively in one place on the internet. A weblog post can be identified uniquely by it's URL, as can a photo on Flickr because they're not only the representation of that concept, they are the thing itself.

But what about films, tv shows, books, people, events? All of these things exist independently of the internet and are likely to have multiple and potentially competing representations online. You need a new concept to link all those representations together - to connect up the data produced in different places - and that's the idea of a unique identifier that represents that concept.

So, if you're working with types of data that already have a canonical or recognised authoritative representation or identifier exposed on the web, then build ways to correlate your identifier with the definitive one. If there aren't definitive URLs or identifiers out there already - which in many cases is more likely - then you will derive huge benefits from defining them or competing with dominant players who have coined them already.

Slide reading: 5 Correlate with external identifier schemes (or coin a new standard).
Slide reading: Brokeback Mountain.
Slide reading: 6 Use readable, reliable and hackable URLs.
Visual from slide 43 of Native to a Web of Data.

In the case of using identifiers - if you’re hoping your objects will be reused around the web - or you’re using an object from elsewhere around the web - make it very easy to correlate annotations around them by exposing the identifier that we just talked about.

Visual from slide 44 of Native to a Web of Data.

Some URL schemes are so elegant and powerful that they really offer themselves up as a major interface to the site itself - even to the extent that they have to be pulled into the page as a design element. This kind of approach started on del.icio.us, but I thought it was really interesting that it found itself over on newsvine.

I’m not sure what I think of this approach, but it certainly shows you how powerful the URL can be in terms of supplementing or extending a site’s navigation

... in addition creating a easy to automate way for a piece of software to connect and explore a site.

Slide reading: Newsvine.
Slide reading: Good URLs are beautiful and a mark of design quality.

We’ve got our core first-order objects, and we’ve made them addressable, with a unique web page representing each one, and we’ve correlated those concepts with identifiers on the wider web. Now we have to think about ways in which you’d navigate between them, and ways in which you can manipulate and fiddle with the data you’ve got at your disposal, which is when we get to list views and ways of manipulating data...

Slide reading: 7 Build list views and batch manipulation interfaces.

There are fundamentally only really THREE CORE TYPES of pages that you need to build a web of data native service - or maybe even that’s overstating it. It’s possible you’ll only need two. The core ones are:

Pages representing your first order concepts - which is what we’ve talked about already, addressable concepts. On these pages, one of the bits of data that you’ll want to have captured is explicit relationships to other first order concepts - ie. next in sequence and stuff like that. But the second type is about higher level views and lists of the first order objects - other ways of exploring that dataspace. The third form is really for convenience. If you’re building a service where data manipulation is more core, and there’s a lot of manipulation to do, then you might need a set of dedicated manipulation interfaces. Flickr has one in Organizr.

The concept of manipulation is really important and interesting one, because user manipulation of data is heavily constrained by the interface widgets at your disposal. Which is where new interface technologies like Flash and Ajax can come in.

The most important thing when using either technology is that you should absolutely not break the web. Each of your destination pages here should be addressable with a clear URL that represents a concept. Similarly each of your list view pages should have its own URL as well. Distinct things get distinct pages. Which means that if you’re using Ajax or Flash on a page that’s about a concept, you should only use it to help people manipulate or edit THAT CONCEPT. It’s only in dedicated batch manipulation interfaces that you can go wild with this technology - because individual resources aren’t necessarily supposed to be referenced in the process of their correction.

Visual from slide 48 of Native to a Web of Data.

Flickr does this extremely well - their destination pages each represent a photo and allow you to rotate the photo, add tags and annotate without refreshing the page. But the pages remain referenceable and part of the web. This kind of componentised use is a complete shift from the Flash / DHTML interfaces of yesterday and is all the better for it.

Slide reading: Odeo Ajax / Flash.

In some places manipulation is really as simple as being able to play a rich media object in place in a way that the web normally might not allow. Odeo does this stuff very very well - they incorporate flash and Ajaxy components into pages but never break the addressability and linkability of the web page.

Ajaxy stuff is also really good for representing state changes. This is a sort of frivolous use for a page that is very rapidly updated - it’s still a list view, but it’s one where one of its core attributes is that it’s fast moving. How better to represent that than by having each new attribute appear in real time. But the page still represents the list. The web is not broken.

Slide reading: Digg Spy.
Slide reading: 8 Create parallel data services using understood standards.

So you’ve got your three particular types of page. As Ray Ozzie put it in a leaked Microsoft memo, RSS is a kind of Unix pipe for the internet; APIs provide another parallel representation.

Visual from slide 53 of Native to a Web of Data.
Slide reading: del.icio.us RSS.
Slide reading: Microformats.

Consistent URL structures that are paralleled between api and real - so api.yoursite.com for example, or a really clearly established way of handling the way you get data - this is where exposed identifiers come in again, or autodiscovery and RSS or link rel: tags or literally just exposing it in place like the microformats crew would advocate. All of these are good, and you should already be in the right place to do it...

Slide reading: 9 Make your data as discoverable as possible.
Slide reading: 10 Give everything an appropriate license.

Appropriate licenses obviously depends on how and where you want to benefit from your data - why you’re doing it. The license for your data can be very different depending on the situation you’re in - you’ll need different licenses for each of these uses if you’re a business - and for most businesses with proprietary data, you’ll probably end up having to be more clear about the terms of your reuse.

Visual from slide 58 of Native to a Web of Data.

But it’s always worth thinking about Creative Commons - if only because the clearer and more standard - and more automatable - the licenses, the easier it will be for data to flow through the wider Aggregate web.

Slide reading: Creative Commons.
Visual from slide 60 of Native to a Web of Data.
Slide reading: If you’ve enjoyed this talk....
Slide reading: plasticbag.org.
Slide reading: developer.yahoo.net.