Silicon Valley Technology Commentary & Archives · Est. 2006 3,031 Posts · 2006–2026

November 11, 2009

November 11, 2009 · 2 MIN READ · BY LOUIS GRAY

Attacking the Web's Beverly Hills and Schenectady Problem

Attacking the Web's Beverly Hills and Schenectady Problem

Not too long ago, every new site you joined on the Web forced you to provide a daunting array of details about you in order to join. Full pages of pull-down menus asking about your date of birth, your marital status, your home address and other information was standard. But over the last few years, with advents such as OpenID, OpenSocial, Facebook Connect, and more recently, Twitter OAuth, personal identities are becoming portable - letting you sign in with a dedicated login to a new site, and reducing your need to store yet another password.

Kevin Marks, vice president of Web services at BT, formerly of Google and Technorati, relayed at the Defrag Conference this afternoon that under the old way, companies, after accumulating a high number of users, would often find they had an extremely high number of users responding they lived in either Beverly Hills or Schenectady, New York. Why? Because they were saying their zip codes were either 90210 or 12345. They were lying - sick of answering page after page of personal data for yet another Web site.

In the years since, thanks to efforts like OpenSocial, we have seen the rise of Web standards that interoperate, letting you pass along your personal information and credentials to new sites without having to create yet another user name and password.

"Over the last two years, we worked out the sanitization of protocols, so it could fetch things from one site to another," Marks said. "In that time, OpenSocial is up to 1 billion users. There are sites all over the world who are using this."

Marks broke down the solution to the real identity problem into four pieces:
  • Me
  • My Friends
  • What We Do
  • The Flow
Tools like OpenID and WebFinger solve for "Me", Portable contacts, through the unification of the Vcard specification, solve for "My Friends", activity streams solve for "What We Do", and new protocols like AtomPub, PubSubHubbub and Salmon are solving the "Flow". As you know, I have been a big proponent of tools like PubSubHubbub, Salmon and tools like Facebook Connect and Twitter OAuth, as they not only pass along data between sites, but also make data pass between sites more quickly. And while they are causing what could be considered a revolution, it is happening through the simple evolution of activity that is already happening.

"All these standards are empirical standards," Marks said. "We first did this with microformats. We asked what people are doing already, and agreed we would do the same thing."

Now, if you do tell companies you live in Beverly HIlls or Schenectady, New York, there's a greater chance that you really do, and maybe we'll believe you.
November 11, 2009 · 2 MIN READ · BY LOUIS GRAY

Search: Less Useful Due to Massive Info Growth, the Flow?

Search: Less Useful Due to Massive Info Growth, the Flow?

In a forward-looking presentation at the Defrag Conference this morning, Stowe Boyd pushed attendees to think about how the Web would look by the year 2019, with the aid of seeing the massive amounts of change that has taken place over the previous decade. One of Boyd's most-aggressive comments stated that the world of search is falling apart, as the problems it initially aimed to solve have been eroded thanks to the information explosion and the corresponding ease of access to social connections in a world of real time. Without saying that social networks would render the established search giants, irrelevant, he suggested, as he has on his blog frequently in the last few years, that the "flow" will replace the world of Web pages - and change the game on search entirely.

Boyd essentially argued that social tools are in the process of changing the culture. He said people were incentivized to discover breaking news from social friends through networks like Twitter and Facebook, which makes the new "real-time Web" interesting. He further suggested that how one interprets this news to define "meaning" is what will replace search.

One of the biggest reasons he thinks meaning will replace search is that the initial argument for search engines was trying to find the few documents on the Web that were relevant to your query, and now, practically any search can deliver millions of results.

"Search is starting to fail because scarcity has been replaced by infinity," Boyd said. "We are heading toward a world where all the critical information is available publicly, and breaking news is a few seconds away - at the most. We will switch to instead relying on finding things through our social connections - engines of meaning, and the source of what is important."

Assuming social elements are going to trump algorithms and crawlers that power today's engines, Boyd said he believes that the most important dimension is now time, not space - and that for the most part, this dimension is shared.

"We are not sharing space online, we are sharing time," he said. "Our time is increasingly not our own. A shared thread of time will be the norm, and how we will get work done."

This new shared thread of time, or "flow", as Stowe referred to it, is poised to become the replacement for today's static Web pages, a new element in today's social Web, which he pontificated could be "the most defining moment of our civilization."
November 11, 2009 · 3 MIN READ · BY LOUIS GRAY

Skepticism Over Current State of Social Web at Defrag

Skepticism Over Current State of Social Web at Defrag

At the Defrag conference in Denver this morning, there was an acknowledgement that social elements are infiltrating practically every aspect of businesses and interpersonal engagement online, but unlike other events, which have seen a practical hugfest over the latest apps or services, the morning's speakers expressed a great deal of frustration over trying to find real benefits and utility to all the activity that is happening online. Speakers suggested today's tools have a stark lack of context, that businesses are too obsessed with having a complete data set and aren't focused enough on the actability on that data, and that many developers are focused on designing apps that simply don't drive benefits.

Eric Marcoullier, CEO of GNIP, was most direct in his comments, saying that "the business world doesn't give a (crap) about your lifestream app," saying that designing yet another application that sorts all your content online is essentially a list of lists - a list of "my stuff" or "my friends' stuff", which is cute, but not necessarily valuable in decision making.

GNIP is best known for offering managing data collection as a service. The company has seen some ups and downs over the last 18 months, culminating in a significant layoff in September that saw the company reduce staff - cutting seven heads from the dozen on their roster. But since the move, Marcoullier said the last few weeks have been "stellar" in terms of productivity, even as his clients aren't necessarily looking for the answers to data - just more data.

He asked, "Is there an opportunity to drive business decisions and revenue for your company?", saying "Data is useless without effort. When you get data, it is a lot of work to do something useful with it, yet market research companies are obsessed with completeness of data."

Similarly, T.A. McCann, CTO of Gist, said that leading social services, like LinkedIn, have curated millions of nodes, tracking millions of relationships. But for most, it hasn't yet been clear how these connections can be leveraged to drive real daily utility - beyond suggesting new connections and companies that should be known due to shared interests.

Much of these shared interests have been displayed in social streams including Twitter and Facebook, which despite their meteoric rise in visibility, are still struggling to provide more than a simple flow of updates and links.

Tim Young of SocialCast complained, "What I find on Twitter is link vomit, or link carpet bombing and swarming about events. During the day, I get all these links, and the issue is I click the link and there isn't a lot of context. Why did they share this and how did it get here?"

Tim called for a new solution to be built that would save traces and paths of content to help communicate new findings to derive value - something made ever more difficult when the most common real time search repository, Twitter search, is now hosting a database that can track as few as only two days.

And despite many people's claims that finding this data ever more quickly is going to make us more productive as a species, Stowe Boyd dumped on that, saying "the myth of increased productivity is a failed world view," adding, "people will trade personal productivity for connectedness, and they will accept an interrupt to help somebody in their social connections."

That's not to say all is dark. Eric of GNIP promised he was still a huge fan of social media, and Stowe pontificated that the rise of the social Web may already be "the most valuable artifact ever created". But from a raft of useless lifestreaming applications and a gap between link visibility and link utility, the speakers seem to agree that we have a long way to go from today's promises to tomorrow's solutions.
November 11, 2009 · 1 MIN READ · BY LOUIS GRAY

Twitter Plucks Data Management Guru from Yahoo!

Twitter Plucks Data Management Guru from Yahoo!

That Twitter is dealing with massive amounts of data flowing through its servers these days would be an understatement, as the service sees strong growth and significant mindshare. With the company having passed what looks to have been its rockiest struggles over the last twelve months, Twitter is now getting to focus on rolling out some significant new features, from Lists to geolocation, trend definitions and retweets. But the microblogging giant looks like it is taking extra steps to harness the power of its rapidly-expanding data set.

If the company's own team list is to be believed, they just picked up Utkarsh Srivastava, a highly respected senior research scientist at Yahoo!, who is best known for his work on building large-scale distributed systems, specifically his efforts with Hadoop.

Hadoop, similar to the Google File System, is a framework that enables applications to work over distributed server nodes and significant data sets - potentially ranging in the petabytes. Yahoo!, Google's off and on competitor, has been the company most associated with Hadoop. While at Yahoo!, Srivastava was one of the original designers of "Pig", an Apache project for analyzing large data sets, which leveraged Hadoop. (See also the research paper: Pig Latin: A Not-So-Foreign Language for Data Processing)

Srivastava, a PhD graduate from Stanford University in Computer Science, has been working at Yahoo! Research since 2006. (See his home page and LinkedIn profile)

Not knowing what aspects of Twitter Srivastava may be working on, it's premature to assume whether his efforts will be primarily focused on new initiatives, or simply helping the company scale its growth. I can dream and hope that he can be the missing piece that brings Twitter's high potential search engine fully online, but that is no doubt a big project indeed.

Update: This hire has been confirmed by Srivastava and also covered by TechCrunch.