Silicon Valley Technology Commentary & Archives · Est. 2006 3,031 Posts · 2006–2026

January 11, 2010

January 11, 2010 · 2 MIN READ · BY LOUIS GRAY

In Case of Real-time Clogs, Check Pipes for Stoppages

In Case of Real-time Clogs, Check Pipes for Stoppages

At the conclusion of 2009, I named Pubsubhubbub as the top new Web service to debut in the year, as it best exemplified our rapid adoption of real time technologies to push data from one site to another, and reduce or eliminate the need for downstream aggregation services to repeatedly request updates from source sites. I have configured my blog to be enabled by Pubsubhubbub, share my Google Reader items downstream using a Pubsubhubbub-enabled tool, and recognize the benefits of the protocol by way of fast updates on social networks including Facebook and FriendFeed. So, having grown accustomed to near-instantaneous flow, hiccups in the system rapidly gain my attention.

For the nearly 400 people who follow my @lgshareditems Twitter account, seeing my own personal filter of technology news, it would appear I have taken a hiatus from sharing the best of Google Reader, despite what some might consider a full-time obsession. But this assumption would not be true, as it is simply a symptom of a break in the flow of data.


How Pubsubhubbub Works, More or Less

Similarly, I noticed that my Google Reader shares were not populating FriendFeed in seconds, or even several minutes to hours, after they were shared. To force the updates to make their way into my stream, I needed to force FriendFeed to poll Google Reader and pull down the latest, something I haven't had to do ever since they supported Pubsubhubbub shortly after its introduction. So something was clearly broken.

In parallel, other Pubsubhubbub-enabled items continued to act as they always had. Blog posts flowed as smoothly as ever, and after asking some techies to look into it, they reported the hub itself was fine, but instead, the "Reader/Ping" bridge that updates the Pubsubhubbub hub had found a glitch.


But There Is Currently a Break In One Pipe

While my natural inclination was to wonder if maybe Pubsub had grown to big for its britches, and stretched under load, it was just one of the upstream services that has hit a bump. It'd be worth overlooking, out of deference to the good folks at Google, but given the near-outage of about 48 hours, albeit on a weekend, it's gained awareness, and I have had others ask me what had caused the lags.

The Internet continues to be "a series of tubes", all interconnected to get us our data from one place to another, and even the best, latest, cool toys are interdependent on other sources, and no product, so far, is failsafe. If you've noticed a lag in your shares to FriendFeed and other networks, or you use Reader2Twitter on Twitter, and saw your updates disappear, yes, there is a disturbance in the force right now, and it should be solved soon, from what I understand.

For more on Pubsubhubbub, check out their new video posted last week:

January 10, 2010

January 10, 2010 · 1 MIN READ · BY LOUIS GRAY

Podcast: Infosmack Episode #33, Talking Storage and Social

Podcast: Infosmack Episode #33, Talking Storage and Social

Aiming to show I'm not just a one-trick pony, I gained the opportunity to participate in the Infosmack Podcast this last Friday, with EMC's Mark Twomey, 3Par's Marc Farley, and Greg Knieriemen of Chi Corporation, behind the popular StorageMonkeys blog, which targets enterprise IT folks and other data geeks.

The podcast managed to talk up recent recent acquisitions and rumors of acquisitions in the storage industry, wrapping up the recently-completed CES conference in Las Vegas, and also, how social media is growing in the enterprise, most specifically in network and storage. This isn't the typical startup fan affair, so if you want to hear some geeks talk storage and social, enjoy below. You can also find the podcast at its original URL or subscribe below.


Infosmack Podcast MP3

Subscribe with iTunes

January 10, 2010 · 2 MIN READ · BY LOUIS GRAY

Searchtastic Enables Export of Twitter Search Results

Searchtastic Enables Export of Twitter Search Results


In October, I first introduced Searchtastic, a new search engine for Twitter, whose claim to fame is archiving Tweets longer than the standard database, as well as letting you refine your search through live addition and deletion of keywords. This weekend, the site upgraded with a minor addition that has major potential - the ability to export search results to Microsoft Excel databases with a single click. This opens up significant opportunity for marketing and sales teams alike who are looking to track who discusses their brand or select keywords, and track this data offline.

As social media aware businesses are looking closely at Twitter and other online networks for potential prospects and influencers, it comes as no surprise that it would be desired that these future customers be stored in the company's CRM, or as part of a marketing database, even if not all the information is complete. Also, as third party social media contractors are often hired to find these folks, there becomes a need to compile frequent mentions of your brand and products and report into Marketing or Sales.



Any Searchtastic Result is Now Exportable to an .xls File

Now, I can go to Searchtastic, enter a keyword or string of words, and see the results, powered by standard Twitter search.

For example, I may search for the word Maytag.

On the right side of the search results from Searchtastic, there is a new function that says "Excel Report". If I click that, I am asked to enter a simple number-based Captcha, and the report is downloaded.

Included in the report are all the relevant details around the mentions, including:
  • Login of the individual
  • Name listed on Twitter
  • Homepage provided on their profile
  • Location (if reported)
  • Followers (by number)
  • Following (by number)
  • Total Tweets
  • Tweet Date (of the mention)
  • Tweet (the tweet in its entirety)
  • In Reply To (the URL of a reply, if there is one)
Then, the smart marketer can sort by location, if they have a geographical focus, or by total number of followers, to see perceived "influence" and reach, by total tweets to figure out how frequently the person updates, or by time, if they want to work the list sequentially.

Since our initial report on Searchtastic, the developer has made a number of behind the scenes enhancements, fixing search string bugs, adding the expansion of URLs from URL shorteners, and migrating to more powerful servers to enable a greater number of tweets for indexing. You can check out Searchtastic at http://www.searchtastic.com. It is obvious to me that this new addition could be very powerful and useful. I will absolutely be using this.

January 9, 2010

January 9, 2010 · 5 MIN READ · BY LOUIS GRAY

The Future: Operating System And Application-Neutral Data

The Future: Operating System And Application-Neutral Data

We are now growing accustomed to the concept of the "cloud", where our data will be increasingly stored in Web services, not on local disk, accessible from any computer, operating system or browser. But, despite the adoption of standards from major players storing our personal data, the choice of services causes serious vendor lock-in, as the data suite, be it from Microsoft, Google, Apple or other providers, is not only interpreted by their offerings, but stored there as well. This storage and management of our data makes migration between services incredibly difficult, and still leaves us at the mercy of a large company, whose priorities may not be the same as our own.

The time has come to start on a path to true ownership of data by the individual, reducing applications and Web services to the role of filters and containers, rather than hosts, who can propagate lock-in as these services spread to mobile devices and tablets from their desktop roots.


What makes a digital device mine, be it a laptop or a cellphone or an music player, is the personal content that is stored, and how that data is translated, stored, presented and categorized. Similarly, when we make a choice as to our preferred Web services or technology providers, we are, for the most part, passing our content to them exclusively. While we may have made the data location independent, it is far from being service independent, and any potential future switching will have dramatic impacts on time and productivity, including:
  • Complicated export and import of personal data
  • Differences in the interpretation of data between similar applications
  • Potential loss of metadata between services
  • Reduced backup stability as differing instances of our data is housed at differing services
This headache is a major part of vendor lock-in. Today, when I make a choice as to what brand computer to buy, or what phone to purchase, while I may be committing to a brand or a suite of applications, all I am really doing is asking this product to provide its own interpretation of my data, including:
  • My Contacts and Relationships
  • My Music Files
  • My Videos
  • My Photos
  • My E-mails and Hierarchy
  • My Documents
  • My Bookmarks and Hierarchy
Increasingly more important than the actual data itself is its metadata, or data around data. How is the data structured, meaning... Do I have my e-mails in folders with subfolders and rules? Do I have photos in specific albums? Do I know how often I have played a specific song or genre? When were documents created or last edited?

Today, for the most part, we are choosing from three major service providers, although there are alternatives. We can select Google, who offers Google Contacts, Calendar, YouTube, Picasa, GMail, Google Apps and Google Chrome for the majority of our needs. We can, instead, select Apple and use Address Book, iTunes, iPhoto, Mail, and Safari. Or, we can stick with Microsoft, and leverage Outlook, Windows Media Player, and Internet Explorer. (Or their online equivalents)

Despite standards adoption, not all programs interoperate well. No doubt there remain issues with meetings from Microsoft Exchange being received by Apple Mail, and the integration of Web browser bookmarks and Web history is not shared between Google Chrome and Safari. These minor problems are greatly magnified when you consider the potential for future switching, as you migrate from one platform to another or one computer to another - largely because in every case, the applications themselves, even those that are Web services, are storing our data for us, and interpreting it in their own way.

I think it is time for a change, that lets us own our own data, turning the situation on its head.

Instead of hosting our own data with the service provider of the day, we should host our own data in a standard format, which will be adopted by the leading providers, whose applications will tap into us directly, and pull down our data and its metadata. If I chose to log in with GMail one day, I would authenticate who I was, and GMail would pull down my e-mail stream, complete with e-mail activity history (such as replies and forwards). The data would not be stored on Gmail, but instead be more like a read-only process, whereby changes to data, including sent items, would not be stored in GMail, but written back to my personal "cloud", if you will. Similarly, if I opted to log in to Microsoft Outlook, tapping in to my own, authenticated, account, I could browse my contacts in their application, through their filter, but the data would reside with me.

Hosting one's own personal cloud with our own data is not an end run around large corporations in fear of Big Brother, but instead, for real, true, portability. In this situation, a longtime iPhone user could pick up an Android phone, enter my own personal ID (be it through OpenID or some other standard), and pull down my details into all of Google's native applications. Similarly, I could log in to any Microsoft, Apple or Google powered device and become me, not with my data hosted on the new machine, but with my data being read, like a Web page, on that device, in their own lens.

Even as we on the Web are rallying around these concepts of standards, and the cloud, we are seeing the concept of vendor lock-in be as true as ever. The switching costs from hardware device to hardware device, OS to OS and Web service to Web service remain completely too high, and the way around this problem is to take back our data, make it personal, and enforce standards that get the major players to come aboard. While we may not all have all the broadband access necessary to make this a solution today, it's 2010, and we should be well beyond the same issues we have been facing in computing for the last 20 years.

So how do we make this happen?