About Me

My photo
Chris works for Autonomy Corporation - the innovative leader behind meaning-based computing.
Showing posts with label Google. Show all posts
Showing posts with label Google. Show all posts

Tuesday, August 17, 2010

Space Travel on DVDs

Get this...according to the latest IDC report, reported by ZDNet, last year the world's digital universe grew by 1.2 million petabytes or 1.2 zettabytes. to give you a feel of what that is, one petabyte is equivalent to a stack of DVDs to the moon and back. Those 1.2 zettabytes of data? They'll easily get you half-way to mars. I get the feeling that soon the whole universe will be in reach, (space shuttles will probably charge more than $50 for that extra bag. I'd pack light.)

While cool, space travel via DVDs is neither here nor there. What it does mean is that all of this data needs to be managed (like setting retention policies...which I'll get to in another post soon), deleted (so we don't have to go Jupiter), held (in case of litigation) and easily searchable (you know, in case you want to find something). Let's take the last case as an example. Google currently has over 21 billion web pages indexed, basically, this is the web.

As an enterprise, why does this matter? Well, large enterprises need to be able to archive and index millions of emails received and sent every day. Big companies can easily reach a million emails a day (if the average user sends/receives only 10 emails and you have 100,000...well). That's over a billion documents every 3 years. And, unlike Google, who can crawl web pages at their leisure and do not get penalized if they miss a page, a true solution will ensure 100% capture of your emails and files. The ability to search through something which is even 1/5 the size of the internet, while ensuring capture (and we're not even talking about files yet) is a necessary part of managing risk and litigation preparation.

This is how CEOs and CFOs stay out of jail, they can prove what they did right (or find employees who violate company policy). Think of BP, and what they've gotta be doing right now to prove that they took the necessary steps and are taking the necessary steps in containing the spill, helping communities, and preventing the same thing from happening elsewhere.
Enhanced by Zemanta

Friday, August 13, 2010

Manual Search...Meet Google

On John Wang's Grokify Blog he states:
Manual ICP is a slow process that increases information risk and can lead to under collection, late collection, and spoliation. On the other hand, automatic collection can enable ECA, fast collection, and Matter-based ICP. There is no question that automated collection holds advantages over manual ICP. Given the risks associated with Manual ICP, the courts and industry thought leaders are correct to ask if manual collections are still relevant and defensible.
Now, there is no doubt that manual collection for eDiscovery is slow and unwieldy. eDiscovery 2.0 concedes the point here, yet they rage on:
While there’s no dispute that the “automated” collection methods available in litigation software referenced above have a number of features that make this approach more efficient, the question is whether a “manual” (i.e., custodian based) collection process is somehow less defensible. If this is truly the case, then many midsized companies without the budget to purchase such e-discovery applications will inherently be found deficient – which is a daunting notion.
There is clearly a fundamental misunderstanding here. Mid-sized companies, with their mid-sized amount of employees will pay mid-sized licensing fees for automated collection, eDiscovery and records management software. The proportion they pay scales linearly (both up and down) with the size of their company.

And the pricing tangent misses the point entirely, which is that a combination of automatic and manual collection will be the most thorough method of eDiscovery. Having the ability to automatically collect documents will be necessary in the near future (if not right now). Without enterprise-wide search and automatic collection, it is like searching the web without Google. Instead, with only manual collection, you would be starting at a website and clicking link to link or typing in random URLs until you find the right site. How thorough is that?
Enhanced by Zemanta