Comment Re:I didn't know (Score 5, Informative) 23
The idea is to give everyone access to crawl data. If you work at a large search company, you have access to crawl data. You can also set up crawlers to get the data yourself, but that is expensive and having countless crawlers doing duplicative work is not ideal. Our idea is that there should be one common repository for crawl data that anyone can use. Researchers are using it for NLP, IR, sentiment analysis and many other things like measuring the adoption of metadata formats http://www.webdatacommons.org/ Educators are using it as a real world dataset to teach big data techniques in the classroom. Developers and entrepreneurs are using it for startups.
Sorry I don't have a car analogy :) Feel free to email me if you have any other questions lisa at commoncrawl dot org