Yahoo Interview Question

Clustering on large, sparse datasets – how would you approach the problem?