Hacker News new | past | comments | ask | show | jobs | submit
How do you cleanse the data at this scale?
By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning/getting more out of existing noisy data.

We have a lot of details in the tech report if you want to go deeper.

loading story #49946839
loading story #49947466