I've been thinking of making a "small web" indexer so I'm curious about that. I'm seeing even tiny websites being behind CloudFlare, Anubis etc. these days. (And everyone complaining they're getting hammered by mysterious distributed HTTP traffic!)
You can look for Common Crawl or Open Web Index for dataset sizes and how many URLs those include to get a sense of baseline storage costs, and then 2x that for minimum usability.
It's honestly a bit tough to accept that we got a report of abuse and are still dealing with the aftermath of that after having a single crawler go haywire for a few hours (because we play nice and identify ourselves properly), but these... "mysterious" bots that keep hitting all the servers everywhere thousands of times per day just go on like nothing's happening and "no one"'s to blame.