Hacker News new | past | comments | ask | show | jobs | submit
It is not just Amazon (although they're #4 on my list of bad actors)[1].

1. https://files.littlebird.com.au/bad-scrapers.png

Perhaps you need a TOS that requires a usage fee be mailed to a PO Box that has a reasonable "free" limit of like 100$ and when they exceed it, you automatically mail them a copy of the TOS, the logs and an invoice.

Just make sure you have a good lawyer.

is there a law they’re breaking?

because idk i could be wrong but some small project vs a 2.5T market cap company is gonna need more than “a good lawyer”

loading story #49155813
loading story #49155757
Yeah it's not just Amazon, but so far Amazon was the only one specifically looking for URLs in source code. Interestingly it ignored URLs in the fake markdown docs.

Another detail: the scraper did not attempt to access the endpoints immediately (as it did for hrefs in htmls) but it did it on the day after, twice.

Was it actually an Amazon bot IP[1] or someone pretending to be on AWS?

1. https://developer.amazon.com/amazonbot/searchbot-ip-addresse...

Yup, already mentioned in a comment below: the IPs are in that list.

They also show up in AbuseIPDB with multiple reports.