Maybe they want Google to be able to crawl their database (which it has clearly done, as you'll see if you do a search.) That also raises some questions...
By a single search engine which probably provides the vast majority of their traffic.
While I don't necessarily agree with the concept of only allowing google to index your site, comparing a search engine which feeds you business to a company reselling your data with no attribution is not really fair in my opinion.
Neither of the two parties mentioned is likely to resell the data directly, both are likely to to create a derivative work which they will exploit commercially as will Google. What part of the ToS make one OK while and the other forbidden? How does this work when a smart scraper can just pull from the Google or archive.org cache?
Probably not, but if they don't want to be scraped, the first place they should notify about this is robots.txt. I was just stating an example that you can allow some bots and not others if you like.
Of course, forging your user agent, disobeying robots.txt or scraping after you were told not to, is wrong ethically.