The googlebot monopoly
<p>TIL that Bell Labs and a whole lot of other websites block archive.org, not to mention most search engines. Turns out I have <a href="https://github.com/danluu/debugging-stories/issues/3">a broken website link</a> in a GitHub repo, caused by the deletion of an old webpage. When I tried to pull the original from archive.org, I found that it's not available because Bell Labs blocks the archive.org crawler in their robots.txt:</p> <p></p> <pre><code>User-agent: Googlebot User-agent: msnbot User-