如何防止网站刮取?

我有一个相当大的音乐网站，有一个很大的艺术家数据库。我一直注意到其他音乐网站在窃取我们网站的数据(我在这里和那里输入假艺人的名字，然后进行谷歌搜索)。

如何防止屏幕刮擦?这可能吗?

当前回答

你可以做一些事情来防止屏幕抓取。有些不是很有效，而另一些(验证码)是，但阻碍可用性。你必须记住，它也可能阻碍合法的网站刮刀，如搜索引擎索引。

然而，我认为如果你不希望它被删除，这意味着你也不希望搜索引擎索引它。

这里有一些你可以尝试的方法:

Show the text in an image. This is quite reliable, and is less of a pain on the user than a CAPTCHA, but means they won't be able to cut and paste and it won't scale prettily or be accessible. Use a CAPTCHA and require it to be completed before returning the page. This is a reliable method, but also the biggest pain to impose on a user. Require the user to sign up for an account before viewing the pages, and confirm their email address. This will be pretty effective, but not totally - a screen-scraper might set up an account and might cleverly program their script to log in for them. If the client's user-agent string is empty, block access. A site-scraping script will often be lazily programmed and won't set a user-agent string, whereas all web browsers will. You can set up a black list of known screen scraper user-agent strings as you discover them. Again, this will only help the lazily-coded ones; a programmer who knows what he's doing can set a user-agent string to impersonate a web browser. Change the URL path often. When you change it, make sure the old one keeps working, but only for as long as one user is likely to have their browser open. Make it hard to predict what the new URL path will be. This will make it difficult for scripts to grab it if their URL is hard-coded. It'd be best to do this with some kind of script.

如果我必须这样做，我可能会结合使用后三种方法，因为它们最大限度地减少了对合法用户的不便。然而，你必须接受这样的事实:你不可能用这种方式屏蔽所有人，一旦有人想出了绕过它的方法，他们就可以永远地刮掉它。我猜你可以在发现他们的时候屏蔽他们的IP地址。

2010-07-02 00:42:10

其他回答

生成HTML, CSS和JavaScript。编写生成器比编写解析器更容易，因此可以以不同的方式生成每个服务页面。这样就不能再使用缓存或静态内容了。

2010-07-02 01:30:54

对不起，这真的很难做到……

我建议你礼貌地要求他们不要使用你的内容(如果你的内容是受版权保护的)。

如果是这样，他们不把它撤下来，那么你可以采取进一步的行动，给他们发一封停止通知信。

一般来说，无论你做什么来防止抓取可能最终会产生更负面的影响，例如可访问性，机器人/蜘蛛等。

2010-07-01 20:50:51

当然，这是可能的。为了100%的成功，让你的网站离线。

在现实中，你可以做一些事情，让抓取变得更加困难。谷歌进行浏览器检查，以确保您不是一个抓取搜索结果的机器人(尽管这和大多数其他事情一样，可以被欺骗)。