{"id":23992,"date":"2008-06-09T00:00:00","date_gmt":"2008-06-09T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/duplicate-content-due-to-scrapers\/"},"modified":"2008-06-09T00:00:00","modified_gmt":"2008-06-09T00:00:00","slug":"duplicate-content-due-to-scrapers","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/duplicate-content-due-to-scrapers\/","title":{"rendered":"Duplicate content due to scrapers"},"content":{"rendered":"<p class=\"gargardate\">Monday, June 09, 2008<\/p>\n<p>\n  Since duplicate content is a hot topic among webmasters, we thought it might be a good time to<br \/>\n  address common questions we get asked regularly at conferences and on the<br \/>\n  <a href=\"https:\/\/support.google.com\/webmasters\/go\/community\" class=\"external-link\">Google Webmaster Help Group<\/a>.\n<\/p>\n<p>\n  Before diving in, I&#8217;d like to briefly touch on a concern webmasters often voice: in most cases a<br \/>\n  webmaster has no influence on third parties that scrape and redistribute content without the<br \/>\n  webmaster&#8217;s consent. We realize that this is not the fault of the affected webmaster, which in<br \/>\n  turn means that identical content showing up on several sites in itself is not inherently regarded<br \/>\n  as a violation of our<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/search-essentials-overview\/\">webmaster guidelines<\/a>.<br \/>\n  This simply leads to further processes with the intent of determining the original source of the<br \/>\n  content&mdash;something Google is quite good at, as in most cases the original content can be<br \/>\n  correctly identified, resulting in no negative effects for the site that originated the content.\n<\/p>\n<p>\n  Generally, we can differentiate between two major scenarios for issues related to duplicate<br \/>\n  content:\n<\/p>\n<ul>\n<li>\n    Within-your-domain-duplicate-content, that is, identical content which (often unintentionally)<br \/>\n    appears in more than one place on your site\n  <\/li>\n<li>\n    Cross-domain-duplicate-content, that is, identical content of your site which appears (again,<br \/>\n    often unintentionally) on different external sites\n  <\/li>\n<\/ul>\n<p>\n  With the first scenario, you can take matters into your own hands to avoid Google indexing<br \/>\n  duplicate content on your site. Check out Adam Lasnik&#8217;s post<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/deftly-dealing-with-duplicate-content\/\">Deftly dealing with duplicate content<\/a><br \/>\n  and Vanessa Fox&#8217;s<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/duplicate-content-summit-at-smx-advanced\/\">Duplicate content summit at SMX Advanced<\/a>,<br \/>\n  both of which give you some great tips on how to resolve duplicate content issues within your<br \/>\n  site. Here&#8217;s one additional tip to help avoid content on your site being crawled as duplicate:<br \/>\n  include the preferred version of your URLs in your Sitemap file. When encountering different pages<br \/>\n  with the same content, this may help raise the likelihood of us serving the version you prefer.<br \/>\n  Some additional information on duplicate content can also be found in our comprehensive<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/specify-canonical\/\">Help Center article<\/a><br \/>\n  discussing this topic.\n<\/p>\n<p>\n  In the second scenario, you might have the case of someone scraping your content to put it on a<br \/>\n  different site, often to try to monetize it. It&#8217;s also common for many web proxies to index parts<br \/>\n  of sites which have been accessed through the proxy. When encountering such duplicate content on<br \/>\n  different sites, we look at various signals to determine which site is the original one, which<br \/>\n  usually works very well. This also means that you shouldn&#8217;t be very concerned about seeing<br \/>\n  negative effects on your site&#8217;s presence on Google if you notice someone scraping your content.\n<\/p>\n<p>\n  In cases when you are syndicating your content but also want to make sure your site is identified<br \/>\n  as the original source, it&#8217;s useful to ask your syndication partners to include a link back to<br \/>\n  your original content. You can find some additional tips on dealing with syndicated content in a<br \/>\n  recent post by Vanessa Fox,<br \/>\n  <a href=\"https:\/\/www.vanessafoxnude.com\/2008\/05\/14\/ranking-as-the-original-source-for-content-you-syndicate\/\" class=\"external-link\">Ranking as the original source for content you syndicate<\/a>.\n<\/p>\n<p>\n  Some webmasters have asked what could cause scraped content to rank higher than the original<br \/>\n  source. That should be a rare case, but if you do find yourself in this situation:\n<\/p>\n<ul>\n<li>\n    Check if your content is still accessible to our crawlers. You might unintentionally have<br \/>\n    blocked access to parts of your content in your robots.txt file.\n  <\/li>\n<li>\n    You can look in your Sitemap file to see if you made changes for the particular content which<br \/>\n    has been scraped.\n  <\/li>\n<li>Check if your site is in line with our webmaster guidelines.<\/li>\n<\/ul>\n<p>\n  To conclude, I&#8217;d like to point out that in the majority of cases, having duplicate content does<br \/>\n  not have negative effects on your site&#8217;s presence in the Google index. It simply gets filtered<br \/>\n  out. If you check out some of the tips mentioned in the resources above, you&#8217;ll basically learn<br \/>\n  how to have greater control about what exactly we&#8217;re crawling and indexing and which versions are<br \/>\n  more likely to appear in the index. Only when there are signals pointing to deliberate and<br \/>\n  malicious intent, occurrences of duplicate content might be considered a violation of the<br \/>\n  webmaster guidelines.\n<\/p>\n<p>\n  If you would like to further discuss this topic, you can visit our<br \/>\n  <a href=\"https:\/\/support.google.com\/webmasters\/go\/community\" class=\"external-link\">Webmaster Help Group<\/a>.\n<\/p>\n<p class=\"byline-author\">Written by Sven Naumann, Search Quality Team<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Monday, June 09, 2008 Since duplicate content is a hot topic among webmasters, we thought it might be a good time to address common questions we get asked regularly at conferences and on the Google Webmaster Help Group. Before diving in, I&#8217;d like to briefly touch on a concern webmasters often voice: in most cases [&hellip;]<\/p>\n","protected":false},"menu_order":85961,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-23992","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/23992","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/23992\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/26601"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=23992"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=23992"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=23992"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}