{"id":23888,"date":"2007-06-13T00:00:00","date_gmt":"2007-06-13T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/duplicate-content-summit-at-smx-advanced\/"},"modified":"2007-06-13T00:00:00","modified_gmt":"2007-06-13T00:00:00","slug":"duplicate-content-summit-at-smx-advanced","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/duplicate-content-summit-at-smx-advanced\/","title":{"rendered":"Duplicate content summit at SMX Advanced"},"content":{"rendered":"<p class=\"gargardate\">Wednesday, June 13, 2007<\/p>\n<p>\n  Last week, I participated in the duplicate content summit at<br \/>\n  <a href=\"https:\/\/searchmarketingexpo.com\/smx_advanced07\/\" class=\"external-link\">SMX Advanced<\/a>.<br \/>\n  I couldn&#8217;t resist the opportunity to show how<br \/>\n  <a href=\"https:\/\/www.vanessafoxnude.com\/2007\/06\/06\/buffy-in-duplicate\/\" class=\"external-link\">Buffy is applicable to the everday Search marketing world<\/a>,<br \/>\n  but mostly I was there to get input from you on the duplicate content issues you face and to<br \/>\n  brainstorm how search engines can help.\n<\/p>\n<p>\n  A few months ago, Adam wrote a great post on<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/deftly-dealing-with-duplicate-content\/\">dealing with duplicate content<\/a>.<br \/>\n  The most important things to know about duplicate content are:\n<\/p>\n<ul>\n<li>\n    Google wants to serve up unique results and does a great job of picking a version of your<br \/>\n    content to show if your sites includes duplication. If you don&#8217;t want to worry about sorting<br \/>\n    through duplication on your site, you can let us worry about it instead.\n  <\/li>\n<li>\n    Duplicate content doesn&#8217;t cause your site to be penalized. If duplicate pages are detected, one<br \/>\n    version will be returned in the search results to ensure variety for searchers.\n  <\/li>\n<li>\n    Duplicate content doesn&#8217;t cause your site to be placed in the supplemental index. Duplication<br \/>\n    may indirectly influence this however, if links to your pages are split among the various<br \/>\n    versions, causing lower per-page PageRank.\n  <\/li>\n<\/ul>\n<p>\n  At the summit at SMX Advanced, we asked what duplicate content issues were most worrisome. Those<br \/>\n  in the audience were concerned about scraper sites, syndication, and internal duplication. We<br \/>\n  discussed lots of potential solutions to these issues and we&#8217;ll definitely consider these options<br \/>\n  along with others as we continue to evolve our toolset. Here&#8217;s the list of some of the potential<br \/>\n  solutions we discussed so that those of you who couldn&#8217;t attend can get in on the conversation.\n<\/p>\n<h2 id=\"specifying-the-preferred-version-of-a-url-in-the-sites-sitemap-file\" tabindex=\"-1\">Specifying the preferred version of a URL in the site&#8217;s Sitemap file<\/h2>\n<p>\n  One thing we discussed was the possibility of specifying the preferred version of a URL in a<br \/>\n  Sitemap file, with the suggestion that if we encountered multiple URLs that point to the same<br \/>\n  content, we could consolidate links to that page and could index the preferred version.\n<\/p>\n<h2 id=\"providing-a-method-for-indicating-parameters-that-should-be-stripped-from-a-url-during-indexing\" tabindex=\"-1\">Providing a method for indicating parameters that should be stripped from a URL during indexing<\/h2>\n<p>\n  We discussed providing this in either an interface such as Webmaster Tools on in the site&#8217;s<br \/>\n  robots.txt file. For instance, if a URL contains sessions IDs, the webmaster could indicate the<br \/>\n  variable for the session ID, which would help search engines index the clean version of the URL<br \/>\n  and consolidate links to it. The audience leaned towards an addition in robots.txt for this.\n<\/p>\n<h2 id=\"providing-a-way-to-authenticate-ownership-of-content\" tabindex=\"-1\">Providing a way to authenticate ownership of content<\/h2>\n<p>\n  This would provide search engines with extra information to help ensure we index the original<br \/>\n  version of an article, rather than a scraped or syndicated version. Note that we do a pretty good<br \/>\n  job of this now and not many people in the audience mentioned this to be a primary issue.<br \/>\n  However, the audience was interested in a way of authenticating content as an extra protection.<br \/>\n  Some suggested using the page with the earliest date, but creation dates aren&#8217;t always reliable.<br \/>\n  Someone also suggested allowing site owners to register content, although that could raise issues<br \/>\n  as well, as non-savvy site owners wouldn&#8217;t know to register content and someone else could take<br \/>\n  the content and register it instead. We currently rely on a number of factors such as the site&#8217;s<br \/>\n  authority and the number of links to the page. If you syndicate content, we suggest that you ask<br \/>\n  the sites who are using your content to block their version with a robots.txt file as part of the<br \/>\n  syndication arrangement to help ensure your version is served in results.\n<\/p>\n<h2 id=\"making-a-duplicate-content-report-available-for-site-owners\" tabindex=\"-1\">Making a duplicate content report available for site owners<\/h2>\n<p>\n  There was great support for the idea of a duplicate content report that would list pages within a<br \/>\n  site that search engines see as duplicate, as well as pages that are seen as duplicates of pages<br \/>\n  on other sites. In addition, we discussed the possibility of adding an alert system to this<br \/>\n  report so site owners could be notified via email or RSS of new duplication issues (particularly<br \/>\n  external duplication).\n<\/p>\n<h2 id=\"working-with-blogging-software-and-content-management-systems-to-address-duplicate-content-issues\" tabindex=\"-1\">Working with blogging software and content management systems to address duplicate content issues<\/h2>\n<p>\n  Some duplicate content issues within a site are due to how the software powering the site<br \/>\n  structures URLs. For instance, a blog may have the same content on the home page, a permalink<br \/>\n  page, a category page, and an archive page. We are definitely open to talking with software<br \/>\n  makers about the best way to provide easy solutions for content creators.\n<\/p>\n<p>\n  In addition to discussing potential solutions to duplicate content issues, the audience had a few<br \/>\n  questions.\n<\/p>\n<p>\n  <b>If I nofollow a substantial number of my internal links to reduce duplicate content issues,<br \/>\n    will this raise a red flag with the search engines?<\/b> The number of <code>nofollow<\/code><br \/>\n  links on a site won&#8217;t raise any red flags, but that is probably not the best method of blocking<br \/>\n  the search engines from crawling duplicate pages, as other sites may link to those pages. A better<br \/>\n  method may be to block pages you don&#8217;t want crawled with a robots.txt file.\n<\/p>\n<p>\n  <b>Are the search engines continuing the Sitemaps alliance?<\/b> We launched<br \/>\n  <a href=\"https:\/\/www.sitemaps.org\/\" class=\"external-link\">sitemaps.org<\/a><br \/>\n  in November of last year and have continued to meet regularly since then. In April, we added the<br \/>\n  ability for you to let us know about your Sitemap in your robots.txt file. We plan to continue to<br \/>\n  work together on initiatives such as this to make the lives of webmasters easier.\n<\/p>\n<p>\n  <b>Many pages on my site primarily consist of graphs. Although the graphs are different on each<br \/>\n    page, how can I ensure that search engines don&#8217;t see these pages as duplicate since they don&#8217;t<br \/>\n    read images?<\/b> To ensure that search engines see these pages as unique, include unique text on<br \/>\n  each page (for instance, a different title, caption, and description for each graph) and include<br \/>\n  unique alt text for each image. (For instance, rather than use <code>alt=\"graph\"<\/code>, use<br \/>\n  something like <code>alt=\"graph that shows Willow's evil trending over time\"<\/code>.\n<\/p>\n<p>\n  <b>I&#8217;ve syndicated my content to many affiliates and now some of those sites are ranking for this<br \/>\n    content rather than my site. What can I do?<\/b> If you&#8217;ve made your content available without<br \/>\n  payment, you may need to enhance and expand the content on your site to make it unique.\n<\/p>\n<p>\n  <b>As a searcher, I want to see duplicates in search results. Can you add this as an option?<\/b><br \/>\n  We&#8217;ve found that most searchers prefer not to have duplicate results. The audience member in<br \/>\n  particular commented that she may not want to get information from one site and would like other<br \/>\n  choices, but for that case, other sites will likely not have identical information and therefore<br \/>\n  will show up in the results. Bear in mind that you can add the &#8220;&filter;=0&#8221; parameter to the end<br \/>\n  of a Google web search URL to see additional results which might be similar.\n<\/p>\n<p>\n  I&#8217;ve brought back all the issues and potential solutions that we discussed at the summit back to<br \/>\n  my team and others within Google and we&#8217;ll continue to work on providing the best search results<br \/>\n  and expanding our partnership with you, the webmaster. If you have additional thoughts, we&#8217;d love<br \/>\n  to hear about them!\n<\/p>\n<p class=\"byline-author\">Posted by <a href=\"https:\/\/www.vanessafox.com\/\" class=\"external-link\">Vanessa Fox<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Wednesday, June 13, 2007 Last week, I participated in the duplicate content summit at SMX Advanced. I couldn&#8217;t resist the opportunity to show how Buffy is applicable to the everday Search marketing world, but mostly I was there to get input from you on the duplicate content issues you face and to brainstorm how search [&hellip;]<\/p>\n","protected":false},"menu_order":86323,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-23888","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/23888","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/23888\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/26518"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=23888"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=23888"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=23888"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}