{"id":24565,"date":"2011-07-22T00:00:00","date_gmt":"2011-07-22T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/improved-handling-of-urls-with-parameters\/"},"modified":"2011-07-22T00:00:00","modified_gmt":"2011-07-22T00:00:00","slug":"improved-handling-of-urls-with-parameters","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/improved-handling-of-urls-with-parameters\/","title":{"rendered":"Improved handling of URLs with parameters"},"content":{"rendered":"<p class=\"gargardate\">Friday, July 22, 2011<\/p>\n<aside class=\"key-point\">It&#8217;s been a while since we published this blog post. Some of the information may be outdated (for example, some images may be missing, and some links may not work anymore).<\/aside>\n<p>\n  You may have noticed that the Parameter Handling feature disappeared from the<br \/>\n  <b>Site configuration <span aria-label=\"and then\">&gt;<\/span> Settings<\/b><br \/>\n  section of Webmaster Tools. Fear not; you can now find it under its new name, URL Parameters!<br \/>\n  Along with renaming it, we refreshed and improved the feature. We hope you&#8217;ll find it even more<br \/>\n  useful. Configuration of URL parameters made in the old version of the feature will be<br \/>\n  automatically visible in the new version. Before we reveal all the cool things you can do with<br \/>\n  URL parameters now, let us remind you (or introduce, if you are new to this feature) of the<br \/>\n  purpose of this feature and when it may come in handy.\n<\/p>\n<h2 id=\"when-to-use\" tabindex=\"-1\">When to use<\/h2>\n<p>\n  URL Parameters helps you control which URLs on your site should be crawled by Googlebot, depending<br \/>\n  on the parameters that appear in these URLs. This functionality provides a simple way to prevent<br \/>\n  crawling duplicate content on your site. Now, your site can be crawled more effectively, reducing<br \/>\n  your bandwidth usage and likely allowing more unique content from your site to be indexed. If you<br \/>\n  suspect that Googlebot&#8217;s crawl coverage of the content on your site could be improved, using this<br \/>\n  feature can be a good idea. But with great power comes great responsibility! You should only use<br \/>\n  this feature if you&#8217;re sure about the behavior of URL parameters on your site. Otherwise you might<br \/>\n  mistakenly prevent some URLs from being crawled, making their content no longer accessible to<br \/>\n  Googlebot.\n<\/p>\n<p><img decoding=\"async\" alt=\"Parameter view for 'page' within the Webmaster Tools URL Parameter tool\" src=\"https:\/\/alienroad.com\/wp-content\/uploads\/kb-gorsel\/g-c023bbf82486.png\" loading=\"lazy\"><\/p>\n<h2 id=\"a-lot-more-to-do\" tabindex=\"-1\">A lot more to do<\/h2>\n<p>\n  Okay, let&#8217;s talk about what&#8217;s new and improved. To begin with, in addition to assigning a crawl<br \/>\n  action to an individual parameter, you can now also describe the behavior of the parameter. You<br \/>\n  start by telling us whether or not the parameter changes the content of the page. If the parameter<br \/>\n  doesn&#8217;t affect the page&#8217;s content then your work is done; Googlebot will choose URLs with a<br \/>\n  representative value of this parameter and will crawl the URLs with this value. Since the<br \/>\n  parameter doesn&#8217;t change the content, any value chosen is equally good. However, if the parameter<br \/>\n  does change the content of a page, you can now assign one of four possible ways for Google to<br \/>\n  crawl URLs with this parameter:\n<\/p>\n<ul>\n<li>Let Googlebot decide<\/li>\n<li>Every URL<\/li>\n<li>Only crawl URLs with value=x<\/li>\n<li>No URLs<\/li>\n<\/ul>\n<p>\n  We also added the ability to provide your own specific value to be used, with the &#8220;Only URLs with<br \/>\n  <code>value=x<\/code>&#8221; option; you&#8217;re no longer restricted to the list of values that we provide.<br \/>\n  Optionally, you can also tell us exactly what the parameter does&mdash;whether it sorts,<br \/>\n  paginates, determines content, etc. One last improvement is that for every parameter, we&#8217;ll try to<br \/>\n  show you a sample of example URLs from your site that Googlebot crawled which contain that<br \/>\n  particular parameter.\n<\/p>\n<p>\n  Of the four crawl options listed above, &#8220;No URLs&#8221; is new and deserves special attention. This<br \/>\n  option is the most restrictive and, for any given URL, takes precedence over settings of other<br \/>\n  parameters in that URL. This means that if the URL contains a parameter that is set to the &#8220;No<br \/>\n  URLs&#8221; option, this URL will never be crawled, even if other parameters in the URL are set to<br \/>\n  &#8220;Every URL.&#8221; You should be careful when using this option. The second most restrictive setting<br \/>\n  is &#8220;Only URLs with <code>value=x<\/code>.&#8221;\n<\/p>\n<h2 id=\"feature-in-use\" tabindex=\"-1\">Feature in use<\/h2>\n<p>Now let&#8217;s do something fun and exercise our brains on an example.<\/p>\n<p>\n  Once upon a time there was an online store, fairyclothes.example.com. The store&#8217;s website used<br \/>\n  parameters in its URLs, and the same content could be reached through multiple URLs. One day the<br \/>\n  store owner noticed, that too many redundant URLs could be preventing Googlebot from crawling the<br \/>\n  site thoroughly. So he sent his assistant CuriousQuestionAsker to The GreatWebWizard to get advice<br \/>\n  on using the URL parameters feature to reduce the duplicate content crawled by Googlebot. The<br \/>\n  Great WebWizard was famous for their wisdom. They looked at the URL parameters and proposed the<br \/>\n  configuration as shown in the following table:\n<\/p>\n<table align=\"center\" border=\"1\" cellpadding=\"3\" cellspacing=\"0\">\n<tbody>\n<tr>\n<th>Parameter name<\/th>\n<th>Effect on content?<\/th>\n<th>What should Googlebot crawl?<\/th>\n<\/tr>\n<tr>\n<td><code>trackingId<\/code><\/td>\n<td>None<\/td>\n<td>One representative URL<\/td>\n<\/tr>\n<tr>\n<td><code>sortOrder<\/code><\/td>\n<td>Sorts<\/td>\n<td>Only URLs with <code>value='lowToHigh'<\/code><\/td>\n<\/tr>\n<tr>\n<td><code>sortBy<\/code><\/td>\n<td>Sorts<\/td>\n<td>Only URLs with <code>value='price'<\/code><\/td>\n<\/tr>\n<tr>\n<td><code>filterByColor<\/code><\/td>\n<td>Narrows<\/td>\n<td>No URLs<\/td>\n<\/tr>\n<tr>\n<td><code>itemId<\/code><\/td>\n<td>Specifies<\/td>\n<td>Every URL<\/td>\n<\/tr>\n<tr>\n<td><code>page<\/code><\/td>\n<td>Paginates<\/td>\n<td>Every URL<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>\n  The CuriousQuestionAsker couldn&#8217;t avoid their nature and started asking questions:\n<\/p>\n<p>\n  <b>CuriousQuestionAsker:<\/b> You&#8217;ve instructed Googlebot to choose a representative URL for<br \/>\n  trackingId (value to be chosen by Googlebot). Why not select the<br \/>\n  <b>Only URLs with <code>value=x<\/code><\/b> option and choose the value myself?<br \/>\n  <b>Great WebWizard:<\/b> While crawling the web Googlebot encountered the following URLs that link<br \/>\n  to your site:\n<\/p>\n<ol>\n<li>fairyclothes.example.com\/skirts\/?trackingId=aaa123<\/li>\n<li>fairyclothes.example.com\/skirts\/?trackingId=aaa124<\/li>\n<li>fairyclothes.example.com\/trousers\/?trackingId=aaa125<\/li>\n<\/ol>\n<p>\n  Imagine that you were to tell Googebot to only crawl URLs where <code>trackingId=aaa125<\/code>.<br \/>\n  In that case Googlebot would not crawl URLs 1 and 2 as neither of them has the value<br \/>\n  <code>aaa125<\/code> for <code>trackingId<\/code>.<br \/>\n  Their content would neither be crawled nor indexed and none of your inventory of fine skirts would<br \/>\n  show up in Google&#8217;s search results. No, for this case choosing a representative URL is the way to<br \/>\n  go. Why? Because that tells Googlebot that when it encounters two URLs on the web that differ only<br \/>\n  in this parameter (as URLs 1 and 2 above do) then it only needs to crawl one of them (either will<br \/>\n  do) and it will still get all the content. In the example above two URLs will be crawled; either 1<br \/>\n  and 3, or 2 and 3. Not a single skirt or trouser will be lost.\n<\/p>\n<p>\n  <b>CuriousQuestionAsker:<\/b> What about the <code>sortOrder<\/code> parameter? I don&#8217;t care if the<br \/>\n  items are listed in ascending or descending order. Why not let Google select a representative<br \/>\n  value?<br \/>\n  <b>Great WebWizard:<\/b> As Googlebot continues to crawl it may find the following URLs:\n<\/p>\n<ol>\n<li>fairyclothes.example.com\/skirts\/?page=1&amp;sortBy=price&amp;sortOrder=&#8217;lowToHigh&#8217;<\/li>\n<li>fairyclothes.example.com\/skirts\/?page=1&amp;sortBy=price&amp;sortOrder=&#8217;highToLow&#8217;<\/li>\n<li>fairyclothes.example.com\/skirts\/?page=2&amp;sortBy=price&amp;sortOrder=&#8217;lowToHigh&#8217;<\/li>\n<li>fairyclothes.example.com\/skirts\/?page=2&amp;sortBy=price&amp;sortOrder=&#8217; highToLow&#8217;<\/li>\n<\/ol>\n<p>\n  Notice how the first pair of URLs (1 and 2) differs only in the value of the<br \/>\n  <code>sortOrder<\/code> parameter<br \/>\n  as do URLs in the second pair (3 and 4). However, URLs 1 and 2 will produce different content:<br \/>\n  the first showing the least expensive of your skirts while the second showing the priciest. That<br \/>\n  should be your first hint that using a single representative value is not a good choice for this<br \/>\n  situation. Moreover, if you let Googlebot choose a single representative from among a set of URLs<br \/>\n  that differ only in their <code>sortOrder<\/code> parameter it might choose a different value each<br \/>\n  time. In the<br \/>\n  example above, from the first pair of URLs, URL 1 might be chosen<br \/>\n  (<code>sortOrder='lowToHigh'<\/code>). Whereas<br \/>\n  from the second pair URL 4 might be picked (<code>sortOrder='highToLow'<\/code>). If that were to<br \/>\n  happen Googlebot would crawl only the least expensive skirts (twice):\n<\/p>\n<ul>\n<li>fairyclothes.example.com\/skirts\/?page=1&amp;sortBy=price&amp;sortOrder=&#8217;lowToHigh&#8217;<\/li>\n<li>fairyclothes.example.com\/skirts\/?page=2&amp;sortBy=price&amp;sortOrder=&#8217; highToLow&#8217;<\/li>\n<\/ul>\n<p>\n  Your most expensive skirts would not be crawled at all! When dealing with sorting parameters<br \/>\n  consistency is key. Always sort the same way.\n<\/p>\n<p>\n  <b>CuriousQuestionAsker:<\/b> How about the <code>sortBy<\/code> value?<br \/>\n  <b>Great WebWizard:<\/b> This is very similar to the <code>sortOrder<\/code> attribute. You want the<br \/>\n  crawled<br \/>\n  URLs of your listing to be sorted consistently throughout all the pages, otherwise some of the<br \/>\n  items may not be visible to Googlebot. However, you should be careful which value you choose. If<br \/>\n  you sell books as well as shoes in your store, it would be better not to select the value<br \/>\n  <code>title<\/code> since URLs pointing to shoes never contain <code>sortBy=title<\/code>, so<br \/>\n  they will not be crawled. Likewise setting <code>sortBy=size<\/code> works well for crawling<br \/>\n  shoes, but not for crawling books. Keep in mind that parameters configuration has influence<br \/>\n  throughout the whole site.\n<\/p>\n<p>\n  <b>CuriousQuestionAsker:<\/b> Why not crawl URLs with parameter <code>filterByColor<\/code>?<br \/>\n  <b>Great WebWizard:<\/b> Imagine that you have a three-page list of skirts. Some of the skirts are<br \/>\n  blue, some of them are red and others are green.\n<\/p>\n<ul>\n<li>fairyclothes.example.com\/skirts\/?page=1<\/li>\n<li>fairyclothes.example.com\/skirts\/?page=2<\/li>\n<li>fairyclothes.example.com\/skirts\/?page=3<\/li>\n<\/ul>\n<p>This list is filterable. When a user selects a color, they get two pages of blue skirts:<\/p>\n<ul>\n<li>fairyclothes.example.com\/skirts\/?page=1&amp;flterByColor=blue<\/li>\n<li>fairyclothes.example.com\/skirts\/?page=2&amp;flterByColor=blue<\/li>\n<\/ul>\n<p>\n  They seem like new pages (the set of items are different from all other pages), but there is<br \/>\n  actually no new content on them, since all the blue skirts were already included in the original<br \/>\n  three pages. There&#8217;s no need to crawl URLs that narrow the content by color, since the content<br \/>\n  served on those URLs was already crawled. There is one important thing to notice here: before you<br \/>\n  disallow some URLs from being crawled by selecting the &#8220;No URLs&#8221; option, make sure that Googlebot<br \/>\n  can access the content in another way. Considering our example, Googlebot needs to be able to<br \/>\n  find the first three links on your site, and there should be no settings that prevent crawling<br \/>\n  them.\n<\/p>\n<p>\n  If your site has<br \/>\n  <a href=\"https:\/\/www.google.com\/support\/webmasters\/bin\/answer.py?answer=1235687\" class=\"external-link\">URL parameters<\/a><br \/>\n  that are potentially creating duplicate content issues then you should check out the new URL<br \/>\n  Parameters feature in Webmaster Tools. Let us know what you think or if you have any questions<br \/>\n  post them to the<br \/>\n  <a href=\"https:\/\/support.google.com\/webmasters\/community\/\" class=\"external-link\">Webmaster Help Forum<\/a>.\n<\/p>\n<p class=\"byline-author\">\n  Written by<br \/>\n  <a href=\"https:\/\/plus.google.com\/109580420505325614989\/about\" rel=\"author\" class=\"external-link\">Kamila Primke<\/a>,<br \/>\n  Software Engineer, Webmaster Tools Team<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Friday, July 22, 2011 It&#8217;s been a while since we published this blog post. Some of the information may be outdated (for example, some images may be missing, and some links may not work anymore). You may have noticed that the Parameter Handling feature disappeared from the Site configuration &gt; Settings section of Webmaster Tools. [&hellip;]<\/p>\n","protected":false},"menu_order":84823,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-24565","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/24565","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/24565\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/26881"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=24565"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=24565"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=24565"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}