{"id":24620,"date":"2011-11-01T00:00:00","date_gmt":"2011-11-01T00:00:00","guid":{"rendered":"https:\/\/alienroad.com\/google-bilgi-bankasi\/get-post-and-safely-surfacing-more-of-the-web\/"},"modified":"2011-11-01T00:00:00","modified_gmt":"2011-11-01T00:00:00","slug":"get-post-and-safely-surfacing-more-of-the-web","status":"publish","type":"ar_kb","link":"https:\/\/alienroad.com\/google-bilgi-bankasi\/get-post-and-safely-surfacing-more-of-the-web\/","title":{"rendered":"GET, POST, and safely surfacing more of the web"},"content":{"rendered":"<p class=\"gargardate\">Tuesday, November 01, 2011<\/p>\n<p>\n  As the web evolves, Google&#8217;s crawling and indexing capabilities also need to progress. We<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/improved-flash-indexing\/\">improved our indexing of Flash<\/a>, built<br \/>\n  a more robust<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/our-new-search-index-caffeine\/\">infrastructure called Caffeine<\/a>,<br \/>\n  and we even started<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/crawling-through-html-forms\/\">crawling forms<\/a> where it makes<br \/>\n  sense. Now, especially with the growing popularity of JavaScript and, with it, AJAX, we&#8217;re<br \/>\n  finding more web pages requiring <code>POST<\/code> requests&mdash;either for the entire content of<br \/>\n  the page or because the pages are missing information and\/or look completely broken without the<br \/>\n  resources returned from <code>POST<\/code>. For Google Search this is less than ideal, because when<br \/>\n  we&#8217;re not properly discovering and indexing content, searchers may not have access to the most<br \/>\n  comprehensive and relevant results.\n<\/p>\n<p>\n  We generally advise to use<br \/>\n  <a href=\"https:\/\/www.google.com\/search?q=GET+POST+HTTP\" class=\"external-link\"><code>GET<\/code><\/a><br \/>\n  for fetching resources a page needs, and this is by far our preferred method of crawling. We&#8217;ve<br \/>\n  started experiments to rewrite <code>POST<\/code> requests to <code>GET<\/code>, and while this<br \/>\n  remains a valid strategy in some cases, often the contents returned by a web server for<br \/>\n  <code>GET<\/code> vs. <code>POST<\/code> are completely different. Additionally, there are<br \/>\n  legitimate reasons to use <code>POST<\/code> (for example, you can attach more<br \/>\n  data to a <code>POST<\/code> request than a <code>GET<\/code>). So, while <code>GET<\/code> requests<br \/>\n  remain far more common, to surface more content on the web, Googlebot may now perform<br \/>\n  <code>POST<\/code> requests when we believe it&#8217;s safe and appropriate.\n<\/p>\n<p>\n  We take precautions to avoid performing any task on a site that could result in executing an<br \/>\n  unintended user action. Our <code>POST<\/code> requests are primarily for crawling resources that<br \/>\n  a page requests automatically, mimicking what a typical user would see when they open the URL in<br \/>\n  their browser. This will evolve over time as we find better heuristics, but that&#8217;s our current<br \/>\n  approach.\n<\/p>\n<p>\n  Let&#8217;s run through a few <code>POST<\/code> request scenarios that demonstrate how we&#8217;re improving<br \/>\n  our crawling and indexing to evolve with the web.\n<\/p>\n<h2 id=\"examples-of-googlebots-post-requests\" tabindex=\"-1\">Examples of Googlebot&#8217;s <code>POST<\/code> requests<\/h2>\n<ul>\n<li><i>Crawling a page via a POST redirect<\/i>\n<div><\/div>\n<\/li>\n<li>\n    <i>Crawling a resource via a <code>POST<\/code> <code>XMLHttpRequest<\/code><\/i>: In this<br \/>\n    step-by-step example, we improve both the indexing of a page and its Instant Preview by<br \/>\n    following the automatic <code>XMLHttpRequest<\/code> generated as the page renders.\n  <\/li>\n<ol>\n<li>Google crawls the URL, yummy-sundae.html.<\/li>\n<li>\n      Google begins indexing yummy-sundae.html and, as a part of this process, decides to attempt<br \/>\n      to render the page to better understand its content and\/or generate the Instant Preview.\n    <\/li>\n<li>\n      During the render, yummy-sundae.html automatically sends an XMLHttpRequest for a resource,<br \/>\n      hot-fudge-info.html, using the <code>POST<\/code> method.<\/p>\n<div><\/div>\n<\/li>\n<li>\n      The URL requested through <code>POST<\/code>, hot-fudge-info.html, along with its data payload,<br \/>\n      is added to Googlebot&#8217;s crawl queue.\n    <\/li>\n<li>Googlebot performs a <code>POST<\/code> request to crawl hot-fudge-info.html.<\/li>\n<li>\n      Google now has an accurate representation of yummy-sundae.html for Instant Previews. In<br \/>\n      certain cases, we may also incorporate the contents of hot-fudge-info.html into<br \/>\n      yummy-sundae.html.\n    <\/li>\n<li>Google completes the indexing of yummy-sundae.html.<\/li>\n<li>User searches for &#8220;hot fudge sundae&#8221;.<\/li>\n<li>\n      Google&#8217;s algorithms can now better determine how yummy-sundae.html is relevant for this query,<br \/>\n      and we can properly display a snapshot of the page for Instant Previews.\n    <\/li>\n<\/ol>\n<\/ul>\n<h2 id=\"improving-your-sites-crawlability-and-indexability\" tabindex=\"-1\">Improving your site&#8217;s crawlability and indexability<\/h2>\n<p>\n  General advice for creating crawlable sites is found in our<br \/>\n  <a href=\"https:\/\/www.google.com\/support\/webmasters\/bin\/answer.py?answer=40349\" class=\"external-link\">Help Center<\/a>.<br \/>\n  For webmasters who want to help Google crawl and index their content and\/or generate the Instant<br \/>\n  Preview, here are a few simple reminders:\n<\/p>\n<ul>\n<li>\n    Prefer <code>GET<\/code> for fetching resources, unless there&#8217;s a specific reason to use<br \/>\n    <code>POST<\/code>.\n  <\/li>\n<li>\n    Verify that we&#8217;re allowed to crawl the resources needed to render your page. In the example<br \/>\n    above, if hot-fudge-info.html is disallowed by<br \/>\n    <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-txt-intro\/\">robots.txt<\/a>,<br \/>\n    Googlebot won&#8217;t fetch it. More subtly, if the JavaScript code that issues the<br \/>\n    <code>XMLHttpRequest<\/code> is located in an external <code>.js<\/code> file disallowed by<br \/>\n    robots.txt, we won&#8217;t see the connection between yummy-sundae.html and hot-fudge-info.html, so<br \/>\n    even if the latter is not disallowed itself, that may not help us much. We&#8217;ve seen even more<br \/>\n    complicated chains of dependencies in the wild. To help Google better understand your site it&#8217;s<br \/>\n    almost always better to allow Googlebot to crawl all resources.<br \/>\n    You can test whether resources are blocked through<br \/>\n    <a href=\"https:\/\/search.google.com\/search-console\" class=\"external-link\">Webmaster Tools<\/a><br \/>\n    <i>Labs<\/i> <span aria-label=\"and then\">&gt;<\/span><br \/>\n    <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/troubleshooting-instant-previews-in-webmaster-tools\/\"><i>Instant Previews<\/i><\/a>.\n  <\/li>\n<li>\n    Make sure to return the same content to Googlebot as is returned to users&#8217; web browsers.<br \/>\n    <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/spam-policies\/#cloaking\">Cloaking<\/a><br \/>\n    (sending different content to Googlebot than to users) is a violation of our<br \/>\n    <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/search-essentials-overview\/\">Webmaster Guidelines<\/a><br \/>\n    because, among other things, it may cause us to provide a searcher with an irrelevant result<br \/>\n    &mdash;the content the user views in their browser may be a complete mismatch from what we<br \/>\n    crawled and indexed. We&#8217;ve seen numerous <code>POST<\/code> request examples where a webmaster<br \/>\n    non-maliciously cloaked (which is still a violation), and their cloaking&mdash;on even the<br \/>\n    smallest of changes&mdash;then caused JavaScript errors that prevented accurate indexing and<br \/>\n    completely defeated their reason for cloaking in the first place. Summarizing, if you want your<br \/>\n    site to be search-friendly, cloaking is an all-around sticky situation that&#8217;s best to avoid.<br \/>\n    <br \/>\n    To verify that you&#8217;re not accidentally cloaking, you can use<br \/>\n    <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/troubleshooting-instant-previews-in-webmaster-tools\/\">Instant Previews<\/a><br \/>\n    within Webmaster Tools, or try setting the User-Agent string in your browser to something like:<\/p>\n<div><\/div>\n<p>    Your site shouldn&#8217;t look any different after such a change. If you see a blank page, a<br \/>\n    JavaScript error, or if parts of the page are missing or different, that means that something&#8217;s<br \/>\n    wrong.\n  <\/li>\n<li>\n    Remember to include important content (that is, the content you&#8217;d like indexed) as text, visible<br \/>\n    directly on the page and without requiring user-action to display. Most search engines are<br \/>\n    text-based and generally work best with text-based content. We&#8217;re always improving our ability<br \/>\n    to crawl and index content published in a variety of ways, but it remains a good practice to<br \/>\n    use text for important information.\n  <\/li>\n<\/ul>\n<h2 id=\"controlling-your-content\" tabindex=\"-1\">Controlling your content<\/h2>\n<p>\n  If you&#8217;d like to prevent content from being crawled or indexed for Google Web Search, traditional<br \/>\n  <a href=\"https:\/\/alienroad.com\/google-bilgi-bankasi\/robots-txt-syntax\/#syntax\">robots.txt rules<\/a><br \/>\n  remain the best method. To prevent the Instant Preview for your page(s), please see our<br \/>\n  <a href=\"https:\/\/sites.google.com\/site\/webmasterhelpforum\/en\/faq-instant-previews\" class=\"external-link\">Instant Previews FAQ<\/a><br \/>\n  which describes the <code>Google Web Preview<\/code> User-Agent and the <code>nosnippet<\/code> <code>meta<\/code> tag.\n<\/p>\n<h2 id=\"moving-forward\" tabindex=\"-1\">Moving forward<\/h2>\n<p>\n  We&#8217;ll continue striving to increase the comprehensiveness of our index so searchers can find more<br \/>\n  relevant information. And we expect our crawling and indexing capability to improve and evolve<br \/>\n  over time, just like the web itself. Please let us know if you have questions or concerns.\n<\/p>\n<p class=\"byline-author\">\n  Written by<br \/>\n  <a href=\"https:\/\/plus.google.com\/103690467358879664235\/about\" rel=\"author\" class=\"external-link\">Pawel Aleksander Fedorynski<\/a>,<br \/>\n  Software Engineer, Indexing Team, and<br \/>\n  <a href=\"https:\/\/developers.google.com\/search\/blog\/authors\/maile-ohye\" rel=\"author\" class=\"external-link\">Maile Ohye<\/a>,<br \/>\n  Developer Programs Tech Lead<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Tuesday, November 01, 2011 As the web evolves, Google&#8217;s crawling and indexing capabilities also need to progress. We improved our indexing of Flash, built a more robust infrastructure called Caffeine, and we even started crawling forms where it makes sense. Now, especially with the growing popularity of JavaScript and, with it, AJAX, we&#8217;re finding more [&hellip;]<\/p>\n","protected":false},"menu_order":84721,"template":"","meta":{"footnotes":""},"ar_kb_kategori":[665],"ar_kb_etiket":[],"class_list":["post-24620","ar_kb","type-ar_kb","status-publish","has-post-thumbnail","hentry","ar_kb_kategori-blog"],"_links":{"self":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/24620","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb"}],"about":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/types\/ar_kb"}],"version-history":[{"count":0,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb\/24620\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media\/26908"}],"wp:attachment":[{"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/media?parent=24620"}],"wp:term":[{"taxonomy":"ar_kb_kategori","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_kategori?post=24620"},{"taxonomy":"ar_kb_etiket","embeddable":true,"href":"https:\/\/alienroad.com\/wp-json\/wp\/v2\/ar_kb_etiket?post=24620"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}